Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

RL Theory

Runlong Zhou*, Zihan Zhang*, Maryam Fazel, Simon S. Du

We provide the first completely horizon-free and asymptotically optimal regret bound for time-homogeneous MDPs.

Abstract