Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence
RL Theory
Runlong Zhou*, Zihan Zhang*, Maryam Fazel, Simon S. Du
We provide the first completely horizon-free and asymptotically optimal regret bound for time-homogeneous MDPs.