No-Regret Exploration in Goal-Oriented Reinforcement Learning

Tarbouriech, Jean; Garcelon, Evrard; Valko, Michal; Pirotta, Matteo; Lazaric, Alessandro

Statistics > Machine Learning

arXiv:1912.03517v1 (stat)

[Submitted on 7 Dec 2019 (this version), latest version 17 Aug 2020 (v3)]

Title:No-Regret Exploration in Goal-Oriented Reinforcement Learning

Authors:Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, Alessandro Lazaric

View PDF

Abstract:Many popular reinforcement learning problems (e.g., navigation in a maze, some Atari games, mountain car) are instances of the so-called episodic setting or stochastic shortest path (SSP) problem, where an agent has to achieve a predefined goal state (e.g., the top of the hill) while maximizing the cumulative reward or minimizing the cumulative cost. Despite its popularity, most of the literature studying the exploration-exploitation dilemma either focused on different problems (i.e., fixed-horizon and infinite-horizon) or made the restrictive loop-free assumption (which implies that no same state can be visited twice during any episode). In this paper, we study the general SSP setting and introduce the algorithm UC-SSP whose regret scales as $\displaystyle \widetilde{O}(c_{\max}^{3/2} c_{\min}^{-1/2} D S \sqrt{ A D K})$ after $K$ episodes for any unknown SSP with $S$ non-terminal states, $A$ actions, an SSP-diameter of $D$ and positive costs in $[c_{\min}, c_{\max}]$. UC-SSP is thus the first learning algorithm with vanishing regret in the theoretically challenging setting of episodic RL.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:1912.03517 [stat.ML]
	(or arXiv:1912.03517v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1912.03517

Submission history

From: Jean Tarbouriech [view email]
[v1] Sat, 7 Dec 2019 15:19:22 UTC (214 KB)
[v2] Thu, 30 Jan 2020 18:23:10 UTC (206 KB)
[v3] Mon, 17 Aug 2020 11:46:12 UTC (2,212 KB)

Statistics > Machine Learning

Title:No-Regret Exploration in Goal-Oriented Reinforcement Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:No-Regret Exploration in Goal-Oriented Reinforcement Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators