Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization

Kalweit, Gabriel; Huegle, Maria; Boedecker, Joschka

Computer Science > Machine Learning

arXiv:1909.13518 (cs)

[Submitted on 30 Sep 2019 (v1), last revised 14 Aug 2020 (this version, v2)]

Title:Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization

Authors:Gabriel Kalweit, Maria Huegle, Joschka Boedecker

View PDF

Abstract:In the past few years, off-policy reinforcement learning methods have shown promising results in their application for robot control. Deep Q-learning, however, still suffers from poor data-efficiency and is susceptible to stochasticity in the environment or reward functions which is limiting with regard to real-world applications. We alleviate these problems by proposing two novel off-policy Temporal-Difference formulations: (1) Truncated Q-functions which represent the return for the first n steps of a target-policy rollout w.r.t. the full action-value and (2) Shifted Q-functions, acting as the farsighted return after this truncated rollout. This decomposition allows us to optimize both parts with their individual learning rates, achieving significant learning speedup. We prove that the combination of these short- and long-term predictions is a representation of the full return, leading to the Composite Q-learning algorithm. We show the efficacy of Composite Q-learning in the tabular case and compare Deep Composite Q-learning with TD3 and TD3(Delta), which we introduce as an off-policy variant of TD(Delta). Moreover, we show that Composite TD3 outperforms TD3 as well as state-of-the-art compositional Q-learning approaches significantly in terms of data-efficiency in multiple simulated robot tasks and that Composite Q-learning is robust to stochastic environments and reward functions.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:1909.13518 [cs.LG]
	(or arXiv:1909.13518v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1909.13518

Submission history

From: Gabriel Kalweit [view email]
[v1] Mon, 30 Sep 2019 08:40:09 UTC (1,094 KB)
[v2] Fri, 14 Aug 2020 08:32:55 UTC (4,640 KB)

Computer Science > Machine Learning

Title:Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators