SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization

Mhamed, Jaafar; Gu, Shangding

Computer Science > Machine Learning

arXiv:2311.00880 (cs)

[Submitted on 1 Nov 2023]

Title:SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization

Authors:Jaafar Mhamed, Shangding Gu

View PDF

Abstract:Incorporating safety is an essential prerequisite for broadening the practical applications of reinforcement learning in real-world scenarios. To tackle this challenge, Constrained Markov Decision Processes (CMDPs) are leveraged, which introduce a distinct cost function representing safety violations. In CMDPs' settings, Lagrangian relaxation technique has been employed in previous algorithms to convert constrained optimization problems into unconstrained dual problems. However, these algorithms may inaccurately predict unsafe behavior, resulting in instability while learning the Lagrange multiplier. This study introduces a novel safe reinforcement learning algorithm, Safety Critic Policy Optimization (SCPO). In this study, we define the safety critic, a mechanism that nullifies rewards obtained through violating safety constraints. Furthermore, our theoretical analysis indicates that the proposed algorithm can automatically balance the trade-off between adhering to safety constraints and maximizing rewards. The effectiveness of the SCPO algorithm is empirically validated by benchmarking it against strong baselines.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2311.00880 [cs.LG]
	(or arXiv:2311.00880v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2311.00880

Submission history

From: Shangding Gu [view email]
[v1] Wed, 1 Nov 2023 22:12:50 UTC (2,011 KB)

Computer Science > Machine Learning

Title:SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators