Traffic Light Control Using Deep Policy-Gradient and Value-Function Based Reinforcement Learning

Mousavi, Seyed Sajad; Schukat, Michael; Howley, Enda

Computer Science > Machine Learning

arXiv:1704.08883 (cs)

[Submitted on 28 Apr 2017 (v1), last revised 27 May 2017 (this version, v2)]

Title:Traffic Light Control Using Deep Policy-Gradient and Value-Function Based Reinforcement Learning

Authors:Seyed Sajad Mousavi, Michael Schukat, Enda Howley

View PDF

Abstract:Recent advances in combining deep neural network architectures with reinforcement learning techniques have shown promising potential results in solving complex control problems with high dimensional state and action spaces. Inspired by these successes, in this paper, we build two kinds of reinforcement learning algorithms: deep policy-gradient and value-function based agents which can predict the best possible traffic signal for a traffic intersection. At each time step, these adaptive traffic light control agents receive a snapshot of the current state of a graphical traffic simulator and produce control signals. The policy-gradient based agent maps its observation directly to the control signal, however the value-function based agent first estimates values for all legal control signals. The agent then selects the optimal control action with the highest value. Our methods show promising results in a traffic network simulated in the SUMO traffic simulator, without suffering from instability issues during the training process.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:1704.08883 [cs.LG]
	(or arXiv:1704.08883v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1704.08883

Submission history

From: Seyed Sajad Mousavi [view email]
[v1] Fri, 28 Apr 2017 11:44:42 UTC (128 KB)
[v2] Sat, 27 May 2017 14:45:56 UTC (128 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2017-04

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Seyed Sajad Mousavi
Michael Schukat
Peter Corcoran
Enda Howley

export BibTeX citation

Computer Science > Machine Learning

Title:Traffic Light Control Using Deep Policy-Gradient and Value-Function Based Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Traffic Light Control Using Deep Policy-Gradient and Value-Function Based Reinforcement Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators