Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations

Sequeira, Pedro; Gervasio, Melinda

doi:10.1016/j.artint.2020.103367

Computer Science > Machine Learning

arXiv:1912.09007 (cs)

[Submitted on 19 Dec 2019 (v1), last revised 19 Aug 2020 (this version, v2)]

Title:Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations

Authors:Pedro Sequeira, Melinda Gervasio

View PDF

Abstract:We propose an explainable reinforcement learning (XRL) framework that analyzes an agent's history of interaction with the environment to extract interestingness elements that help explain its behavior. The framework relies on data readily available from standard RL algorithms, augmented with data that can easily be collected by the agent while learning. We describe how to create visual summaries of an agent's behavior in the form of short video-clips highlighting key interaction moments, based on the proposed elements. We also report on a user study where we evaluated the ability of humans to correctly perceive the aptitude of agents with different characteristics, including their capabilities and limitations, given visual summaries automatically generated by our framework. The results show that the diversity of aspects captured by the different interestingness elements is crucial to help humans correctly understand an agent's strengths and limitations in performing a task, and determine when it might need adjustments to improve its performance.

Comments:	To appear in: Artificial Intelligence
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (stat.ML)
Cite as:	arXiv:1912.09007 [cs.LG]
	(or arXiv:1912.09007v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1912.09007
Related DOI:	https://doi.org/10.1016/j.artint.2020.103367

Submission history

From: Pedro Sequeira [view email]
[v1] Thu, 19 Dec 2019 03:46:22 UTC (1,974 KB)
[v2] Wed, 19 Aug 2020 03:25:14 UTC (2,022 KB)

Computer Science > Machine Learning

Title:Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators