Offline RL for Natural Language Generation with Implicit Language Q Learning

Snell, Charlie; Kostrikov, Ilya; Su, Yi; Yang, Mengjiao; Levine, Sergey

Computer Science > Computation and Language

arXiv:2206.11871 (cs)

[Submitted on 5 Jun 2022 (v1), last revised 1 May 2023 (this version, v2)]

Title:Offline RL for Natural Language Generation with Implicit Language Q Learning

Authors:Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, Sergey Levine

View PDF

Abstract:Large language models distill broad knowledge from text corpora. However, they can be inconsistent when it comes to completing user specified tasks. This issue can be addressed by finetuning such models via supervised learning on curated datasets, or via reinforcement learning. In this work, we propose a novel offline RL method, implicit language Q-learning (ILQL), designed for use on language models, that combines both the flexible utility maximization framework of RL algorithms with the ability of supervised learning to leverage previously collected data, as well as its simplicity and stability. Our method employs a combination of value conservatism alongside an implicit dataset support constraint in learning value functions, which are then used to guide language model generations towards maximizing user-specified utility functions. In addition to empirically validating ILQL, we present a detailed empirical analysis of situations where offline RL can be useful in natural language generation settings, demonstrating how it can be a more effective utility optimizer than prior approaches for end-to-end dialogue, and how it can effectively optimize high variance reward functions based on subjective judgement, such as whether to label a comment as toxic or not.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2206.11871 [cs.CL]
	(or arXiv:2206.11871v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2206.11871

Submission history

From: Charlie Snell [view email]
[v1] Sun, 5 Jun 2022 18:38:42 UTC (1,185 KB)
[v2] Mon, 1 May 2023 04:42:27 UTC (1,224 KB)

Computer Science > Computation and Language

Title:Offline RL for Natural Language Generation with Implicit Language Q Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Offline RL for Natural Language Generation with Implicit Language Q Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators