Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Xu, Weijia; Agrawal, Sweta; Briakou, Eleftheria; Martindale, Marianna J.; Carpuat, Marine

Computer Science > Computation and Language

arXiv:2301.07779 (cs)

[Submitted on 18 Jan 2023 (v1), last revised 25 Feb 2023 (this version, v2)]

Title:Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Authors:Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat

View PDF

Abstract:Neural sequence generation models are known to "hallucinate", by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds.

Comments:	Accepted at TACL
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2301.07779 [cs.CL]
	(or arXiv:2301.07779v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2301.07779

Submission history

From: Weijia Xu [view email]
[v1] Wed, 18 Jan 2023 20:43:13 UTC (192 KB)
[v2] Sat, 25 Feb 2023 04:23:12 UTC (192 KB)

Computer Science > Computation and Language

Title:Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators