Automatic Correction of Human Translations

Lin, Jessy; Kovacs, Geza; Shastry, Aditya; Wuebker, Joern; DeNero, John

Computer Science > Computation and Language

arXiv:2206.08593 (cs)

[Submitted on 17 Jun 2022]

Title:Automatic Correction of Human Translations

Authors:Jessy Lin, Geza Kovacs, Aditya Shastry, Joern Wuebker, John DeNero

View PDF

Abstract:We introduce translation error correction (TEC), the task of automatically correcting human-generated translations. Imperfections in machine translations (MT) have long motivated systems for improving translations post-hoc with automatic post-editing. In contrast, little attention has been devoted to the problem of automatically correcting human translations, despite the intuition that humans make distinct errors that machines would be well-suited to assist with, from typos to inconsistencies in translation conventions. To investigate this, we build and release the Aced corpus with three TEC datasets. We show that human errors in TEC exhibit a more diverse range of errors and far fewer translation fluency errors than the MT errors in automatic post-editing datasets, suggesting the need for dedicated TEC models that are specialized to correct human errors. We show that pre-training instead on synthetic errors based on human errors improves TEC F-score by as much as 5.1 points. We conducted a human-in-the-loop user study with nine professional translation editors and found that the assistance of our TEC system led them to produce significantly higher quality revised translations.

Comments:	NAACL 2022. Dataset available at: this https URL
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2206.08593 [cs.CL]
	(or arXiv:2206.08593v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2206.08593

Submission history

From: Jessy Lin [view email]
[v1] Fri, 17 Jun 2022 07:30:55 UTC (6,653 KB)

Computer Science > Computation and Language

Title:Automatic Correction of Human Translations

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Automatic Correction of Human Translations

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators