To Tune or Not To Tune? How About the Best of Both Worlds?

Wang, Ran; Su, Haibo; Wang, Chunye; Ji, Kailin; Ding, Jupeng

Computer Science > Computation and Language

arXiv:1907.05338 (cs)

[Submitted on 9 Jul 2019]

Title:To Tune or Not To Tune? How About the Best of Both Worlds?

Authors:Ran Wang, Haibo Su, Chunye Wang, Kailin Ji, Jupeng Ding

View PDF

Abstract:The introduction of pre-trained language models has revolutionized natural language research communities. However, researchers still know relatively little regarding their theoretical and empirical properties. In this regard, Peters et al. perform several experiments which demonstrate that it is better to adapt BERT with a light-weight task-specific head, rather than building a complex one on top of the pre-trained language model, and freeze the parameters in the said language model. However, there is another option to adopt. In this paper, we propose a new adaptation method which we first train the task model with the BERT parameters frozen and then fine-tune the entire model together. Our experimental results show that our model adaptation method can achieve 4.7% accuracy improvement in semantic similarity task, 0.99% accuracy improvement in sequence labeling task and 0.72% accuracy improvement in the text classification task.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:1907.05338 [cs.CL]
	(or arXiv:1907.05338v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1907.05338

Submission history

From: Ran Wang [view email]
[v1] Tue, 9 Jul 2019 04:46:31 UTC (50 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2019-07

Change to browse by:

cs
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Ran Wang
Haibo Su
Chunye Wang
Kailin Ji
Jupeng Ding

export BibTeX citation

Computer Science > Computation and Language

Title:To Tune or Not To Tune? How About the Best of Both Worlds?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:To Tune or Not To Tune? How About the Best of Both Worlds?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators