ReMask: A Robust Information-Masking Approach for Domain Counterfactual Generation

Hong, Pengfei; Bhardwaj, Rishabh; Majumdar, Navonil; Aditya, Somak; Poria, Soujanya

Computer Science > Computation and Language

arXiv:2305.02858 (cs)

[Submitted on 4 May 2023]

Title:ReMask: A Robust Information-Masking Approach for Domain Counterfactual Generation

Authors:Pengfei Hong, Rishabh Bhardwaj, Navonil Majumdar, Somak Aditya, Soujanya Poria

View PDF

Abstract:Domain shift is a big challenge in NLP, thus, many approaches resort to learning domain-invariant features to mitigate the inference phase domain shift. Such methods, however, fail to leverage the domain-specific nuances relevant to the task at hand. To avoid such drawbacks, domain counterfactual generation aims to transform a text from the source domain to a given target domain. However, due to the limited availability of data, such frequency-based methods often miss and lead to some valid and spurious domain-token associations. Hence, we employ a three-step domain obfuscation approach that involves frequency and attention norm-based masking, to mask domain-specific cues, and unmasking to regain the domain generic context. Our experiments empirically show that the counterfactual samples sourced from our masked text lead to improved domain transfer on 10 out of 12 domain sentiment classification settings, with an average of 2% accuracy improvement over the state-of-the-art for unsupervised domain adaptation (UDA). Further, our model outperforms the state-of-the-art by achieving 1.4% average accuracy improvement in the adversarial domain adaptation (ADA) setting. Moreover, our model also shows its domain adaptation efficacy on a large multi-domain intent classification dataset where it attains state-of-the-art results. We release the codes publicly at \url{this https URL}.

Comments:	12 pages, 1 figure, 8 tables, ACL 2023 Long Paper (Findings)
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2305.02858 [cs.CL]
	(or arXiv:2305.02858v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2305.02858

Submission history

From: Somak Aditya [view email]
[v1] Thu, 4 May 2023 14:19:02 UTC (2,023 KB)

✅2024-10-01: arxiv.org is back to normal.✅

Computer Science > Computation and Language

Title:ReMask: A Robust Information-Masking Approach for Domain Counterfactual Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

✅2024-10-01: arxiv.org is back to normal.✅

Computer Science > Computation and Language

Title:ReMask: A Robust Information-Masking Approach for Domain Counterfactual Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators