Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain

Shon, Suwon; Ali, Ahmed; Glass, James

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:1812.01501 (eess)

[Submitted on 4 Dec 2018 (v1), last revised 6 May 2019 (this version, v2)]

Title:Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain

Authors:Suwon Shon, Ahmed Ali, James Glass

View PDF

Abstract:End-to-end deep learning language or dialect identification systems operate on the spectrogram or other acoustic feature and directly generate identification scores for each class. An important issue for end-to-end systems is to have some knowledge of the application domain, because the system can be vulnerable to use cases that were not seen in the training phase; such a scenario is often referred to as a domain mismatched condition. In general, we assume that there is enough variation in the training dataset to expose the system to multiple domains. In this work, we study how to best make use a training dataset in order to have maximum effectiveness on unknown target domains. Our goal is to process the input without any knowledge of the target domain while preserving robust performance on other domains as well. To accomplish this objective, we propose a domain attentive fusion approach for end-to-end dialect/language identification systems. To help with experimentation, we collect a dataset from three different domains, and create experimental protocols for a domain mismatched condition. The results of our proposed approach, which were tested on a variety of broadcast and YouTube data, shows significant performance gain compared to traditional approaches, even without any prior target domain information.

Comments:	ICASSP 2019, revised typos
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:1812.01501 [eess.AS]
	(or arXiv:1812.01501v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.1812.01501

Submission history

From: Suwon Shon [view email]
[v1] Tue, 4 Dec 2018 16:05:35 UTC (443 KB)
[v2] Mon, 6 May 2019 15:36:50 UTC (495 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators