Pre-training for low resource speech-to-intent applications

Wang, Pu; Van hamme, Hugo

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2103.16674 (eess)

[Submitted on 30 Mar 2021]

Title:Pre-training for low resource speech-to-intent applications

Authors:Pu Wang, Hugo Van hamme

View PDF

Abstract:Designing a speech-to-intent (S2I) agent which maps the users' spoken commands to the agents' desired task actions can be challenging due to the diverse grammatical and lexical preference of different users. As a remedy, we discuss a user-taught S2I system in this paper. The user-taught system learns from scratch from the users' spoken input with action demonstration, which ensure it is fully matched to the users' way of formulating intents and their articulation habits. The main issue is the scarce training data due to the user effort involved. Existing state-of-art approaches in this setting are based on non-negative matrix factorization (NMF) and capsule networks. In this paper we combine the encoder of an end-to-end ASR system with the prior NMF/capsule network-based user-taught decoder, and investigate whether pre-training methodology can reduce training data requirements for the NMF and capsule network. Experimental results show the pre-trained ASR-NMF framework significantly outperforms other models, and also, we discuss limitations of pre-training with different types of command-and-control(C&C) applications.

Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Sound (cs.SD)
Cite as:	arXiv:2103.16674 [eess.AS]
	(or arXiv:2103.16674v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2103.16674

Submission history

From: Pu Wang [view email]
[v1] Tue, 30 Mar 2021 20:44:29 UTC (478 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Pre-training for low resource speech-to-intent applications

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Pre-training for low resource speech-to-intent applications

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators