Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

Du, Zhihao; Zhang, Shiliang; Zheng, Siqi; Yan, Zhijie

Computer Science > Sound

arXiv:2211.10243 (cs)

[Submitted on 18 Nov 2022]

Title:Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

Authors:Zhihao Du, Shiliang Zhang, Siqi Zheng, Zhijie Yan

View PDF

Abstract:Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped speaker diarization as a multi-label classification problem, where speaker dependency and overlaps are not well considered. To overcome the disadvantages, we reformulate overlapped speaker diarization task as a single-label prediction problem via the proposed power set encoding (PSE). Through this formulation, speaker dependency and overlaps can be explicitly modeled. To fully leverage this formulation, we further propose the speaker overlap-aware neural diarization (SOND) model, which consists of a context-independent (CI) scorer to model global speaker discriminability, a context-dependent scorer (CD) to model local discriminability, and a speaker combining network (SCN) to combine and reassign speaker activities. Experimental results show that using the proposed formulation can outperform the state-of-the-art methods based on target speaker voice activity detection, and the performance can be further improved with SOND, resulting in a 6.30% relative diarization error reduction.

Comments:	Accepted by EMNLP 2022
Subjects:	Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2211.10243 [cs.SD]
	(or arXiv:2211.10243v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2211.10243

Submission history

From: Zhihao Du [view email]
[v1] Fri, 18 Nov 2022 14:03:26 UTC (1,165 KB)

Computer Science > Sound

Title:Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators