Visually Supervised Speaker Detection and Localization via Microphone Array

Berghi, Davide; Hilton, Adrian; Jackson, Philip J. B.

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2203.03291 (eess)

[Submitted on 7 Mar 2022]

Title:Visually Supervised Speaker Detection and Localization via Microphone Array

Authors:Davide Berghi, Adrian Hilton, Philip J. B. Jackson

View PDF

Abstract:Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of consecutive face crops) and the respective monaural audio. However, their recall rate is often low as only the visible faces are included in the set of candidates. Monaural audio may successfully detect the presence of speech activity but fails in localizing the speaker due to the lack of spatial cues. Our solution extends the audio front-end using a microphone array. We train an audio convolutional neural network (CNN) in combination with beamforming techniques to regress the speaker's horizontal position directly in the video frames. We propose to generate weak labels using a pre-trained active speaker detector on pre-extracted face tracks. Our pipeline embraces the "student-teacher" paradigm, where a trained "teacher" network is used to produce pseudo-labels visually. The "student" network is an audio network trained to generate the same results. At inference, the student network can independently localize the speaker in the visual frames directly from the audio input. Experimental results on newly collected data prove that our approach significantly outperforms a variety of other baselines as well as the teacher network itself. It results in an excellent speech activity detector too.

Comments:	Erratum: Due to a bug in the evaluation script, the correct average distance (aD) metric is here reported in yellow. The analysis remains unchanged from the original paper as the trend between the old and new measures are perfectly monotonic. The bug was caused by an incorrect normalization factor
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD); Image and Video Processing (eess.IV)
Cite as:	arXiv:2203.03291 [eess.AS]
	(or arXiv:2203.03291v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2203.03291
Journal reference:	IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), 2021

Submission history

From: Davide Berghi Mr [view email]
[v1] Mon, 7 Mar 2022 11:12:39 UTC (26,228 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Visually Supervised Speaker Detection and Localization via Microphone Array

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Visually Supervised Speaker Detection and Localization via Microphone Array

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators