Next Article in Journal
Compact Broadband Antenna with Vicsek Fractal Slots for WLAN and WiMAX Applications
Next Article in Special Issue
Non-Parallel Articulatory-to-Acoustic Conversion Using Multiview-Based Time Warping
Previous Article in Journal
Molecular Dynamics of Solids at Constant Pressure and Stress Using Anisotropic Stochastic Cell Rescaling
Previous Article in Special Issue
Cascade or Direct Speech Translation? A Case Study
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Multimodal Diarization Systems by Training Enrollment Models as Identity Representations †

ViVoLab, Aragón Institute for Engineering Research (I3A), University of Zaragoza, 50018 Zaragoza, Spain
*
Authors to whom correspondence should be addressed.
This paper is an extended version of our paper published in the conference IberSPEECH2020.
Appl. Sci. 2022, 12(3), 1141; https://doi.org/10.3390/app12031141
Submission received: 23 December 2021 / Revised: 17 January 2022 / Accepted: 19 January 2022 / Published: 21 January 2022

Abstract

This paper describes a post-evaluation analysis of the system developed by ViVoLAB research group for the IberSPEECH-RTVE 2020 Multimodal Diarization (MD) Challenge. This challenge focuses on the study of multimodal systems for the diarization of audiovisual files and the assignment of an identity to each segment where a person is detected. In this work, we implemented two different subsystems to address this task using the audio and the video from audiovisual files separately. To develop our subsystems, we used the state-of-the-art speaker and face verification embeddings extracted from publicly available deep neural networks (DNN). Different clustering techniques were also employed in combination with the tracking and identity assignment process. Furthermore, we included a novel back-end approach in the face verification subsystem to train an enrollment model for each identity, which we have previously shown to improve the results compared to the average of the enrollment data. Using this approach, we trained a learnable vector to represent each enrollment character. The loss function employed to train this vector was an approximated version of the detection cost function (aDCF) which is inspired by the DCF widely used metric to measure performance in verification tasks. In this paper, we also focused on exploring and analyzing the effect of training this vector with several configurations of this objective loss function. This analysis allows us to assess the impact of the configuration parameters of the loss in the amount and type of errors produced by the system.
Keywords: enrollment models; face recognition; aDCF loss; speaker recognition; deep neural networks; spectral clustering; video processing enrollment models; face recognition; aDCF loss; speaker recognition; deep neural networks; spectral clustering; video processing

Share and Cite

MDPI and ACS Style

Mingote, V.; Viñals, I.; Gimeno, P.; Miguel, A.; Ortega, A.; Lleida, E. Multimodal Diarization Systems by Training Enrollment Models as Identity Representations. Appl. Sci. 2022, 12, 1141. https://doi.org/10.3390/app12031141

AMA Style

Mingote V, Viñals I, Gimeno P, Miguel A, Ortega A, Lleida E. Multimodal Diarization Systems by Training Enrollment Models as Identity Representations. Applied Sciences. 2022; 12(3):1141. https://doi.org/10.3390/app12031141

Chicago/Turabian Style

Mingote, Victoria, Ignacio Viñals, Pablo Gimeno, Antonio Miguel, Alfonso Ortega, and Eduardo Lleida. 2022. "Multimodal Diarization Systems by Training Enrollment Models as Identity Representations" Applied Sciences 12, no. 3: 1141. https://doi.org/10.3390/app12031141

APA Style

Mingote, V., Viñals, I., Gimeno, P., Miguel, A., Ortega, A., & Lleida, E. (2022). Multimodal Diarization Systems by Training Enrollment Models as Identity Representations. Applied Sciences, 12(3), 1141. https://doi.org/10.3390/app12031141

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop