Team HYU ASML ROBOVOX SP Cup 2024 System Description
This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognition using a mobile robot in noisy and reverberant conditions. Our solution combines the result of deep residual neural networks and time-delay neural network-based speaker embedding models. These models were trained on a diverse dataset that includes French speech. To account for the challenging evaluation environment characterized by high noise, reverberation, and short speech conditions, we focused on data augmentation and training speech duration for the speaker embedding model. Our submission achieved second place on the SP Cup 2024 public leaderboard, with a detection cost function of 0.5245 and an equal error rate of 6.46%.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeaker RecognitionSimilar Papers 제목 키워드 기반
Physics-Informed Inference Time Scaling via Simulation-Calibrated Scientific Machine Learning
High-dimensional partial differential equations (PDEs) pose significant computational challenges across fields ranging from quantum chemistry to economics and finance. Although scientific machine learning (SciML) techniq…
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
Large language models have shown impressive performance in various domains, including code generation across diverse open-source domains. However, their applicability in proprietary industrial settings, where domain-spec…
Code GenerationStructured Prediction for Conditional Meta-Learning
The goal of optimization-based meta-learning is to find a single initialization shared across a distribution of tasks to speed up the process of learning new tasks. Conditional meta-learning seeks task-specific initializ…
Few-Shot LearningMeta-LearningPredictionStructured PredictionAttention-Set based Metric Learning for Video Face Recognition
Face recognition has made great progress with the development of deep learning. However, video face recognition (VFR) is still an ongoing task due to various illumination, low-resolution, pose variations and motion blur.…
Face RecognitionMetric LearningTEAM HUB@LT-EDI-EACL2021: Hope Speech Detection Based On Pre-trained Language Model
This article introduces the system description of TEAM_HUB team participating in LT-EDI 2021: Hope Speech Detection. This shared task is the first task related to the desired voice detection. The data set in the shared t…
Hope Speech DetectionLanguage ModelingLanguage Modellingtext-classification+1