paper-with-me

홈 › Papers

Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness

2025-04-26 · Erfan Loweimi, Mengjie Qian, Kate Knill, Mark Gales

There is a growing abundance of publicly available or company-owned audio/video archives, highlighting the increasing importance of efficient access to desired content and information retrieval from these archives. This paper investigates the challenges, solutions, effectiveness, and robustness of speaker retrieval systems developed "in the wild" which involves addressing two primary challenges: extraction of task-relevant labels from limited metadata for system development and evaluation, as well as the unconstrained acoustic conditions encountered in the archive, ranging from quiet studios to adverse noisy environments. While we focus on the publicly-available BBC Rewind archive (spanning 1948 to 1979), our framework addresses the broader issue of speaker retrieval on extensive and possibly aged archives with no control over the content and acoustic conditions. Typically, these archives offer a brief and general file description, mostly inadequate for specific applications like speaker retrieval, and manual annotation of such large-scale archives is unfeasible. We explore various aspects of system development (e.g., speaker diarisation, embedding extraction, query selection) and analyse the challenges, possible solutions, and their functionality. To evaluate the performance, we conduct systematic experiments in both clean setup and against various distortions simulating real-world applications. Our findings demonstrate the effectiveness and robustness of the developed speaker retrieval systems, establishing the versatility and scalability of the proposed framework for a wide range of applications beyond the BBC Rewind corpus.

📄 PDF Abstract BibTeX arXiv:2504.18950

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

2026-03-17 · Liangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang 외 arxiv

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore op…

Speaker Diarization

AVA-AVD: Audio-Visual Speaker Diarization in the Wild

2021-11-29 · Eric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 외

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news …

Relation Networkspeaker-diarizationSpeaker Diarization

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

2024-09-02 · Jeong Hun Yeo, Chae Won Kim, Hyunjun Kim, Hyeongseop Rha 외

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variat…

Lip ReadingSentence

Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition

2023-09-21 · Shuai Wang, Qibing Bai, Qi Liu, Jianwei Yu 외

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to tr…

Speaker Recognition

WASD: A Wilder Active Speaker Detection Dataset

2023-03-09 · Tiago Roxo, Joana C. Costa, Pedro R. M. Inácio, Hugo Proença

Current Active Speaker Detection (ASD) models achieve great results on AVA-ActiveSpeaker (AVA), using only sound and facial features. Although this approach is applicable in movie setups (AVA), it is not suited for less …

Active Speaker Detection