paper-with-me

Speaker Recognition

1개 벤치마크 · 논문 458편 · 이 태스크의 논문 보기 →

Benchmarks

VoxCeleb1

결과 2개

Most implemented

Papers

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

2026-09-09 · Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He 외 arxiv

The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…

Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion Answering

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

2026-08-08 · Tianle Yang, Cuiling Zhang, Chengzhe Sun, Siwei Lyu 외 arxiv

In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the assumption that a person's voice constitutes a stable and unique biom…

Speaker Recognition

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

2026-07-02 · Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu 외 arxiv

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken u…

Speaker Recognition

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

2026-06-23 · Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat 외 arxiv

Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…

Speaker Recognition

Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

2026-06-22 · Yanze Xu, Mark D. Plumbley, Wenwu Wang arxiv

Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of explainable AI (XAI). Analogous to the human …

Speaker Recognition

Speaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

2026-06-07 · Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans arxiv

Sophisticated generative speech technology can undermined the reliability of voice biometrics. While spoofing detection systems excel when assessed under in-domain conditions, generalisation to out-of-domain settings is …

Representation LearningSpeaker Recognition

전체 458편 보기 →