Speaker Recognition
1개 벤치마크 · 논문 458편 · 이 태스크의 논문 보기 →
Benchmarks
VoxCeleb1
Most implemented
Speaker Recognition from Raw Waveform with SincNet
Deep Speaker: an End-to-End Neural Speaker Embedding System
Utterance-level Aggregation For Speaker Recognition In The Wild
Papers
SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…
Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion AnsweringThe Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints
In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the assumption that a person's voice constitutes a stable and unique biom…
Speaker RecognitionReasoning LLM Improves Speaker Recognition in Long-form TV Dramas
Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken u…
Speaker RecognitionVieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency
Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…
Speaker RecognitionExplainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of explainable AI (XAI). Analogous to the human …
Speaker RecognitionSpeaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck
Sophisticated generative speech technology can undermined the reliability of voice biometrics. While spoofing detection systems excel when assessed under in-domain conditions, generalisation to out-of-domain settings is …
Representation LearningSpeaker Recognition