Speaker Identification
4개 벤치마크 · 논문 268편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Speaker Recognition from Raw Waveform with SincNet
Deep Speaker: an End-to-End Neural Speaker Embedding System
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
Masked Autoencoders that Listen
ATST: Audio Representation Learning with Teacher-Student Transformer
Papers
AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification
Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions. In practical scenarios, background multi-speak…
Speaker IdentificationMAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content. However, their performance in sentiment analysis exhibits acute sensitivity to prompt design…
Speaker IdentificationSentiment AnalysisMultimodal Speaker Identification in Classroom Environments
Automated analysis of K-12 classroom dynamics faces challenges due to background noise and variable child speech, often confounding acoustic-only models. This study evaluates a multimodal speaker identification framework…
Speaker IdentificationSpeaker Group Encoding in Self-supervised Speech Recognition Models
We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identification (SID), finetuned on automatic spee…
Speaker IdentificationSpeech RecognitionOmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, au…
Sign Language RecognitionComputational EfficiencySpeaker IdentificationWhere Do Self-Supervised Speech Models Become Unfair?
Speech encoder models are known to model members of some speaker groups (SGs) better than others. However, there has been little work in establishing why this occurs on a technological level. To our knowledge, we present…
Speaker IdentificationSpeech Recognition