paper-with-me

Speaker Identification

4개 벤치마크 · 논문 268편 · 이 태스크의 논문 보기 →

Benchmarks

VoxCeleb1

결과 12개

EVI en-GB

결과 1개

EVI fr-FR

결과 1개

EVI pl-PL

결과 1개

Most implemented

Masked Autoencoders that Listen

2022-07-13 · 구현 4개

Papers

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

2026-06-28 · Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang 외 arxiv

Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions. In practical scenarios, background multi-speak…

Speaker Identification

MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs

2026-06-14 · Hangling Xie arxiv

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content. However, their performance in sentiment analysis exhibits acute sensitivity to prompt design…

Speaker IdentificationSentiment Analysis

Multimodal Speaker Identification in Classroom Environments

2026-06-10 · Michael L. Chrzan, Meghavarshini Krishnaswamy, Robert Gibboni, Katie Wetstone 외 arxiv

Automated analysis of K-12 classroom dynamics faces challenges due to background noise and variable child speech, often confounding acoustic-only models. This study evaluates a multimodal speaker identification framework…

Speaker Identification

Speaker Group Encoding in Self-supervised Speech Recognition Models

2026-06-09 · Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre, François Portet arxiv

We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identification (SID), finetuned on automatic spee…

Speaker IdentificationSpeech Recognition

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

2026-05-02 · Detao Bai, Shimin Yao, Weixuan Chen, Chengen Lai 외 arxiv

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, au…

Sign Language RecognitionComputational EfficiencySpeaker Identification

Where Do Self-Supervised Speech Models Become Unfair?

2026-04-20 · Felix Herron, Maja Hjuler, Solange Rossato, Alexandre Allauzen 외 arxiv

Speech encoder models are known to model members of some speaker groups (SGs) better than others. However, there has been little work in establishing why this occurs on a technological level. To our knowledge, we present…

Speaker IdentificationSpeech Recognition

전체 268편 보기 →