paper-with-me

Papers

Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment

2025-10-23 · Zhiyu Lin, Jingwen Yang, Jiale Zhao, Meng Liu, Sunzhu Li, Benyou Wang arxiv

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic features, and emotion recognition are costly, limited, or incomplete. To address this, we present DeEAR (Decoding the Expressive Preference of eAR), a framework that converts human preference for speech expressiveness into an objective score. Grounded in phonetics and psychology, DeEAR evaluates speech across three dimensions: Emotion, Prosody, and Spontaneity, achieving strong alignment with human perception (Spearman's Rank Correlation Coefficient, SRCC = 0.86) using fewer than 500 annotated samples. Beyond reliable scoring, DeEAR enables fair benchmarking and targeted data curation. It not only distinguishes expressiveness gaps across S2S models but also selects 14K expressive utterances to form ExpressiveSpeech, which improves the expressive score (from 2.0 to 23.4 on a 100-point scale) of S2S models. Demos and codes are available at https://github.com/FreedomIntelligence/ExpressiveSpeech

📄 PDF Abstract BibTeX arXiv:2510.20513

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognition

Similar Papers 제목 키워드 기반

Detecting Objectifying Language in Online Professor Reviews

2020-10-16 · EMNLP (WNUT) 2020 11 · Angie Waller, Kyle Gorman

Student reviews often make reference to professors' physical appearances. Until recently RateMyProfessors.com, the website of this study's focus, used a design feature to encourage a "hot or not" rating of college profes…

Drift: Decoding-time Personalized Alignments with Implicit User Preferences

2025-02-20 · Minbeom Kim, Kang-il Lee, Seongho Joo, Hwaran Lee 외

Personalized alignments for individual users have been a long-standing goal in large language models (LLMs). We introduce Drift, a novel framework that personalizes LLMs at decoding time with implicit user preferences. T…

Objectifying Women? A Syntactic Bias in French and English Corpora.

2022-06-01 · LATERAISSE (LREC) 2022 6 · Yanis da Cunha, Anne Abeillé

Gender biases in syntax have been documented for languages with grammatical gender for cases where mixed-gender coordination structures take masculine agreement, or with male-first preference in the ordering of pairs (Ad…

HAPI: A Model for Learning Robot Facial Expressions from Human Preferences

2025-03-21 · Dongsheng Yang, Qianying Liu, Wataru Sato, Takashi Minato 외

Automatic robotic facial expression generation is crucial for human-robot interaction, as handcrafted methods based on fixed joint configurations often yield rigid and unnatural behaviors. Although recent automated techn…

Bayesian OptimizationFacial expression generationLearning-To-Rank

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models

2025-08-26 · Chenxu Yang, Qingyi Si, Zheng Lin arxiv

Grounding responses in external knowledge represents an effective strategy for mitigating hallucinations in Large Language Models (LLMs). However, current LLMs struggle to seamlessly integrate knowledge while simultaneou…