paper-with-me

홈 › Papers

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation

2026-06-30 · Dominika Woszczyk, Andreas Triantafyllopoulos, Jura Miniota, Éva Székely, Bjoern Schuller arxiv

Text-to-speech (TTS) evaluation is an open challenge. While the primary target was "naturalness," recent fidelity gains shifted focus toward "appropriateness" and whether speech is correct for its context. In this work, we examine how perception changes when the expected downstream use varies. We measure the appropriateness and human-likeness of five SOTA TTS systems across five domains: AI assistant, reader, actor, animated character, and spontaneous speaker. Results show appropriateness varies across domains independently of naturalness. While systems shine at reading, expressive domains remain challenging, and optimizing for one can degrade others. Furthermore, naturalness scores tend to penalize stylized speech while rewarding spontaneity. Finally, our study also highlights blind spots in one-size-fits-all evaluation metrics across more expressive domains. We demonstrate that TTS performance is not "solved" but depends on the target domain, requiring context-aware evaluation.

📄 PDF Abstract BibTeX arXiv:2606.31729

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features

2022-11-01 · Alexandra Vioni, Georgia Maniati, Nikolaos Ellinas, June Sig Sung 외

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS t…

POSPredictionSelf-Supervised Learningtext-to-speech+1

Rethinking Sentiment Style Transfer

2021-11-01 · Findings (EMNLP) 2021 11 · Ping Yu, Yang Zhao, Chunyuan Li, Changyou Chen

Though remarkable efforts have been made in non-parallel text style transfer, the evaluation system is unsatisfactory. It always evaluates over samples from only one checkpoint of the model and compares three metrics, i.…

AttributeStyle TransferText Style Transfer

Multilingual Dialogue Generation and Localization with Dialogue Act Scripting

2025-09-26 · Justin Vasselli, Eunike Andriani Kardinata, Yusuke Sakai, Taro Watanabe arxiv

Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that reduce their naturalness and cultural appr…

Dialogue Generation

Prompt Guided Copy Mechanism for Conversational Question Answering

2023-08-07 · Yong Zhang, Zhitao Li, Jianzong Wang, Yiming Gao 외

Conversational Question Answering (CQA) is a challenging task that aims to generate natural answers for conversational flow questions. In this paper, we propose a pluggable approach for extractive methods that introduces…

Conversational Question AnsweringQuestion Answering

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

2026-07-06 · Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar 외 arxiv

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations,…