LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations
Reliable evaluation of phrase break annotations is crucial, as subtle variations in prosodic boundaries directly affect the clarity and naturalness of speech. However, existing approaches exhibit major limitations: single-reference evaluation assumes a unique gold phrasing for an utterance despite multiple valid phrasings, while human judgment, though flexible, is labor-intensive and unscalable. To address these, we propose LLM-based Multi-Reference Evaluation (LMRE) for phrase break annotations that models the one-to-many nature of prosodic phrasing and generates multiple valid phrasings from minimal demonstrations. On a Korean testbed of 1,356 annotations covering five strategies, LMRE shows stronger alignment with human judgment than single-reference evaluation in both acceptance behavior and score correlation. Our findings demonstrate that LMRE effectively achieves both scalability and multi-reference support, highlighting the potential of LLMs for evaluation in the speech domain.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Assessing Phrase Break of ESL Speech with Pre-trained Language Models and Large Language Models
This work introduces approaches to assessing phrase breaks in ESL learners' speech using pre-trained language models (PLMs) and large language models (LLMs). There are two tasks: overall assessment of phrase break for a …
text-classificationText ClassificationAssessing Phrase Break of ESL speech with Pre-trained Language Models
This work introduces an approach to assessing phrase break in ESL learners' speech with pre-trained language models (PLMs). Different with traditional methods, this proposal converts speech to token sequences, and then l…
text-classificationText ClassificationAn investigation of phrase break prediction in an End-to-End TTS system
Purpose: This work explores the use of external phrase break prediction models to enhance listener comprehension in End-to-End Text-to-Speech (TTS) systems. Methods: The effectiveness of these models is evaluated based o…
Predictiontext-to-speechText to SpeechHICEScore: A Hierarchical Metric for Image Captioning Evaluation
Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant …
DescriptiveImage CaptioningGenerating Questions and Multiple-Choice Answers using Semantic Analysis of Texts
We present a novel approach to automated question generation that improves upon prior work both from a technology perspective and from an assessment perspective. Our system is aimed at engaging language learners by gener…
coreference-resolutionCoreference ResolutionMultiple-choiceQuestion Generation+3