Evaluating Appropriateness Of System Responses In A Spoken CALL Game
We describe an experiment carried out using a French version of CALL-SLT, a web-enabled CALL game in which students at each turn are prompted to give a semi-free spoken response which the system then either accepts or rejects. The central question we investigate is whether the response is appropriate; we do this by extracting pairs of utterances where both members of the pair are responses by the same student to the same prompt, and where one response is accepted and one rejected. When the two spoken responses are presented in random order, native speakers show a reasonable degree of agreement in judging that the accepted utterance is better than the rejected one. We discuss the significance of the results and also present a small study supporting the claim that native speakers are nearly always recognised by the system, while non-native speakers are rejected a significant proportion of the time.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSpeech RecognitionSimilar Papers 제목 키워드 기반
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
Short feedback responses, such as backchannels, play an important role in spoken dialogue. So far, most of the modeling of feedback responses has focused on their timing, often neglecting how their lexical and prosodic f…
Contrastive LearningTELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited…
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard acoustic cues during speech-to-text conver…
Dialogue GenerationResponse GenerationSpeech SynthesisAutomated Scoring of Clinical Expressive Language Evaluation Tasks
Many clinical assessment instruments used to diagnose language impairments in children include a task in which the subject must formulate a sentence to describe an image using a specific target word. Because producing se…
DiversityMachine TranslationSentenceTransfer Learning+2Chat or Learn: a Data-Driven Robust Question-Answering System
We present a voice-based conversational agent which combines the robustness of chatbots and the utility of question answering (QA) systems. Indeed, while data-driven chatbots are typically user-friendly but not goal-orie…
ArticlesChatbotcoreference-resolutionCoreference Resolution+2