paper-with-me

홈 › Papers

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

2026-06-03 · Zhihan Li, Hankun Wang, Yiwei Guo, Bohan Li, Xie Chen, Kai Yu arxiv

Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on internal confidence estimation or auxiliary language models. We propose READ (Reference-free Hypothesis Evaluation with Acoustic Discrepancy), a novel metric that evaluates ASR hypotheses directly from the speech signal. READ emphasizes the acoustic grounding of hypotheses. It uses a pretrained auto-regressive TTS model to compute the conditional likelihood of speech tokens given a text hypothesis, to measure fine-grained acoustic discrepancy between speech and text. Without additional training, READ can be applied for hypothesis refinement. Experiments show that READ correlates with specific recognition errors and improves ASR outputs, achieving up to 20\% relative error rate reduction, with particularly strong gains under noisy conditions.

📄 PDF Abstract BibTeX arXiv:2606.04680

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Do we read what we hear? Modeling orthographic influences on spoken word recognition

2021-04-01 · EACL 2021 2 · Nicole Macher, Badr M. Abdullah, Harm Brouwer, Dietrich Klakow

Theories and models of spoken word recognition aim to explain the process of accessing lexical knowledge given an acoustic realization of a word form. There is consensus that phonological and semantic information is cruc…

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations

2026-06-07 · Lei Lin, Ronghao Wang, Chunbao Zhou, Jue Wang 외 arxiv

Modern artificial intelligence excels at prediction but cannot explain. From large language models to AI-for-science systems, today's machines answer what by recombining patterns already present in the human literature, …

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

2026-07-22 · Siqian Tong, Xuan Li, Chaozhuo Li, Baolong Bi 외 arxiv

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-t…

HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

2026-09-09 · Jianzhi Shen, Keyu Mao, Minghao Shao, Chuanyang Jin 외 arxiv

Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved…

Whole Heart Mesh Generation For Image-Based Computational Simulations By Learning Free-From Deformations

2021-07-22 · Fanwei Kong, Shawn C. Shadden

Image-based computer simulation of cardiac function can be used to probe the mechanisms of (patho)physiology, and guide diagnosis and personalized treatment of cardiac diseases. This paradigm requires constructing simula…