paper-with-me

홈 › Papers

Voice Memory for Agentic Speech Recognition

2026-07-29 · Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg arxiv

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

📄 PDF Abstract BibTeX arXiv:2607.26410

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

2026-01-14 · Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye 외 arxiv

We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintui…

Speech RecognitionQuestion Answering

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

2025-10-09 · Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni 외 arxiv

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities su…

Adversarial RobustnessQuestion AnsweringVoice Conversion

VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition

2020-09-09 · Quan Wang, Ignacio Lopez Moreno, Mert Saglam, Kevin Wilson 외

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a …

CPUspeech-recognitionSpeech Recognition

A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder

2016-10-13 · Hans Krupakar, Keerthika Rajvel, Bharathi B, Angel Deborah S 외

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice…

DecoderSentenceSpeech SynthesisSurvey+1

Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition

2022-04-08 · Shaojin Ding, Rajeev Rikhye, Qiao Liang, Yanzhang He 외

Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this …

Action DetectionActivity DetectionCPUspeech-recognition+1