paper-with-me

Papers

EvolveCaptions: Empowering DHH Users Through Real-Time Collaborative Captioning

2025-10-02 · Liang-Yuan Wu, Dhruv Jain arxiv

Automatic Speech Recognition (ASR) systems often fail to accurately transcribe speech from Deaf and Hard of Hearing (DHH) individuals, especially during real-time conversations. Existing personalization approaches typically require extensive pre-recorded data and place the burden of adaptation on the DHH speaker. We present EvolveCaptions, a real-time, collaborative ASR adaptation system that supports in-situ personalization with minimal effort. Hearing participants correct ASR errors during live conversations. Based on these corrections, the system generates short, phonetically targeted prompts for the DHH speaker to record, which are then used to fine-tune the ASR model. In a study with 12 DHH and six hearing participants, EvolveCaptions reduced Word Error Rate (WER) across all DHH users within one hour of use, using only five minutes of recording time on average. Participants described the system as intuitive, low-effort, and well-integrated into communication. These findings demonstrate the promise of collaborative, real-time ASR adaptation for more equitable communication.

📄 PDF Abstract BibTeX arXiv:2510.02181

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents

2024-10-15 · Bolun Sun, Yifan Zhou, Haiyun Jiang

This paper presents a novel application of large language models (LLMs) to enhance user comprehension of privacy policies through an interactive dialogue agent. We demonstrate that LLMs significantly outperform tradition…

ManagementQuestion Answering

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

2024-12-30 · Yifei HUANG, Jilan Xu, Baoqi Pei, Yuping He 외

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always …

Language ModelingLanguage ModellingTask PlanningVideo Generation

Agile Modeling: From Concept to Classifier in Minutes

2023-02-25 · ICCV 2023 1 · Otilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan 외

The application of computer vision to nuanced subjective use cases is growing. While crowdsourcing has served the vision community well for most objective tasks (such as labeling a "zebra"), it now falters on tasks where…

image-classificationImage Classification

Transformer Explainer: Interactive Learning of Text-Generative Models

2024-08-08 · Aeree Cho, Grace C. Kim, Alexander Karpekov, Alec Helbling 외

Transformers have revolutionized machine learning, yet their inner workings remain opaque to many. We present Transformer Explainer, an interactive visualization tool designed for non-experts to learn about Transformers …

Gaussian Material Synthesis

2018-04-23 · Károly Zsolnai-Fehér, Peter Wonka, Michael Wimmer

We present a learning-based system for rapid mass-scale material synthesis that is useful for novice and expert users alike. The user preferences are learned via Gaussian Process Regression and can be easily sampled for …