paper-with-me

홈 › Papers

Talking to Your TV: Context-Aware Voice Search with Hierarchical Recurrent Neural Networks

2017-05-13 · Rao Jinfeng, Ture Ferhan, He Hua, Jojic Oliver, Lin Jimmy

We tackle the novel problem of navigational voice queries posed against an entertainment system, where viewers interact with a voice-enabled remote controller to specify the program to watch. This is a difficult problem for several reasons: such queries are short, even shorter than comparable voice queries in other domains, which offers fewer opportunities for deciphering user intent. Furthermore, ambiguity is exacerbated by underlying speech recognition errors. We address these challenges by integrating word- and character-level representations of the queries and by modeling voice search sessions to capture the contextual dependencies in query sequences. Both are accomplished with a probabilistic framework in which recurrent and feedforward neural network modules are organized in a hierarchical manner. From a raw dataset of 32M voice queries from 2.5M viewers on the Comcast Xfinity X1 entertainment system, we extracted data to train and test our models. We demonstrate the benefits of our hybrid representation and context-aware model, which significantly outperforms models without context as well as the current deployed product.

📄 PDF Abstract BibTeX arXiv:1705.04892

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication

2025-03-21 · Yiwen Xu, Monideep Chakraborti, Tianyi Zhang, Katelyn Eng 외

In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers

2026-01-29 · John Flynn, Wolfgang Paier, Dimitar Dinev, Sam Nhut Nguyen 외 arxiv

Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserv…

Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement

2023-12-15 · Xiaofeng Zhang, Zishan Xu, Hao Tang, Chaochen Gu 외

Low-light image enhancement is a crucial visual task, and many unsupervised methods tend to overlook the degradation of visible information in low-light scenes, which adversely affects the fusion of complementary informa…

Image EnhancementLow-Light Image Enhancement

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

2025-07-25 · Fang Kang, Yin Cao, Haoyu Chen arxiv

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more cha…

Talking Face GenerationFace Alignment

Generate Your Talking Avatar from Video Reference

2026-04-30 · Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li 외 arxiv

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient…

Reinforcement Learning