paper-with-me

Papers

Streaming Intended Query Detection using E2E Modeling for Continued Conversation

2022-08-29 · Shuo-Yiin Chang, Guru Prakash, Zelin Wu, Qiao Liang, Tara N. Sainath, Bo Li, Adam Stambler, Shyam Upadhyay, Manaal Faruqui, Trevor Strohman

In voice-enabled applications, a predetermined hotword isusually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each timeintroduces a cognitive burden in continued conversations. Toavoid repeating a hotword, we propose a streaming end-to-end(E2E) intended query detector that identifies the utterancesdirected towards the device and filters out other utterancesnot directed towards device. The proposed approach incor-porates the intended query detector into the E2E model thatalready folds different components of the speech recognitionpipeline into one neural network.The E2E modeling onspeech decoding and intended query detection also allows us todeclare a quick intended query detection based on early partialrecognition result, which is important to decrease latencyand make the system responsive. We demonstrate that theproposed E2E approach yields a 22% relative improvement onequal error rate (EER) for the detection accuracy and 600 mslatency improvement compared with an independent intendedquery detector. In our experiment, the proposed model detectswhether the user is talking to the device with a 8.7% EERwithin 1.4 seconds of median latency after user starts speaking.

📄 PDF Abstract BibTeX arXiv:2208.13322

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Streaming Detection of Queried Event Start

2024-12-04 · Cristobal Eyzaguirre, Eric Tang, Shyamal Buch, Adrien Gaidon 외

Robotics, autonomous driving, augmented reality, and many embodied computer vision applications must quickly react to user-defined events unfolding in real time. We address this setting by proposing a novel task for mult…

Autonomous Drivingparameter-efficient fine-tuningTransfer LearningVideo Understanding

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

2026-05-08 · Ke Ma, Jiaqi Tang, Bin Guo, Xueting Han 외 arxiv

Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, query-agnostic modeling of visual evidence. …

Scene Graph Generation

Real-Time Energy Pricing in New Zealand: An Evolving Stream Analysis

2024-08-29 · Yibin Sun, Heitor Murilo Gomes, Bernhard Pfahringer, Albert Bifet

This paper introduces a group of novel datasets representing real-time time-series and streaming data of energy prices in New Zealand, sourced from the Electricity Market Information (EMI) website maintained by the New Z…

Anomaly DetectionDrift DetectionPrediction Intervalsregression+1

Stream Query Denoising for Vectorized HD Map Construction

2024-01-17 · Shuo Wang, Fan Jia, Yingfei Liu, Yucheng Zhao 외

To enhance perception performance in complex and extensive scenarios within the realm of autonomous driving, there has been a noteworthy focus on temporal modeling, with a particular emphasis on streaming methods. The pr…

Autonomous DrivingDenoising

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

2026-03-04 · Zehua Fan, Wenqi Lyu, Wenxuan Song, Linge Zhao 외 arxiv

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but also predictive modeling of environment dyn…

Vision-Language NavigationRepresentation Learning