paper-with-me

Papers

VoiSeR: A New Benchmark for Voice-Based Search Refinement

2021-04-01 · EACL 2021 2 · Simone Filice, Giuseppe Castellucci, Marcus Collins, Eugene Agichtein, Oleg Rokhlenko

Voice assistants, e.g., Alexa or Google Assistant, have dramatically improved in recent years. Supporting voice-based search, exploration, and refinement are fundamental tasks for voice assistants, and remain an open challenge. For example, when using voice to search an online shopping site, a user often needs to refine their search by some aspect or facet. This common user intent is usually available through a {``}filter-by{''} interface on online shopping websites, but is challenging to support naturally via voice, as the intent of refinements must be interpreted in the context of the original search, the initial results, and the available product catalogue facets. To our knowledge, no benchmark dataset exists for training or validating such contextual search understanding models. To bridge this gap, we introduce the first large-scale dataset of voice-based search refinements, VoiSeR, consisting of about 10,000 search refinement utterances, collected using a novel crowdsourcing task. These utterances are intended to refine a previous search, with respect to a search facet or attribute (e.g., brand, color, review rating, etc.), and are manually annotated with the specific intent. This paper reports qualitative and empirical insights into the most common and challenging types of refinements that a voice-based conversational search system must support. As we show, VoiSeR can support research in conversational query understanding, contextual user intent prediction, and other conversational search topics to facilitate the development of conversational search systems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeConversational Search

Similar Papers 제목 키워드 기반

Lost in Transcription: How Speech-to-Text Errors Derail Code Understanding

2026-01-20 · Jayant Havare, Ashish Mittal, Srikanth Tamilselvam, Ganesh Ramakrishnan arxiv

Code understanding is a foundational capability in software engineering tools and developer workflows. However, most existing systems are designed for English-speaking users interacting via keyboards, which limits access…

Speech RecognitionQuestion Answering

Multi-interaction TTS toward professional recording reproduction

2025-07-01 · Hiroki Kanagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima arxiv

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has b…

Text-To-Speech Synthesis

Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching

2026-01-01 · Minhyeok Yun, Yong-Hoon Choi arxiv

Singing voice synthesis (SVS) aims to generate natural and expressive singing waveforms from symbolic musical scores. In cVAE-based SVS, however, a mismatch arises because the decoder is trained with latent representatio…

Voice Interaction With Conversational AI Could Facilitate Thoughtful Reflection and Substantive Revision in Writing

2025-04-11 · Jiho Kim, Philippe Laban, Xiang 'Anthony' Chen, Kenneth C. Arnold

Writing well requires not only expressing ideas but also refining them through revision, a process facilitated by reflection. Prior research suggests that feedback delivered through dialogues, such as those in writing ce…

Your Voice Cloning System is Secretly a Voice Anonymizer

2026-08-27 · Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu arxiv

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speec…

Voice Conversion