paper-with-me

홈 › Papers

Backtracing: Retrieving the Cause of the Query

2024-03-06 · Rose E. Wang, Pawan Wirawarn, Omar Khattab, Noah Goodman, Dorottya Demszky

Many online content portals allow users to ask questions to supplement their understanding (e.g., of lectures). While information retrieval (IR) systems may provide answers for such user queries, they do not directly assist content creators -- such as lecturers who want to improve their content -- identify segments that _caused_ a user to ask those questions. We introduce the task of backtracing, in which systems retrieve the text segment that most likely caused a user query. We formalize three real-world domains for which backtracing is important in improving content delivery and communication: understanding the cause of (a) student confusion in the Lecture domain, (b) reader curiosity in the News Article domain, and (c) user emotion in the Conversation domain. We evaluate the zero-shot performance of popular information retrieval methods and language modeling methods, including bi-encoder, re-ranking and likelihood-based methods and ChatGPT. While traditional IR systems retrieve semantically relevant information (e.g., details on "projection matrices" for a query "does projecting multiple times still lead to the same point?"), they often miss the causally relevant context (e.g., the lecturer states "projecting twice gets me the same answer as one projection"). Our results show that there is room for improvement on backtracing and it requires new retrieval approaches. We hope our benchmark serves to improve future retrieval systems for backtracing, spawning systems that refine content generation and identify linguistic triggers influencing user queries. Our code and data are open-sourced: https://github.com/rosewang2008/backtracing.

📄 PDF Abstract BibTeX arXiv:2403.03956

Code (1)

rosewang2008/backtracing 공식 구현 pytorch

Tasks

Information RetrievalLanguage ModelingLanguage ModellingRe-RankingRetrieval

Similar Papers 제목 키워드 기반

Dynamically Retrieving Knowledge via Query Generation for Informative Dialogue Generation

2022-07-30 · Zhongtian Hu, Lifang Wang, Yangqi Chen, Yushuang Liu 외

Knowledge-driven dialog system has recently made remarkable breakthroughs. Compared with general dialog systems, superior knowledge-driven dialog systems can generate more informative and knowledgeable responses with pre…

Dialogue GenerationResponse Generation

Retrieving Time-Series Differences Using Natural Language Queries

2025-03-27 · Kota Dohi, Tomoya Nishida, Harsh Purohit, Takashi Endo 외

Effectively searching time-series data is essential for system analysis; however, traditional methods often require domain expertise to define search criteria. Recent advancements have enabled natural language-based sear…

Contrastive LearningNatural Language QueriesTime Series

Query-by-example Spoken Term Detection using Attention-based Multi-hop Networks

2017-09-01 · Chia-Wei Ao, Hung-Yi Lee

Retrieving spoken content with spoken queries, or query-by- example spoken term detection (STD), is attractive because it makes possible the matching of signals directly on the acoustic level without transcribing them in…

Distributed In-Context Learning under Non-IID Among Clients

2024-07-31 · Siqi Liang, Sumyeong Ahn, Jiayu Zhou

Advancements in large language models (LLMs) have shown their effectiveness in multiple complicated natural language reasoning tasks. A key challenge remains in adapting these models efficiently to new or unfamiliar task…

In-Context Learning

Query Decomposition for RAG: Balancing Exploration-Exploitation

2025-10-21 · Roxana Petcu, Kenton Murray, Daniel Khashabi, Evangelos Kanoulas 외 arxiv

Retrieval-augmented generation (RAG) systems address complex user requests by decomposing them into subqueries, retrieving potentially relevant documents for each, and then aggregating them to generate an answer. Efficie…