paper-with-me

Papers

Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents

2024-10-01 · Shiwei Wu, Chen Zhang, Yan Gao, Qimeng Wang, Tong Xu, Yao Hu, Enhong Chen

Instructional documents are rich sources of knowledge for completing various tasks, yet their unique challenges in conversational question answering (CQA) have not been thoroughly explored. Existing benchmarks have primarily focused on basic factual question-answering from single narrative documents, making them inadequate for assessing a model`s ability to comprehend complex real-world instructional documents and provide accurate step-by-step guidance in daily life. To bridge this gap, we present InsCoQA, a novel benchmark tailored for evaluating large language models (LLMs) in the context of CQA with instructional documents. Sourced from extensive, encyclopedia-style instructional content, InsCoQA assesses models on their ability to retrieve, interpret, and accurately summarize procedural guidance from multiple documents, reflecting the intricate and multi-faceted nature of real-world instructional tasks. Additionally, to comprehensively assess state-of-the-art LLMs on the InsCoQA benchmark, we propose InsEval, an LLM-assisted evaluator that measures the integrity and accuracy of generated responses and procedural instructions.

📄 PDF Abstract BibTeX arXiv:2410.00526

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingConversational Question AnsweringQuestion Answering

Similar Papers 제목 키워드 기반

MMCoQA: Conversational Question Answering over Text, Tables, and Images

2022-05-01 · ACL 2022 5 · Yongqi Li, Wenjie Li, Liqiang Nie

The rapid development of conversational assistants accelerates the study on conversational question answering (QA). However, the existing conversational QA systems usually answer users’ questions with a single knowledge …

BenchmarkingConversational Question AnsweringQuestion AnsweringRetrieval

Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models

2025-04-01 · Rafael Giebisch, Ken E. Friedl, Lev Sorokin, Andrea Stocco

In-car conversational systems bring the promise to improve the in-vehicle user experience. Modern conversational systems are based on Large Language Models (LLMs), which makes them prone to errors such as hallucinations,…

BenchmarkingConversational Question AnsweringQuestion Answering

Disambiguation in Conversational Question Answering in the Era of LLM: A Survey

2025-05-18 · Md Mehrab Tanjim, Yeonjun In, Xiang Chen, Victor S. Bursztyn 외

Ambiguity remains a fundamental challenge in Natural Language Processing (NLP) due to the inherent complexity and flexibility of human language. With the advent of Large Language Models (LLMs), addressing ambiguity has b…

BenchmarkingConversational Question AnsweringQuestion Answering

PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models

2024-04-08 · Pardis Moradbeiki, Nasser Ghadiri

Smart cities need the involvement of their residents to enhance quality of life. Conversational query-answering is an emerging approach for user engagement. There is an increasing demand of an advanced conversational que…

Conversational Question AnsweringKeyword ExtractionQuestion Answering

Speeding Up Question Answering Task of Language Models via Inverted Index

2022-10-24 · Xiang Ji, Yesim Sungu-Eryilmaz, Elaheh Momeni, Reza Rawassizadeh

Natural language processing applications, such as conversational agents and their question-answering capabilities, are widely used in the real world. Despite the wide popularity of large language models (LLMs), few real-…

Question Answering