paper-with-me

홈 › Papers

Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering

2025-08-17 · Eviatar Nachshoni, Arie Cattan, Shmuel Amar, Ori Shapira, Ido Dagan arxiv

Large Language Models (LLMs) have demonstrated strong performance in question answering (QA) tasks. However, Multi-Answer Question Answering (MAQA), where a question may have several valid answers, remains challenging. Traditional QA settings often assume consistency across evidences, but MAQA can involve conflicting answers. Constructing datasets that reflect such conflicts is costly and labor-intensive, while existing benchmarks often rely on synthetic data, restrict the task to yes/no questions, or apply unverified automated annotation. To advance research in this area, we extend the conflict-aware MAQA setting to require models not only to identify all valid answers, but also to detect specific conflicting answer pairs, if any. To support this task, we introduce a novel cost-effective methodology for leveraging fact-checking datasets to construct NATCONFQA, a new benchmark for realistic, conflict-aware MAQA, enriched with detailed conflict labels, for all answer pairs. We evaluate eight high-end LLMs on NATCONFQA, revealing their fragility in handling various types of conflicts and the flawed strategies they employ to resolve them.

📄 PDF Abstract BibTeX arXiv:2508.12355

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Conflict-Aware Federated Fine-Tuning of Large Language Models with Mixture-of-Experts

2026-06-14 · Yijun Lu, Zihan Fang, Pengpeng Qiao, Zheng Lin 외 arxiv

The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation. While federated learn…

Federated Learning

Trust-based Multiagent Consensus or Weightings Aggregation

2020-04-06 · Bruno Yun, Madalina Croitoru

We introduce a framework for reaching a consensus amongst several agents communicating via a trust network on conflicting information about their environment. We formalise our approach and provide an empirical and theore…

Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

2026-01-07 · Zheng Wu, Xingyu Lou, Xinbei Ma, Yansi Li 외 arxiv

Large Language Model (LLM)-based agents significantly extend the utility of LLMs by interacting with dynamic environments. However, enabling agents to continually learn new tasks without catastrophic forgetting remains a…

Continual Learning

Recognizing Conflict Opinions in Aspect-level Sentiment Classification with Dual Attention Networks

2019-11-01 · IJCNLP 2019 11 · Xingwei Tan, Yi Cai, Changxi Zhu

Aspect-level sentiment classification, which is a fine-grained sentiment analysis task, has received lots of attention these years. There is a phenomenon that people express both positive and negative sentiments towards …

ClassificationGeneral ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+2

A Rule-Based Relational XML Access Control Model in the Presence of Authorization Conflicts

2019-09-24 · Ali Alwehaibi, Mustafa Atay

There is considerable amount of sensitive XML data stored in relational databases. It is a challenge to enforce node level fine-grained authorization policies for XML data stored in relational databases which typically s…