paper-with-me

Answer Selection

6개 벤치마크 · 논문 210편 · 이 태스크의 논문 보기 →

Benchmarks

ASNQ

결과 3개

CICERO

결과 2개

TrecQA

결과 1개

WikiQA

결과 1개

Most implemented

Papers

Tracing Audio Grounding and Answer Selection in Audio LLMs

2026-09-04 · Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung arxiv

Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is …

Answer Selection

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

2026-08-26 · Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She 외 arxiv

Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules…

Answer Selection

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

2026-08-21 · Haorui Xu, Yuzhou Zhu, Liyuan Gao arxiv

Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confid…

Mathematical ReasoningAnswer Selection

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

2026-08-13 · Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe arxiv

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfide…

Question AnsweringAnswer Selection

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

2026-08-12 · Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang 외 arxiv

Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the …

Reinforcement LearningQuestion AnsweringAnswer Selection

Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering

2026-08-04 · Khaled Ziani arxiv

Large language models can generate fluent responses to Islamic questions while introducing factual errors that are difficult to identify. This paper presents our system for \textsc{HalluScoring 2026} Task 2.1, \textit{Is…

Question AnsweringAnswer Selection

전체 210편 보기 →