paper-with-me

Papers Answer Selection

“Answer Selection” 태그가 달린 논문 210편 · 필터 해제

Tracing Audio Grounding and Answer Selection in Audio LLMs

2026-09-04 · Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung arxiv

Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is …

Answer Selection

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

2026-08-26 · Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She 외 arxiv

Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules…

Answer Selection

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

2026-08-21 · Haorui Xu, Yuzhou Zhu, Liyuan Gao arxiv

Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confid…

Mathematical ReasoningAnswer Selection

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

2026-08-13 · Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe arxiv

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfide…

Question AnsweringAnswer Selection

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

2026-08-12 · Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang 외 arxiv

Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the …

Reinforcement LearningQuestion AnsweringAnswer Selection

Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering

2026-08-04 · Khaled Ziani arxiv

Large language models can generate fluent responses to Islamic questions while introducing factual errors that are difficult to identify. This paper presents our system for \textsc{HalluScoring 2026} Task 2.1, \textit{Is…

Question AnsweringAnswer Selection

STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

2026-07-12 · Xinkang Li, Rong Jiang, Xin Song, Ye Wang 외 arxiv

In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach to knowledge-intensive QA by combining retrieval with reasoning. Existing methods mainly improve open-domain multi-hop …

Multi-hop Question AnsweringAnswer Selection

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

2026-07-10 · Nirjhar Das, Md. Al-Mamun Provath arxiv

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions…

Relational ReasoningQuestion AnsweringAnswer Selection

Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

2026-06-27 · Shahnewaz Karim Sakib, Anindya Bijoy Das arxiv

Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial predictions, and support a final decision. However, such communication may also intro…

General KnowledgeAnswer Selection

Formalize Once, Edit the Rest: Efficient Lean-Based Answer Selection for Math Reasoning

2026-06-14 · Ji Feng, Zhouxing Shi arxiv

With large language models (LLMs) increasingly applied to mathematical reasoning, formal proof assistants such as Lean can be leveraged to verify reasoning outputs with machine-checkable rigor, enabling use cases such as…

Mathematical ReasoningAnswer Selection

Boosting Self-Consistency with Ranking

2026-06-03 · Maria Marina, Daniil Moskovskiy, Sergey Pletenev, Mikhail Salnikov 외 arxiv

Self-consistency improves large language models by sampling multiple reasoning paths and selecting the most frequent answer, but majority voting often fails to recover correct answers that are already present among the s…

Question AnsweringAnswer Selection

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

2026-05-29 · Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng 외 arxiv

Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe what it sees and name the underlying pattern, yet still fail to choos…

Answer SelectionVisual Reasoning

Can LLM Teams Play What? Where? When?

2026-05-28 · Anastasia Kotelnikova, Viktor Byzov, Maria Dolzhenkova, Evgeny Kotelnikov arxiv

Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigate whether team-based interaction improves LLM performance in What? W…

Answer Selection

The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection

2026-05-26 · Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang 외 arxiv

LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate stude…

Answer Selection

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

2026-05-26 · Hadi Bayrami Asl Tekanlou, Mahdi Bakhtiyarzadeh, Jafar Razmara arxiv

Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, they may face challenges with culturally grounded knowledge within la…

Semantic SimilarityQuestion AnsweringAnswer Selection

Inference Time Optimization with Confidence Dynamics

2026-05-24 · Yu Wang, Minghao Liu, Jiayun Wang, Jinrui Huang 외 arxiv

Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, the critical role of model uncertainty remains largely u…

Answer Selection

SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning

2026-05-16 · Yongfeng Huang, Ruiying Chen, James Cheng arxiv

Retrieval-Augmented Generation (RAG) is widely employed to mitigate risks such as hallucinations and knowledge obsolescence in medical question answering, yet its predominantly single-round, static retrieval paradigm mis…

Question AnsweringAnswer Selection

Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding

2026-05-11 · Anton Bazdyrev, Ivan Bashtovyi, Ivan Havlytskyi, Oleksandr Kharytonov 외 arxiv

We participated in the Fifth UNLP shared task on multi-domain document understanding, where systems must answer Ukrainian multiple-choice questions from PDF collections and localize the supporting document and page. We p…

Answer GenerationAnswer SelectionPassage Ranking

VecCISC: Improving Confidence-Informed Self-Consistency with Reasoning Trace Clustering and Candidate Answer Selection

2026-05-08 · James Petullo, Sonny George, Dylan Cashman, Nianwen Xue arxiv

A standard technique for scaling inference-time reasoning is Self-Consistency, whereby multiple candidate answers are sampled from an LLM and the most common answer is selected. More recently, it has been shown that weig…

Semantic SimilarityAnswer Selection

A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering

2026-05-08 · Zhanliang Wang, Jiancong Xiao, Ruochen Jin, Shu Yang 외 arxiv

Calibration measures whether a model's predicted confidence aligns with its empirical accuracy, and is central to the reliable deployment of large language models (LLMs) in high-stakes domains such as medicine and law. W…

Question AnsweringAnswer Selection
1–20 / 210 다음 →