Papers Question Selection
“Question Selection” 태그가 달린 논문 40편 · 필터 해제
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire respon…
Explanation GenerationQuestion SelectionAnswer GenerationAsk4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA
Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstream interpretation. However, many medical VQA questions are generic, …
Visual Question AnsweringQuestion SelectionAnswer GenerationCA-BED: Conversation-Aware Bayesian Experimental Design
Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning. A key challenge lies in selecti…
Question SelectionOptimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake
Psychiatric intake is a sequential, high-stakes information-gathering process in which clinicians must decide what to ask, in what order, and how to interpret incomplete or ambiguous responses under limited time. Despite…
Question SelectionAMIGO: Agentic Multi-Image Grounding Oracle Benchmark
Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark…
Question SelectionSemantic Labeling for Third-Party Cybersecurity Risk Assessment: A Semi-Supervised Approach to Intent-Aware Question Retrieval
Third-Party Risk Assessment (TPRA) relies on large repositories of cybersecurity compliance questions used to assess external suppliers against standards such as ISO/IEC 27001 and NIST. In practice, not all questions are…
Semantic SimilarityQuestion SelectionEfficient Evaluation of LLM Performance with Statistical Guarantees
Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query budget, seek tight confidence intervals …
Question SelectionHybridQuestion: Human-AI Collaboration for Identifying High-Impact Research Questions
The "AI Scientist" paradigm is transforming scientific research by automating key stages of the research process, from idea generation to scholarly writing. This shift is expected to accelerate discovery and expand the s…
Question SelectionStructured Uncertainty guided Clarification for LLM Agents
LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and task failures. Existing approaches operate in unstructured language spaces, ge…
Reinforcement LearningQuestion SelectionTestAgent: An Adaptive and Intelligent Expert for Human Assessment
Accurately assessing internal human states is key to understanding preferences, offering personalized services, and identifying challenges in real-world applications. Originating from psychometrics, adaptive testing has …
Large Language ModelQuestion SelectionSociologyAutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introd…
BenchmarkingQuestion SelectionQA-prompting: Improving Summarization with Large Language Models using Question-Answering
Language Models (LMs) have revolutionized natural language processing, enabling high-quality text generation through prompting and in-context learning. However, models often struggle with long-context summarization due t…
In-Context LearningQuestion AnsweringQuestion SelectionText GenerationDecentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard practices nowadays face fundamental trade-of…
BenchmarkingChatbotMMLUQuestion SelectionAdaptive political surveys and GPT-4: Tackling the cold start problem with simulated user interactions
Adaptive questionnaires dynamically select the next question for a survey participant based on their previous answers. Due to digitalisation, they have become a viable alternative to traditional surveys in application ar…
Question SelectionPSCon: Product Search Through Conversations
Conversational Product Search ( CPS ) systems interact with users via natural language to offer personalized and context-aware product lists. However, most existing research on CPS is limited to simulated conversations, …
Intent DetectionKeyword ExtractionQuestion SelectionResponse GenerationActive Task Disambiguation with LLMs
Despite the impressive performance of large language models (LLMs) across various benchmarks, their ability to address ambiguously specified problems--frequent in real-world interactions--remains underexplored. To addres…
Experimental DesignQuestion SelectionDiffusion-Inspired Cold Start with Sufficient Prior in Computerized Adaptive Testing
Computerized Adaptive Testing (CAT) aims to select the most appropriate questions based on the examinee's ability and is widely used in online education. However, existing CAT systems often lack initial understanding of …
Question SelectionAn incremental preference elicitation-based approach to learning potentially non-monotonic preferences in multi-criteria sorting
This paper introduces a novel incremental preference elicitation-based approach to learning potentially non-monotonic preferences in multi-criteria sorting (MCS) problems, enabling decision makers to progressively provid…
Active LearningQuestion SelectionFast and Adaptive Questionnaires for Voting Advice Applications
The effectiveness of Voting Advice Applications (VAA) is often compromised by the length of their questionnaires. To address user fatigue and incomplete responses, some applications (such as the Swiss Smartvote) offer a …
DecoderMissing ValuesQuestion SelectionSurvey of Computerized Adaptive Testing: A Machine Learning Perspective
Computerized Adaptive Testing (CAT) provides an efficient and tailored method for assessing the proficiency of examinees, by dynamically adjusting test questions based on their performance. Widely adopted across diverse …
cognitive diagnosisQuestion SelectionSociologySurvey