paper-with-me

Papers Question Selection

“Question Selection” 태그가 달린 논문 40편 · 필터 해제

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

2026-08-04 · Salah Eddine Bekhouche, Abdessalam Bouchekif, Hichem Telli, Mohammed-En-Nadhir Zighem 외 arxiv

Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire respon…

Explanation GenerationQuestion SelectionAnswer Generation

Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA

2026-05-31 · Xiaorong Zhu, Qiang Li, Zibo Xu, Weijie Wang 외 arxiv

Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstream interpretation. However, many medical VQA questions are generic, …

Visual Question AnsweringQuestion SelectionAnswer Generation

CA-BED: Conversation-Aware Bayesian Experimental Design

2026-05-31 · Daniel Arnould, Rashad Aziz, Zixuan Kang, Tanav Changal 외 arxiv

Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning. A key challenge lies in selecti…

Question Selection

Optimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake

2026-04-23 · Guan Gui, Peter Zandi, Jacob Taylor, Ananya Joshi arxiv

Psychiatric intake is a sequential, high-stakes information-gathering process in which clinicians must decide what to ask, in what order, and how to interpret incomplete or ambiguous responses under limited time. Despite…

Question Selection

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

2026-03-30 · Min Wang, Ata Mahjoubfar arxiv

Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark…

Question Selection

Semantic Labeling for Third-Party Cybersecurity Risk Assessment: A Semi-Supervised Approach to Intent-Aware Question Retrieval

2026-02-09 · Ali Nour Eldin, Mohamed Sellami, Mehdi Acheli, Walid Gaaloul 외 arxiv

Third-Party Risk Assessment (TPRA) relies on large repositories of cybersecurity compliance questions used to assess external suppliers against standards such as ISO/IEC 27001 and NIST. In practice, not all questions are…

Semantic SimilarityQuestion Selection

Efficient Evaluation of LLM Performance with Statistical Guarantees

2026-01-28 · Skyler Wu, Yash Nair, Emmanuel J. Candès arxiv

Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query budget, seek tight confidence intervals …

Question Selection

HybridQuestion: Human-AI Collaboration for Identifying High-Impact Research Questions

2025-12-18 · Keyu Zhao, Fengli Xu, Yong Li, Tie-Yan Liu arxiv

The "AI Scientist" paradigm is transforming scientific research by automating key stages of the research process, from idea generation to scholarly writing. This shift is expected to accelerate discovery and expand the s…

Question Selection

Structured Uncertainty guided Clarification for LLM Agents

2025-11-11 · Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt 외 arxiv

LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and task failures. Existing approaches operate in unstructured language spaces, ge…

Reinforcement LearningQuestion Selection

TestAgent: An Adaptive and Intelligent Expert for Human Assessment

2025-06-03 · Junhao Yu, Yan Zhuang, Yuxuan Sun, Weibo Gao 외

Accurately assessing internal human states is key to understanding preferences, offering personalized services, and identifying challenges in real-world applications. Originating from psychometrics, adaptive testing has …

Large Language ModelQuestion SelectionSociology

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

2025-05-27 · Xuanwen Ding, Chengjun Pan, Zejun Li, Jiwen Zhang 외

Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introd…

BenchmarkingQuestion Selection

QA-prompting: Improving Summarization with Large Language Models using Question-Answering

2025-05-20 · Neelabh Sinha

Language Models (LMs) have revolutionized natural language processing, enabling high-quality text generation through prompting and in-context learning. However, models often struggle with long-context summarization due t…

In-Context LearningQuestion AnsweringQuestion SelectionText Generation

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

2025-05-19 · Yanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang 외

The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard practices nowadays face fundamental trade-of…

BenchmarkingChatbotMMLUQuestion Selection

Adaptive political surveys and GPT-4: Tackling the cold start problem with simulated user interactions

2025-03-12 · Fynn Bachmann, Daan van der Weijden, Lucien Heitz, Cristina Sarasua 외

Adaptive questionnaires dynamically select the next question for a survey participant based on their previous answers. Due to digitalisation, they have become a viable alternative to traditional surveys in application ar…

Question Selection

PSCon: Product Search Through Conversations

2025-02-19 · Jie Zou, Mohammad Aliannejadi, Evangelos Kanoulas, Shuxi Han 외

Conversational Product Search ( CPS ) systems interact with users via natural language to offer personalized and context-aware product lists. However, most existing research on CPS is limited to simulated conversations, …

Intent DetectionKeyword ExtractionQuestion SelectionResponse Generation

Active Task Disambiguation with LLMs

2025-02-06 · Katarzyna Kobalczyk, Nicolas Astorga, Tennison Liu, Mihaela van der Schaar

Despite the impressive performance of large language models (LLMs) across various benchmarks, their ability to address ambiguously specified problems--frequent in real-world interactions--remains underexplored. To addres…

Experimental DesignQuestion Selection

Diffusion-Inspired Cold Start with Sufficient Prior in Computerized Adaptive Testing

2024-11-19 · Haiping Ma, Aoqing Xia, Changqian Wang, Hai Wang 외

Computerized Adaptive Testing (CAT) aims to select the most appropriate questions based on the examinee's ability and is widely used in online education. However, existing CAT systems often lack initial understanding of …

Question Selection

An incremental preference elicitation-based approach to learning potentially non-monotonic preferences in multi-criteria sorting

2024-09-04 · Zhuolin Li, Zhen Zhang, Witold Pedrycz

This paper introduces a novel incremental preference elicitation-based approach to learning potentially non-monotonic preferences in multi-criteria sorting (MCS) problems, enabling decision makers to progressively provid…

Active LearningQuestion Selection

Fast and Adaptive Questionnaires for Voting Advice Applications

2024-04-02 · Fynn Bachmann, Cristina Sarasua, Abraham Bernstein

The effectiveness of Voting Advice Applications (VAA) is often compromised by the length of their questionnaires. To address user fatigue and incomplete responses, some applications (such as the Swiss Smartvote) offer a …

DecoderMissing ValuesQuestion Selection

Survey of Computerized Adaptive Testing: A Machine Learning Perspective

2024-03-31 · Qi Liu, Yan Zhuang, Haoyang Bi, Zhenya Huang 외

Computerized Adaptive Testing (CAT) provides an efficient and tailored method for assessing the proficiency of examinees, by dynamically adjusting test questions based on their performance. Widely adopted across diverse …

cognitive diagnosisQuestion SelectionSociologySurvey
1–20 / 40 다음 →