Question Selection
1개 벤치마크 · 논문 40편 · 이 태스크의 논문 보기 →
Benchmarks
BIG-bench
Most implemented
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Training Compute-Optimal Large Language Models
BOBCAT: Bilevel Optimization-Based Computerized Adaptive Testing
Asking Clarifying Questions in Open-Domain Information-Seeking Conversations
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Papers
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire respon…
Explanation GenerationQuestion SelectionAnswer GenerationAsk4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA
Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstream interpretation. However, many medical VQA questions are generic, …
Visual Question AnsweringQuestion SelectionAnswer GenerationCA-BED: Conversation-Aware Bayesian Experimental Design
Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning. A key challenge lies in selecti…
Question SelectionOptimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake
Psychiatric intake is a sequential, high-stakes information-gathering process in which clinicians must decide what to ask, in what order, and how to interpret incomplete or ambiguous responses under limited time. Despite…
Question SelectionAMIGO: Agentic Multi-Image Grounding Oracle Benchmark
Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark…
Question SelectionSemantic Labeling for Third-Party Cybersecurity Risk Assessment: A Semi-Supervised Approach to Intent-Aware Question Retrieval
Third-Party Risk Assessment (TPRA) relies on large repositories of cybersecurity compliance questions used to assess external suppliers against standards such as ISO/IEC 27001 and NIST. In practice, not all questions are…
Semantic SimilarityQuestion Selection