paper-with-me

Question Selection

1개 벤치마크 · 논문 40편 · 이 태스크의 논문 보기 →

Benchmarks

BIG-bench

결과 2개

Most implemented

Papers

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

2026-08-04 · Salah Eddine Bekhouche, Abdessalam Bouchekif, Hichem Telli, Mohammed-En-Nadhir Zighem 외 arxiv

Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire respon…

Explanation GenerationQuestion SelectionAnswer Generation

Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA

2026-05-31 · Xiaorong Zhu, Qiang Li, Zibo Xu, Weijie Wang 외 arxiv

Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstream interpretation. However, many medical VQA questions are generic, …

Visual Question AnsweringQuestion SelectionAnswer Generation

CA-BED: Conversation-Aware Bayesian Experimental Design

2026-05-31 · Daniel Arnould, Rashad Aziz, Zixuan Kang, Tanav Changal 외 arxiv

Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning. A key challenge lies in selecti…

Question Selection

Optimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake

2026-04-23 · Guan Gui, Peter Zandi, Jacob Taylor, Ananya Joshi arxiv

Psychiatric intake is a sequential, high-stakes information-gathering process in which clinicians must decide what to ask, in what order, and how to interpret incomplete or ambiguous responses under limited time. Despite…

Question Selection

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

2026-03-30 · Min Wang, Ata Mahjoubfar arxiv

Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark…

Question Selection

Semantic Labeling for Third-Party Cybersecurity Risk Assessment: A Semi-Supervised Approach to Intent-Aware Question Retrieval

2026-02-09 · Ali Nour Eldin, Mohamed Sellami, Mehdi Acheli, Walid Gaaloul 외 arxiv

Third-Party Risk Assessment (TPRA) relies on large repositories of cybersecurity compliance questions used to assess external suppliers against standards such as ISO/IEC 27001 and NIST. In practice, not all questions are…

Semantic SimilarityQuestion Selection

전체 40편 보기 →