paper-with-me

Generative Question Answering

2개 벤치마크 · 논문 48편 · 이 태스크의 논문 보기 →

Benchmarks

CICERO

결과 4개

CoQA

결과 3개

Most implemented

Papers

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

2026-05-29 · Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar 외 arxiv

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucina…

Generative Question Answering

Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation

2026-02-26 · Aashish Anantha Ramakrishnan, Ardavan Saeedi, Hamid Reza Hassanzadeh, Fazlolah Mohaghegh 외 arxiv

With the widespread adoption of pre-trained Large Language Models (LLM), there exists a high demand for task-specific test sets to benchmark their performance in domains such as healthcare and biomedicine. However, the c…

Generative Question Answering

AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

2025-09-04 · Aisha Alansari, Hamzah Luqman arxiv

Recently, extensive research on the hallucination of the large language models (LLMs) has mainly focused on the English language. Despite the growing number of multilingual and Arabic-specific LLMs, evaluating LLMs' hall…

Generative Question Answering

Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs

2025-02-18 · Zixiao Wang, Duzhen Zhang, Ishita Agrawal, Shen Gao 외

Previous approaches to persona simulation large language models (LLMs) have typically relied on learning basic biographical information, or using limited role-play dialogue datasets to capture a character's responses. Ho…

Generative Question AnsweringMultiple-choiceQuestion AnsweringStyle Transfer

EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering

2025-01-22 · Chang Zong, Jian Wan, Siliang Tang, Lei Zhang

When addressing professional questions in the biomedical domain, humans typically acquire multiple pieces of information as evidence and engage in multifaceted evidence analysis to provide high-quality answers. Current L…

Answer GenerationGenerative Question AnsweringLanguage ModelingLanguage Modelling+2

Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering

2024-08-27 · Haowei Du, Huishuai Zhang, Dongyan Zhao

To address the hallucination in generative question answering (GQA) where the answer can not be derived from the document, we propose a novel evidence-enhanced triplet generation framework, EATQA, encouraging the model t…

Generative Question AnsweringHallucinationQuestion AnsweringTriplet

전체 48편 보기 →