Question Answering
144개 벤치마크 · 논문 12,370편 · 이 태스크의 논문 보기 →
Benchmarks
GPQA Diamond
Humanity's Last Exam
SQuAD2.0
SQuAD1.1
HotpotQA
PIQA
BoolQ
COPA
TriviaQA
SQuAD1.1 dev
Natural Questions
OpenBookQA
WebQuestions
PubMedQA
TruthfulQA
MedQA
MultiRC
CronQuestions
WikiQA
SIQA
StoryCloze
DaNetQA
TimeQuestions
Quora Question Pairs
NewsQA
CNN / Daily Mail
DROP Test
bAbi
Natural Questions (long)
SQuAD2.0 dev
StrategyQA
TrecQA
MultiTQ
Bamboogle
NarrativeQA
CoQA
OBQA
TIQ
WikiHop
BioASQ
Children's Book Test
FEVER
SQA3D
TempQuestions
FQuAD
FinQA
KILT: ELI5
QASent
Quasart-T
RACE
Story Cloze
YahooCQA
DROP
FriendsQA
NQ (BEIR)
PeerQA
SemEvalCQA
AI2 Kaggle Dataset
BLURB
CheGeKa
Complex-CronQuestions
EgoTaskQA
FairytaleQA
FiQA-2018 (BEIR)
HotpotQA (BEIR)
HybridQA
MS MARCO
Molweni
MultiQ
NaturalQA
QuALITY
RuOpenBookQA
catbAbI LM-mode
catbAbI QA-mode
CaseHOLD
ConditionalQA
ConvFinQA
DuoRC
Mathematics Dataset
OTT-QA
ReClor
SCDE
SberQuAD
TweetQA
VNHSGE-English
WikiTableQuestions
AGI Eval
CODAH
COMPLEXQUESTIONS
CliCR
GeoQuestions1089
MCTest-500
MRQA
MapEval-API
MuLD (HotpotQA)
MuLD (NarrativeQA)
PopQA
PubChemQA
QuAC
Reverb
SQuAD
TempQA-WD
Torque
UniProtQA
VNHSGE Mathematics
VNHSGE-Biology
VNHSGE-Chemistry
VNHSGE-Civic
VNHSGE-Geography
VNHSGE-History
VNHSGE-Literature
VNHSGE-Physics
WikiSQL
AviationQA
BBH
ComplexWebQuestions
EfficientQA dev
EfficientQA test
GraphQuestions
HellaSwag
JaQuAD
KQA Pro
MCTest-160
MML
MRQA out-of-domain
MapEval-Textual
MedMCQA Dev
MetaQA
MultiSpanQA
QASPER
RecipeQA
SWAG
SchizzoSQUAD
SimpleQuestions
StepGame
TAT-QA
WebQuestionsSP
WebSRC
Most implemented
Attention Is All You Need
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Graph Attention Networks
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
LLaMA: Open and Efficient Foundation Language Models
Papers
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an age…
Question AnsweringWhy Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their …
Question AnsweringFrom Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora
We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA f…
Question AnsweringPutting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from th…
Question AnsweringVideo CaptioningSEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…
Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion AnsweringReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as…
Question Answering