paper-with-me

Question Answering

144개 벤치마크 · 논문 12,370편 · 이 태스크의 논문 보기 →

Benchmarks

GPQA Diamond

결과 617개

Humanity's Last Exam

결과 613개

SQuAD2.0

결과 286개

SQuAD1.1

결과 213개

HotpotQA

결과 73개

PIQA

결과 67개

BoolQ

결과 66개

COPA

결과 60개

TriviaQA

결과 56개

SQuAD1.1 dev

결과 55개

Natural Questions

결과 47개

OpenBookQA

결과 45개

WebQuestions

결과 37개

PubMedQA

결과 35개

TruthfulQA

결과 33개

MedQA

결과 32개

MultiRC

결과 30개

CronQuestions

결과 29개

WikiQA

결과 25개

SIQA

결과 24개

StoryCloze

결과 23개

DaNetQA

결과 22개

TimeQuestions

결과 21개

Quora Question Pairs

결과 19개

NewsQA

결과 18개

CNN / Daily Mail

결과 16개

DROP Test

결과 16개

bAbi

결과 14개

SQuAD2.0 dev

결과 13개

StrategyQA

결과 13개

TrecQA

결과 13개

MultiTQ

결과 11개

Bamboogle

결과 10개

NarrativeQA

결과 10개

CoQA

결과 9개

OBQA

결과 9개

TIQ

결과 9개

WikiHop

결과 9개

BioASQ

결과 8개

Children's Book Test

결과 8개

FEVER

결과 8개

SQA3D

결과 8개

TempQuestions

결과 8개

FQuAD

결과 7개

FinQA

결과 7개

KILT: ELI5

결과 7개

QASent

결과 7개

Quasart-T

결과 7개

RACE

결과 7개

Story Cloze

결과 7개

YahooCQA

결과 7개

DROP

결과 6개

FriendsQA

결과 6개

NQ (BEIR)

결과 6개

PeerQA

결과 6개

SemEvalCQA

결과 5개

AI2 Kaggle Dataset

결과 4개

BLURB

결과 4개

CheGeKa

결과 4개

Complex-CronQuestions

결과 4개

EgoTaskQA

결과 4개

FairytaleQA

결과 4개

FiQA-2018 (BEIR)

결과 4개

HotpotQA (BEIR)

결과 4개

HybridQA

결과 4개

MS MARCO

결과 4개

Molweni

결과 4개

MultiQ

결과 4개

NaturalQA

결과 4개

QuALITY

결과 4개

RuOpenBookQA

결과 4개

catbAbI LM-mode

결과 4개

catbAbI QA-mode

결과 4개

CaseHOLD

결과 3개

ConditionalQA

결과 3개

ConvFinQA

결과 3개

DuoRC

결과 3개

Mathematics Dataset

결과 3개

OTT-QA

결과 3개

ReClor

결과 3개

SCDE

결과 3개

SberQuAD

결과 3개

TweetQA

결과 3개

VNHSGE-English

결과 3개

WikiTableQuestions

결과 3개

AGI Eval

결과 2개

CODAH

결과 2개

COMPLEXQUESTIONS

결과 2개

CliCR

결과 2개

GeoQuestions1089

결과 2개

MCTest-500

결과 2개

MRQA

결과 2개

MapEval-API

결과 2개

MuLD (HotpotQA)

결과 2개

MuLD (NarrativeQA)

결과 2개

PopQA

결과 2개

PubChemQA

결과 2개

QuAC

결과 2개

Reverb

결과 2개

SQuAD

결과 2개

TempQA-WD

결과 2개

Torque

결과 2개

UniProtQA

결과 2개

VNHSGE Mathematics

결과 2개

VNHSGE-Biology

결과 2개

VNHSGE-Chemistry

결과 2개

VNHSGE-Civic

결과 2개

VNHSGE-Geography

결과 2개

VNHSGE-History

결과 2개

VNHSGE-Literature

결과 2개

VNHSGE-Physics

결과 2개

WikiSQL

결과 2개

AviationQA

결과 1개

BBH

결과 1개

ComplexWebQuestions

결과 1개

EfficientQA dev

결과 1개

EfficientQA test

결과 1개

GraphQuestions

결과 1개

HellaSwag

결과 1개

JaQuAD

결과 1개

KQA Pro

결과 1개

MCTest-160

결과 1개

MML

결과 1개

MRQA out-of-domain

결과 1개

MapEval-Textual

결과 1개

MedMCQA Dev

결과 1개

MetaQA

결과 1개

MultiSpanQA

결과 1개

QASPER

결과 1개

RecipeQA

결과 1개

SWAG

결과 1개

SchizzoSQUAD

결과 1개

SimpleQuestions

결과 1개

StepGame

결과 1개

TAT-QA

결과 1개

WebQuestionsSP

결과 1개

WebSRC

결과 1개

Most implemented

Attention Is All You Need

2017-06-12 · 구현 595개

Graph Attention Networks

2017-10-30 · 구현 93개

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

Papers

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

2026-09-14 · Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi 외 hf

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an age…

Question Answering

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

2026-09-09 · Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi arxiv

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their …

Question Answering

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

2026-09-09 · Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler 외 arxiv

We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA f…

Question Answering

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

2026-09-09 · Zizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li 외 arxiv

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from th…

Question AnsweringVideo Captioning

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

2026-09-09 · Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He 외 arxiv

The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…

Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion Answering

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

2026-09-09 · Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen 외 hf

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as…

Question Answering

전체 12,370편 보기 →