paper-with-me

Question Answering 벤치마크

Question Answering on FEVER

8개 결과 · ⬇ CSV · JSON

EM

50 54.73 59.45 64.18 68.9 2019-02 2026-09 Zero-shot — 50.0 (2019-02-14) Self-Ask — 64.2 (2022-10-07) DSP — 62.2 (2023-10-05) CoA — 68.9 (2024-03-26) Self-Ask — 64.2 (2024-03-26) DSP — 62.2 (2024-03-26) CoA w/o actions — 54.2 (2024-03-26) Zero-shot — 50.0 (2024-03-26) Zero-shot — 50.0 (2019-02-14) Self-Ask — 64.2 (2022-10-07) CoA — 68.9 (2024-03-26)
RankModel EM PaperCodeYear
1 CoA 68.9 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models MAGICS-LAB/Chain-of-Actions 2024
2 Self-Ask 64.2 Measuring and Narrowing the Compositionality Gap in Language Models ofirpress/self-ask 2022
2 Self-Ask 64.2 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models MAGICS-LAB/Chain-of-Actions 2024
4 DSP 62.2 DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines stanfordnlp/dsp · stanfordnlp/dspy · codelion/optillm 2023
4 DSP 62.2 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models MAGICS-LAB/Chain-of-Actions 2024
6 CoA w/o actions 54.2 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models MAGICS-LAB/Chain-of-Actions 2024
7 Zero-shot 50 Language Models are Unsupervised Multitask Learners huggingface/transformers · openai/gpt-2 · PaddlePaddle/PaddleNLP · +18 2019
7 Zero-shot 50 Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models MAGICS-LAB/Chain-of-Actions 2024
1–8 / 8 페이지당 10 20 50 100