paper-with-me

Question Answering 벤치마크

Question Answering on MultiRC

30개 결과 · ⬇ CSV · JSON

F1

18.8 36.62 54.45 72.27 90.1 2018-10 2026-09 BERT-large(single model) — 70.0 (2018-10-11) T5-XXL 11B (fine-tuned) — 88.1 (2019-10-23) GPT-3 175B (Few-Shot) — 75.4 (2020-05-28) DeBERTa-1.5B — 88.2 (2020-06-05) FLAN 137B (prompt-tuned) — 83.4 (2021-09-03) FLAN 137B (zero-shot) — 77.5 (2021-09-03) FLAN 137B (1-shot) — 72.1 (2021-09-03) KELM (finetuning BERT-large based single model) — 70.8 (2021-09-09) ST-MoE-32B 269B (fine-tuned) — 89.6 (2022-02-17) ST-MoE-L 4.1B (fine-tuned) — 86.0 (2022-02-17) PaLM 540B (finetuned) — 90.1 (2022-04-05) N-Grammer 343M — 62.0 (2022-07-13) AlexaTM 20B — 59.6 (2022-08-02) Neo-6B (QA + WS) — 63.8 (2022-10-05) Neo-6B (few-shot) — 60.8 (2022-10-05) Neo-6B (QA) — 58.8 (2022-10-05) Turing NLR v5 XXL 5.4B (fine-tuned) — 88.4 (2022-12-04) Vega v2 6B (fine-tuned) — 88.2 (2022-12-04) Bloomberg GPT 50B (1-shot) — 62.3 (2023-03-30) BLOOM 176B (1-shot) — 26.7 (2023-03-30) GPT-NeoX 20B (1-shot) — 22.9 (2023-03-30) OPT 66B (1-shot) — 18.8 (2023-03-30) PaLM 2-L (one-shot) — 88.2 (2023-05-17) PaLM 2-M (one-shot) — 84.1 (2023-05-17) PaLM 2-S (one-shot) — 84.0 (2023-05-17) BERT-large(single model) — 70.0 (2018-10-11) T5-XXL 11B (fine-tuned) — 88.1 (2019-10-23) DeBERTa-1.5B — 88.2 (2020-06-05) ST-MoE-32B 269B (fine-tuned) — 89.6 (2022-02-17) PaLM 540B (finetuned) — 90.1 (2022-04-05)
RankModel F1EM PaperCodeYear
21 AlexaTM 20B 59.6 AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model amazon-science/alexa-teacher-models 2022
22 Neo-6B (QA) 58.8 Ask Me Anything: A simple strategy for prompting language models hazyresearch/ama_prompting · simran-arora/privacy_fm · simran-arora/focus 2022
23 BLOOM 176B (1-shot) 26.7 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
24 GPT-NeoX 20B (1-shot) 22.9 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
25 OPT 66B (1-shot) 18.8 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
26 T5-11B 63.3 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer huggingface/transformers · PaddlePaddle/PaddleNLP · google-research/text-to-text-transfer-transformer · +54 2019
27 Hybrid H3 355M (3-shot, logit scoring) 59.7 Hungry Hungry Hippos: Towards Language Modeling with State Space Models hazyresearch/safari · hazyresearch/h3 · lindermanlab/S5 2022
28 Hybrid H3 355M (0-shot, logit scoring) 59.5 Hungry Hungry Hippos: Towards Language Modeling with State Space Models hazyresearch/safari · hazyresearch/h3 · lindermanlab/S5 2022
29 Hybrid H3 125M (0-shot, logit scoring) 51.4 Hungry Hungry Hippos: Towards Language Modeling with State Space Models hazyresearch/safari · hazyresearch/h3 · lindermanlab/S5 2022
30 Hybrid H3 125M (3-shot, logit scoring) 48.9 Hungry Hungry Hippos: Towards Language Modeling with State Space Models hazyresearch/safari · hazyresearch/h3 · lindermanlab/S5 2022
← 이전 21–30 / 30 페이지당 10 20 50 100