paper-with-me

Common Sense Reasoning 벤치마크

Common Sense Reasoning on CommonsenseQA

38개 결과 · ⬇ CSV · JSON

Accuracy

20.9 38.81 56.72 74.63 92.54 2018-11 2026-09 BERT-LARGE — 55.9 (2018-11-02) CAGE-reasoning — 64.7 (2019-06-06) RoBERTa-Large 355M — 72.1 (2019-07-26) BERT_CSlarge — 62.2 (2019-08-19) KagNet — 58.9 (2019-09-04) XLNet+GraphReason — 75.3 (2019-09-09) Albert Lan et al. (2020) (ensemble) — 76.5 (2019-09-26) RoBERTa+HyKAS Ma et al. (2019) — 73.2 (2019-10-30) UnifiedQA 11B (fine-tuned) — 79.1 (2020-05-02) T5-XXL 11B (fine-tuned) — 78.1 (2020-05-02) UnifiedQA 11B (zero-shot) — 76.2 (2020-05-02) UnifiedQA 440M (fine-tuned) — 64.0 (2020-05-02) BART-large 440M (fine-tuned) — 62.5 (2020-05-02) DEKCOR — 83.3 (2020-12-09) MUPPET Roberta Large — 79.2 (2021-01-26) Unicorn 11B (fine-tuned) — 79.3 (2021-03-24) QA-GNN — 76.1 (2021-04-13) DeBERTaV3-large+KEAR — 91.2 (2021-12-06) KEAR — 89.4 (2021-12-06) GPT-3 Direct Finetuned — 73.0 (2021-12-06) Chain of thought ASDiv — 28.6 (2022-01-28) STaR (on GPT-J) — 72.3 (2022-03-28) STaR without Rationalization (on GPT-J) — 68.8 (2022-03-28) GPT-J Direct Finetuned — 60.0 (2022-03-28) Few-shot CoT LaMDA 137B — 55.6 (2022-03-28) Few-shot CoT GPT-J — 36.6 (2022-03-28) Few-shot Direct GPT-J — 20.9 (2022-03-28) UL2 20B (chain-of-thought + self-consistency) — 55.7 (2022-05-10) UL2 20B (chain-of-thought) — 51.4 (2022-05-10) UL2 20B (zero-shot) — 34.2 (2022-05-10) DRAGON — 78.2 (2022-10-17) GrapeQA: PEGA — 73.5 (2023-03-22) OPT 66B (1-shot) — 66.4 (2023-03-30) Bloomberg GPT 50B (1-shot) — 65.5 (2023-03-30) BLOOM 176B (1-shot) — 64.2 (2023-03-30) GPT-NeoX 20B (1-shot) — 60.4 (2023-03-30) PaLM 2 (few‑shot, CoT, SC) — 90.4 (2023-05-17) GPT-4o (HPT) — 92.54 (2024-06-18) BERT-LARGE — 55.9 (2018-11-02) CAGE-reasoning — 64.7 (2019-06-06) RoBERTa-Large 355M — 72.1 (2019-07-26) XLNet+GraphReason — 75.3 (2019-09-09) Albert Lan et al. (2020) (ensemble) — 76.5 (2019-09-26) UnifiedQA 11B (fine-tuned) — 79.1 (2020-05-02) DEKCOR — 83.3 (2020-12-09) DeBERTaV3-large+KEAR — 91.2 (2021-12-06) GPT-4o (HPT) — 92.54 (2024-06-18)
RankModel Accuracy Extra Training Data PaperCodeYear
21 OPT 66B (1-shot) 66.4 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
22 Bloomberg GPT 50B (1-shot) 65.5 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
23 CAGE-reasoning 64.7 Explain Yourself! Leveraging Language Models for Commonsense Reasoning salesforce/cos-e 2019
24 BLOOM 176B (1-shot) 64.2 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
25 UnifiedQA 440M (fine-tuned) 64 UnifiedQA: Crossing Format Boundaries With a Single QA System allenai/unifiedqa · facebookresearch/metaicl 2020
26 BART-large 440M (fine-tuned) 62.5 UnifiedQA: Crossing Format Boundaries With a Single QA System allenai/unifiedqa · facebookresearch/metaicl 2020
27 BERT_CSlarge 62.2 Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models 2019
28 GPT-NeoX 20B (1-shot) 60.4 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
29 GPT-J Direct Finetuned 60.0 STaR: Bootstrapping Reasoning With Reasoning ezelikman/STaR 2022
30 KagNet 58.9 KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning INK-USC/KagNet · INK-USC/MHGRN 2019
31 BERT-LARGE 55.9 CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge jonathanherzig/commonsenseqa · xlang-ai/batch-prompting · hkunlp/batch-prompting · +1 2018
32 UL2 20B (chain-of-thought + self-consistency) 55.7 UL2: Unifying Language Learning Paradigms google-research/google-research · opennlg/openba-v2 2022
33 Few-shot CoT LaMDA 137B 55.6 STaR: Bootstrapping Reasoning With Reasoning ezelikman/STaR 2022
34 UL2 20B (chain-of-thought) 51.4 UL2: Unifying Language Learning Paradigms google-research/google-research · opennlg/openba-v2 2022
35 Few-shot CoT GPT-J 36.6 STaR: Bootstrapping Reasoning With Reasoning ezelikman/STaR 2022
36 UL2 20B (zero-shot) 34.2 UL2: Unifying Language Learning Paradigms google-research/google-research · opennlg/openba-v2 2022
37 Chain of thought ASDiv 28.6 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models microsoft/guidance · guidance-ai/guidance · thudm/chatglm2-6b · +16 2022
38 Few-shot Direct GPT-J 20.9 STaR: Bootstrapping Reasoning With Reasoning ezelikman/STaR 2022
← 이전 21–38 / 38 페이지당 10 20 50 100