paper-with-me

Multi-task Language Understanding

6개 벤치마크 · 논문 57편 · 이 태스크의 논문 보기 →

Benchmarks

MMLU-Pro

결과 349개

MML

결과 88개

BBH-nlp

결과 30개

MGSM

결과 24개

BBH-alg

결과 14개

MMLU (5-Shot)

결과 2개

Most implemented

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

Papers

Measuring Hong Kong Massive Multi-Task Language Understanding

2025-05-04 · Chuxue Cao, Zhenghao Zhu, Junqi Zhu, Guoying Lu 외

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional …

MMLUMulti-task Language Understanding

Effectiveness of Zero-shot-CoT in Japanese Prompts

2025-03-09 · Shusuke Takayama, Ian Frank

We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think ste…

Abstract AlgebraCollege MathematicsMMLUMulti-task Language Understanding

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages

2025-02-16 · Jafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova 외

Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quali…

Machine TranslationMMLUMulti-task Language Understanding

IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

2025-01-27 · Sankalp KJ, Ashutosh Kumar, Laxmaan Balaji, Nikunj Kotecha 외

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic…

BenchmarkingDiversityMMLUMulti-task Language Understanding

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

2025-01-22 · DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang 외

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary st…

Mathematical ReasoningMulti-task Language UnderstandingQuestion AnsweringReinforcement Learning (RL)

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

2024-12-19 · QiHao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 외

Multiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). Howe…

MMLUMultiple-choiceMulti-task Language UnderstandingWorld Knowledge

전체 57편 보기 →