Multi-task Language Understanding
6개 벤치마크 · 논문 57편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
LLaMA: Open and Efficient Foundation Language Models
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Language Models are Unsupervised Multitask Learners
Llama 2: Open Foundation and Fine-Tuned Chat Models
Papers
Measuring Hong Kong Massive Multi-Task Language Understanding
Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional …
MMLUMulti-task Language UnderstandingEffectiveness of Zero-shot-CoT in Japanese Prompts
We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think ste…
Abstract AlgebraCollege MathematicsMMLUMulti-task Language UnderstandingTUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quali…
Machine TranslationMMLUMulti-task Language UnderstandingIndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic…
BenchmarkingDiversityMMLUMulti-task Language UnderstandingDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary st…
Mathematical ReasoningMulti-task Language UnderstandingQuestion AnsweringReinforcement Learning (RL)MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Multiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). Howe…
MMLUMultiple-choiceMulti-task Language UnderstandingWorld Knowledge