Papers Multi-task Language Understanding
“Multi-task Language Understanding” 태그가 달린 논문 57편 · 필터 해제
Measuring Hong Kong Massive Multi-Task Language Understanding
Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional …
MMLUMulti-task Language UnderstandingEffectiveness of Zero-shot-CoT in Japanese Prompts
We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think ste…
Abstract AlgebraCollege MathematicsMMLUMulti-task Language UnderstandingTUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quali…
Machine TranslationMMLUMulti-task Language UnderstandingIndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic…
BenchmarkingDiversityMMLUMulti-task Language UnderstandingDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary st…
Mathematical ReasoningMulti-task Language UnderstandingQuestion AnsweringReinforcement Learning (RL)MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Multiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). Howe…
MMLUMultiple-choiceMulti-task Language UnderstandingWorld KnowledgeLlama 3 Meets MoE: Efficient Upcycling
Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternative, increasing capacity without a propor…
Mixture-of-ExpertsMMLUMulti-task Language UnderstandingGPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data
Large language models require vast amounts of high-quality training data, but effective filtering of web-scale datasets remains a significant challenge. This paper demonstrates that GPT-4o is remarkably effective at iden…
Active LearningLanguage ModelingLanguage ModellingMulti-task Language Understanding+3Reasoning Beyond Bias: A Study on Counterfactual Prompting and Chain of Thought Reasoning
Language models are known to absorb biases from their training data, leading to predictions driven by statistical regularities rather than semantic relevance. We investigate the impact of these biases on answer choice pr…
counterfactualMMLUMulti-task Language UnderstandingThe Llama 3 Herd of Models
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, cod…
answerability predictionLanguage ModelingLanguage ModellingMulti-task Language Understanding+3Claude 3.5 Sonnet Model Card Addendum
This addendum to our Claude 3 Model Card describes Claude 3.5 Sonnet, a new model which outperforms our previous most capable model, Claude 3 Opus, while operating faster and at a lower cost. Claude 3.5 Sonnet offers i…
Code GenerationMMR totalmodelMulti-task Language Understanding+2Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
Assessing the effectiveness of large language models (LLMs) in performing different tasks is crucial for understanding their strengths and weaknesses. This paper presents Hierarchical Prompting Taxonomy (HPT), grounded o…
Arithmetic ReasoningCode GenerationCommon Sense ReasoningGSM8K+8Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
Ensembling multiple models has always been an effective approach to push the limits of existing performance and is widely used in classification tasks by simply averaging the classification probability vectors from multi…
Arithmetic ReasoningLanguage ModelingLanguage ModellingLarge Language Model+2MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering tasks with modified terms. We reasoned that …
Mathematical ReasoningMMLUMulti-task Language UnderstandingNatural Language Understanding+1MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning ac…
MMLUMulti-task Language UnderstandingBranch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
We investigate efficient methods for training Large Language Models (LLMs) to possess capabilities in multiple specialized domains, such as coding, math reasoning and world knowledge. Our method, named Branch-Train-MiX (…
Arithmetic ReasoningCode GenerationCommon Sense ReasoningMath+5The Claude 3 Model Family: Opus, Sonnet, Haiku
We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus, our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku, our fastest and least expe…
1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationCommon Sense Reasoning+8ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Ar…
ArabicMMLULanguage Model EvaluationLanguage ModelingLanguage Modelling+2Routoo: Learning to Route to Large Language Models Effectively
LLMs with superior response quality--particularly larger or closed-source models--often come with higher inference costs, making their deployment inefficient and costly. Meanwhile, developing foundational LLMs from scrat…
MMLUMulti-task Language UnderstandingMixtral of Experts
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts). For e…
Code GenerationCommon Sense ReasoningLanguage ModelingLanguage Modelling+4