MMLU
1개 벤치마크 · 논문 340편 · 이 태스크의 논문 보기 →
Benchmarks
MMLU-Pro
Most implemented
Scaling Instruction-Finetuned Language Models
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Qwen2 Technical Report
tinyBenchmarks: evaluating LLMs with fewer examples
DataComp-LM: In search of the next generation of training sets for language models
Are We Done with MMLU?
Papers
Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning
The performance of finetuned large language models (LLMs) hinges critically on the composition of the training mixture. However, selecting an optimal blend of task datasets remains a largely manual, heuristic driven proc…
DiversityMMLUStep-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs
We present Step-wise Policy for Rare-tool Knowledge (SPaRK), a novel reinforcement learning framework that teaches large language models to explore diverse tool usage patterns beyond conventional high-temperature samplin…
DiversityMMLUOffline RLreinforcement-learning+1Lizard: An Efficient Linearization Framework for Large Language Models
We propose Lizard, a linearization framework that transforms pretrained Transformer-based Large Language Models (LLMs) into flexible, subquadratic architectures for infinite-context generation. Transformer-based LLMs fac…
Language ModelingLanguage ModellingMMLUIntegrating External Tools with Large Language Models to Improve Accuracy
This paper deals with improving querying large language models (LLMs). It is well-known that without relevant contextual information, LLMs can provide poor quality responses or tend to hallucinate. Several initiatives ha…
Mathematical ReasoningMMLUThe Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
Improvements in language models are often driven by improving the quality of the data we train them on, which can be limiting when strong supervision is scarce. In this work, we show that paired preference data consistin…
MathMMLUGrowing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
The prevailing paradigm for scaling large language models (LLMs) involves monolithic, end-to-end training, a resource-intensive process that lacks flexibility. This paper explores an alternative, constructive approach to…
Continual LearningMixture-of-ExpertsMMLU