paper-with-me

MMLU

1개 벤치마크 · 논문 340편 · 이 태스크의 논문 보기 →

Benchmarks

MMLU-Pro

결과 2개

Most implemented

Qwen2 Technical Report

2024-07-15 · 구현 6개

Are We Done with MMLU?

2024-06-06 · 구현 3개

Papers

Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning

2025-07-16 · Prateek Chanda, Saral Sureka, Parth Pratim Chatterjee, KrishnaTeja Killamsetty 외

The performance of finetuned large language models (LLMs) hinges critically on the composition of the training mixture. However, selecting an optimal blend of task datasets remains a largely manual, heuristic driven proc…

DiversityMMLU

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs

2025-07-15 · Gabriel Bo, Koa Chang, Justin Gu

We present Step-wise Policy for Rare-tool Knowledge (SPaRK), a novel reinforcement learning framework that teaches large language models to explore diverse tool usage patterns beyond conventional high-temperature samplin…

DiversityMMLUOffline RLreinforcement-learning+1

Lizard: An Efficient Linearization Framework for Large Language Models

2025-07-11 · Chien Van Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur 외

We propose Lizard, a linearization framework that transforms pretrained Transformer-based Large Language Models (LLMs) into flexible, subquadratic architectures for infinite-context generation. Transformer-based LLMs fac…

Language ModelingLanguage ModellingMMLU

Integrating External Tools with Large Language Models to Improve Accuracy

2025-07-09 · Nripesh Niketan, Hadj Batatia

This paper deals with improving querying large language models (LLMs). It is well-known that without relevant contextual information, LLMs can provide poor quality responses or tend to hallucinate. Several initiatives ha…

Mathematical ReasoningMMLU

The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains

2025-07-08 · Scott Geng, Hamish Ivison, Chun-Liang Li, Maarten Sap 외

Improvements in language models are often driven by improving the quality of the data we train them on, which can be limiting when strong supervision is scarce. In this work, we show that paired preference data consistin…

MathMMLU

Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate

2025-07-08 · A. Bochkov

The prevailing paradigm for scaling large language models (LLMs) involves monolithic, end-to-end training, a resource-intensive process that lacks flexibility. This paper explores an alternative, constructive approach to…

Continual LearningMixture-of-ExpertsMMLU

전체 340편 보기 →