paper-with-me

홈 › Papers

When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following

2025-09-25 · Keno Harada, Yudai Yamazaki, Masachika Taniguchi, Edison Marrese-Taylor, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo arxiv

As large language models (LLMs) are increasingly applied to real-world scenarios, it becomes crucial to understand their ability to follow multiple instructions simultaneously. To systematically evaluate these capabilities, we introduce two specialized benchmarks for fundamental domains where multiple instructions following is important: Many Instruction-Following Eval (ManyIFEval) for text generation with up to ten instructions, and Style-aware Mostly Basic Programming Problems (StyleMBPP) for code generation with up to six instructions. Our experiments with the created benchmarks across ten LLMs reveal that performance consistently degrades as the number of instructions increases. Furthermore, given the fact that evaluating all the possible combinations of multiple instructions is computationally impractical in actual use cases, we developed three types of regression models that can estimate performance on both unseen instruction combinations and different numbers of instructions which are not used during training. We demonstrate that a logistic regression model using instruction count as an explanatory variable can predict performance of following multiple instructions with approximately 10% error, even for unseen instruction combinations. We show that relatively modest sample sizes (500 for ManyIFEval and 300 for StyleMBPP) are sufficient for performance estimation, enabling efficient evaluation of LLMs under various instruction combinations.

📄 PDF Abstract BibTeX arXiv:2509.21051

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationText Generation

Similar Papers 제목 키워드 기반

TernGEMM: GEneral Matrix Multiply Library with Ternary Weights for Fast DNN Inference

2021-11-13 · 2021 IEEE Workshop on Signal Processing Systems (SiPS) 2021 11 · Seokhyeon Choi, Kyuhong Shim, Jungwook Choi, Wonyong Sung 외

Efficient implementation of deep neural networks on CPU-based systems is very critical because applications proliferate to embedded and Internet of Things (IoT) systems. Many CPUs for personal computers and embedded syst…

CPU

MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark

2024-09-26 · Elliot L. Epstein, Kaisheng Yao, Jing Li, Xinyi Bai 외

Evaluating instruction following capabilities for multimodal, multi-turn dialogue is challenging. With potentially multiple instructions in the input model context, the task is time-consuming for human raters and we show…

Instruction Following

A matrix math facility for Power ISA(TM) processors

2021-04-07 · José E. Moreira, Kit Barton, Steven Battle, Peter Bergner 외

Power ISA(TM) Version 3.1 has introduced a new family of matrix math instructions, collectively known as the Matrix-Multiply Assist (MMA) facility. The instructions in this facility implement numerical linear algebra ope…

Math

Masked Matrix Multiplication for Emergent Sparsity

2024-02-21 · Brian Wheatman, Meghana Madhyastha, Randal Burns

Artificial intelligence workloads, especially transformer models, exhibit emergent sparsity in which computations perform selective sparse access to dense data. The workloads are inefficient on hardware designed for dens…

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

2026-02-23 · David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, Maksym Andriushchenko arxiv

LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized third-party code, knowledge, and instruc…