paper-with-me

홈 › Papers

Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report

2025-07-09 · Li Du, Hanyu Zhao, Yiming Ju, Tengfei Pan arxiv

Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction of high-quality instruction datasets is crucial for enhancing model performance and generalizability. Although current instruction datasets have reached tens of millions of samples, models finetuned on them may still struggle with complex instruction following and tasks in rare domains. This is primarily due to limited expansion in both `coverage'' (coverage of task types and knowledge areas) and `depth'' (instruction complexity) of the instruction set. To address this issue, we propose a systematic instruction data construction framework, which integrates a hierarchical tagging system, an informative seed selection algorithm, an evolutionary data synthesis process, and a model deficiency diagnosis with targeted data generation. These components form an iterative closed-loop to continuously enhance the coverage and depth of instruction data. Based on this framework, we construct Infinity Instruct Subject, a high-quality dataset containing $\sim$1.5 million instructions. Experiments on multiple foundation models and benchmark tasks demonstrate its effectiveness in improving instruction-following capabilities. Further analyses suggest that Infinity Instruct Subject shows enlarged coverage and depth compared to comparable synthesized instruction datasets. Our work lays a theoretical and practical foundation for the efficient, continuous evolution of instruction datasets, moving from data quantity expansion to qualitative improvement.

📄 PDF Abstract BibTeX arXiv:2507.06968

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

2024-12-05 · CVPR 2025 1 · Jian Han, Jinlai Liu, Yi Jiang, Bin Yan 외

We present Infinity, a Bitwise Visual AutoRegressive Modeling capable of generating high-resolution, photorealistic images following language instruction. Infinity redefines visual autoregressive model under a bitwise to…

Image Generation

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

2024-10-24 · Shuhao Gu, Jialing Zhang, Siyuan Zhou, Kevin Yu 외

Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabilities. Despite the availability of several …

Image GenerationQuestion GenerationQuestion-GenerationVisual Question Answering (VQA)

The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

2026-05-28 · Zeli Su, Zhankai Xu, Tianlei Chen, Longfei Zheng 외 arxiv

Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specified tasks over externally provided reference text. In practice, such …

Reinforcement Learning

The Value of Insider Information for Super--Replication with Quadratic Transaction Costs

2019-10-22 · Yan Dolinsky, Jonathan Zouari

We study super--replication of European contingent claims in an illiquid market with insider information. Illiquidity is captured by quadratic transaction costs and insider information is modeled by an investor who can p…

On semiparametric estimation of the intercept of the sample selection model: a kernel approach

2023-02-10 · Zhewen Pan

This paper presents a new perspective on the identification at infinity for the intercept of the sample selection model as identification at the boundary via a transformation of the selection index. This perspective sugg…

regression