paper-with-me

홈 › Papers

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

2026-07-28 · Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao, Haocheng Lu, Mengkang Hu, Michael Qizhe Shieh arxiv

Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and learning from checkpoint feedback. Can LLM agents automate this loop? Existing benchmarks entangle research decisions with optimization, serving, evaluation, and systems implementation, obscuring agents' research capability. We introduce RSIBench-Data, a controlled benchmark of LLM agents as data-centric researchers with a fixed post-training stack. Agents iteratively revise training-data strategies for a fixed target model; training and serving use Tinker-backed services, official evaluation runs through Harbor and E2B sandboxes, and budgets are fixed across agents. We evaluate four frontier agents on six benchmarks across software engineering, terminal use, scientific question answering, and mathematics. Agents demonstrate core data-centric research capabilities: in 58.33\% of settings, they improve upon the first valid attempt by refining strategies from feedback. However, improvement is inconsistent. Among searches continuing after the best observed score, 78.26\% end with a lower-scoring final attempt, while the rest only recover the same peak. A strong candidate may therefore appear early or midway through a run even as later revisions fail. Trajectory analysis identifies four patterns in stronger runs: accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preservation of strong checkpoints. These findings suggest that current agents can make useful data-centric discoveries but cannot yet translate feedback into consistent improvements. RSIBench-Data provides a measurable, auditable testbed for the research capabilities required for recursive self-improvement. We open-source our code at https://github.com/evolvent-ai/RSIBench-Data.

📄 PDF Abstract BibTeX arXiv:2607.25886

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos

2025-12-03 · Wenliang Guo, Yu Kong arxiv

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task…

HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

2022-03-03 · CVPR 2022 1 · Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu 외

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences c…

Action SegmentationBenchmarkingHuman-Object Interaction DetectionMotion Segmentation+5

Real Time Egocentric Object Segmentation: THU-READ Labeling and Benchmarking Results

2021-06-09 · E. Gonzalez-Sosa, G. Robledo, D. Gonzalez-Morin, P. Perez-Garcia 외

Egocentric segmentation has attracted recent interest in the computer vision community due to their potential in Mixed Reality (MR) applications. While most previous works have been focused on segmenting egocentric human…

BenchmarkingMixed RealityReal-Time Semantic SegmentationSegmentation+1

Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset

2025-06-04 · ZiRui Wang, Wenjing Bian, Xinghui Li, Yifu Tao 외

We introduce Oxford Day-and-Night, a large-scale, egocentric dataset for novel view synthesis (NVS) and visual relocalisation under challenging lighting conditions. Existing datasets often lack crucial combinations of fe…

3D geometryBenchmarkingNovel View Synthesis

HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding

2024-10-09 · Keliang Li, Zaifei Yang, Jiahe Zhao, Hongze Shen 외

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centr…

BenchmarkingInstruction Following