paper-with-me

Papers

DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science

2026-02-27 · Fan Shu, Yite Wang, Ruofan Wu, Boyi Liu, Zhewei Yao, Yuxiong He, Feng Yan arxiv

The fast-growing demands in using Large Language Models (LLMs) to tackle complex multi-step data science tasks create an emergent need for accurate benchmarking. There are two major gaps in existing benchmarks: (i) the lack of standardized, process-aware evaluation that captures instruction adherence and process fidelity, and (ii) the scarcity of accurately labeled training data. To bridge these gaps, we introduce DARE-bench, a benchmark designed for machine learning modeling and data science instruction following. Unlike many existing benchmarks that rely on human- or model-based judges, all tasks in DARE-bench have verifiable ground truth, ensuring objective and reproducible evaluation. To cover a broad range of tasks and support agentic tools, DARE-bench consists of 6,300 Kaggle-derived tasks and provides both large-scale training data and evaluation sets. Extensive evaluations show that even highly capable models such as gpt-o4-mini struggle to achieve good performance, especially in machine learning modeling tasks. Using DARE-bench training tasks for fine-tuning can substantially improve model performance. For example, supervised fine-tuning boosts Qwen3-32B's accuracy by 1.83x and reinforcement learning boosts Qwen3-4B's accuracy by more than 8x. These significant improvements verify the importance of DARE-bench both as an accurate evaluation benchmark and critical training data.

📄 PDF Abstract BibTeX arXiv:2602.24288

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningInstruction Following

Similar Papers 제목 키워드 기반

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

2026-02-09 · Yu Shang, Zhuohang Li, Yiding Ma, Weikang Su 외 arxiv

While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their evaluation remains fragmented. Current eval…

Video Generation

DARE: Diffusion Language Model Activation Reuse for Efficient Inference

2026-05-01 · Natalia Frumkin, Bokun Wang, Hung-Yueh Chiang, Chi-Chih Chang 외 arxiv

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to auto-regressive (AR) models, offering greater expressive capacity and potential for parallel generation and faster inference. However, op…

BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction

2025-10-18 · Tian Xia, Tianrun Gao, Wenhao Deng, Long Wei 외 arxiv

Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrated reasoning under strict physical constraints. While modern LLMs possess…

DaRePlane: Direction-aware Representations for Dynamic Scene Reconstruction

2024-10-18 · Ange Lou, Benjamin Planche, Zhongpai Gao, Yamin Li 외

Numerous recent approaches to modeling and re-rendering dynamic scenes leverage plane-based explicit representations, addressing slow training times associated with models like neural radiance fields (NeRF) and Gaussian …

NeRFNovel View Synthesis

DaReNeRF: Direction-aware Representation for Dynamic Scenes

2024-03-04 · CVPR 2024 1 · Ange Lou, Benjamin Planche, Zhongpai Gao, Yamin Li 외

Addressing the intricate challenge of modeling and re-rendering dynamic scenes, most recent approaches have sought to simplify these complexities using plane-based explicit representations, overcoming the slow training t…

NeRFNovel View Synthesis