paper-with-me

홈 › Papers

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

2026-05-06 · Santosh Premi Adhikari, Radu Timofte, Dmitry Ignatov arxiv

Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations from scratch -- computationally expensive and yielding verbose code. We propose Delta-Code Generation, where fine-tuned LLMs generate compact unified diffs (deltas) to refine baseline architectures rather than synthesizing entire models. Our pipeline iteratively fine-tunes the LLM via LoRA on curated architectures from the LEMUR dataset, with MinHash-Jaccard novelty filtering for structural diversity. We evaluate three 7B-class LLMs -- DeepSeek-Coder-7B, Qwen2.5-Coder-7B, and Mistral-7B -- across six datasets (CIFAR-10, CIFAR-100, MNIST, SVHN, ImageNette, CelebA) using a 22-cycle protocol (1,100 candidates per LLM). All three substantially surpass the full-generation baseline (50.6% valid rate, 42.3% mean first-epoch accuracy): DeepSeek-Coder reaches 75.3% valid rate and 65.8% mean accuracy; Qwen2.5-Coder 72.1%/64.6%; Mistral 66.6%/66.1%. On CIFAR-10, best first-epoch accuracies reach 85.5% (Mistral), 85.2% (DeepSeek), 80.6% (Qwen) -- well above 63.98% full generation and 71.5% for the concurrent approach of Gu et al. Output lengths are 30-50 lines versus 200+ for full generation (75-85% reduction). A 50-epoch study confirms the 1-epoch proxy preserves rankings (Mistral: Spearman $ρ$ = 0.926). Delta-based generation is a token-efficient, multi-domain, LLM-agnostic alternative to full-model synthesis for LLM-driven NAS.

📄 PDF Abstract BibTeX arXiv:2605.04903

Code (0)

등록된 구현이 없습니다.

Tasks

Neural Architecture SearchCode Generation

Results from the Paper

RankTaskDatasetModelMetrics
#83 Neural Architecture Search CIFAR-10 Delta-Code Accuracy (% ): 66.1

Similar Papers 제목 키워드 기반

OpenDelta: A Plug-and-play Library for Parameter-efficient Adaptation of Pre-trained Models

2023-07-05 · Shengding Hu, Ning Ding, Weilin Zhao, Xingtai Lv 외

The scale of large pre-trained models (PTMs) poses significant challenges in adapting to downstream tasks due to the high optimization overhead and storage costs associated with full-parameter fine-tuning. To address thi…

Sparse Structure Search for Delta Tuning

2022-11-01 · NIPS 2022 11 · Shengding Hu, Zhen Zhang, Ning Ding, Yadao Wang 외

Adapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a smal…

Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models

2022-03-14 · Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei 외

Despite the success, the process of fine-tuning large-scale PLMs brings prohibitive adaptation costs. In fact, fine-tuning all the parameters of a colossal model and retaining separate instances for different tasks are p…

Text Classification

Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes

2026-02-16 · Aly Kassem, Thomas Jiralerspong, Negar Rostamzadeh, Golnoosh Farnadi arxiv

Model diffing methods aim to identify how fine-tuning changes a model's internal representations. Crosscoders approach this by learning shared dictionaries of interpretable latent directions between base and fine-tuned m…

Delta Activations: A Representation for Finetuned Large Language Models

2025-09-04 · Zhiqiu Xu, Amish Sethi, Mayur Naik, Ser-Nam Lim arxiv

The success of powerful open source Large Language Models (LLMs) has enabled the community to create a vast collection of post-trained models adapted to specific tasks and domains. However, navigating and understanding t…