paper-with-me

Papers

SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning

2026-01-29 · Jinjun Peng, Magnus Saebo, Tianjun Zhong, Yi-Jie Cheng, Junfeng Yang, Baishakhi Ray, Simin Chen, Yangruibo Ding arxiv

The deployment of coding agents in privacy-sensitive and resource-constrained environments drives the demand for capable open-weight Small Language Models (SLMs). However, they suffer from a fundamental capability gap: unlike frontier large models, they lack the inference-time strong generalization to work with complicated, unfamiliar codebases. We identify that the prevailing Task-Centric Learning (TCL) paradigm, which scales exposure across disparate repositories, fails to address this limitation. In response, we propose Repository-Centric Learning (RCL), a paradigm shift that prioritizes vertical repository depth over horizontal task breadth, suggesting SLMs must internalize the "physics" of a target software environment through parametric knowledge acquisition, rather than attempting to recover it via costly inference-time search. Following this new paradigm, we design a four-unit Repository-Centric Experience, transforming static codebases into interactive learning signals, to train SWE-Spot-4B, a family of highly compact models built as repo-specialized experts that breaks established scaling trends, outperforming open-weight models up to larger (e.g., CWM by Meta, Qwen3-Coder-30B) and surpassing/matching efficiency-focused commercial models (e.g., GPT-4.1-mini, GPT-5-nano) across multiple SWE tasks. Further analysis reveals that RCL yields higher training sample efficiency and lower inference costs, emphasizing that for building efficient intelligence, repository mastery is a distinct and necessary dimension that complements general coding capability.

📄 PDF Abstract BibTeX arXiv:2601.21649

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

S-pot - a benchmark in spotting signs within continuous signing

2014-05-01 · LREC 2014 5 · Ville Viitaniemi, Tommi Jantunen, Leena Savolainen, Matti Karppa 외

In this paper we present S-pot, a benchmark setting for evaluating the performance of automatic spotting of signs in continuous sign language videos. The benchmark includes 5539 video files of Finnish Sign Language, grou…

A Repository for the Sustainable Management of Research Data

2012-05-01 · LREC 2012 5 · Emanuel Dima, Verena Henrich, Erhard Hinrichs, Marie Hinrichs 외

This paper presents the system architecture as well as the underlying workflow of the Extensible Repository System of Digital Objects (ERDO) which has been developed for the sustainable archiving of language resources wi…

Management

Repo2Vec: A Comprehensive Embedding Approach for Determining Repository Similarity

2021-07-11 · Md Omar Faruk Rokon, Pei Yan, Risul Islam, Michalis Faloutsos

How can we identify similar repositories and clusters among a large online archive, such as GitHub? Determiningrepository similarity is an essential building block in studying the dynamics and the evolution of such softw…

Building and benchmarking an Arabic Speech Commands dataset for small-footprint keyword spotting

2021-05-07 · Engineering Applications of Artificial Intelligence 2021 5 · Abdulkader Ghandoura, Farouk Hjabo, Oumayma Al Dakkak

The introduction of the Google Speech Commands dataset accelerated research and resulted in a variety of new deep learning approaches that address keyword spotting tasks. The main contribution of this work is the buildin…

BenchmarkingDeep LearningKeyword SpottingSmall-Footprint Keyword Spotting

Building a Public Domain Voice Database for Odia

2022-08-16 · WWW '22: Companion Proceedings of the Web Conference 2022 8 · Subhashish Panigrahi

Projects like Mozilla Common Voice were born to address the challenges of unavailability of voice data or the high cost of available data for use in speech technology such as Automatic Speech Recognition (ASR) research a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition