paper-with-me

홈 › Papers

Meta-RAG on Large Codebases Using Code Summarization

2025-08-04 · Vali Tawosi, Salwa Alamir, Xiaomo Liu, Manuela Veloso arxiv

Large Language Model (LLM) systems have been at the forefront of applied Artificial Intelligence (AI) research in a multitude of domains. One such domain is software development, where researchers have pushed the automation of a number of code tasks through LLM agents. Software development is a complex ecosystem, that stretches far beyond code implementation and well into the realm of code maintenance. In this paper, we propose a multi-agent system to localize bugs in large pre-existing codebases using information retrieval and LLMs. Our system introduces a novel Retrieval Augmented Generation (RAG) approach, Meta-RAG, where we utilize summaries to condense codebases by an average of 79.8\%, into a compact, structured, natural language representation. We then use an LLM agent to determine which parts of the codebase are critical for bug resolution, i.e. bug localization. We demonstrate the usefulness of Meta-RAG through evaluation with the SWE-bench Lite dataset. Meta-RAG scores 84.67 % and 53.0 % for file-level and function-level correct localization rates, respectively, achieving state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2508.02611

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

2026-07-01 · Yongjian Tang, Ezgi Sarikayak, Doruk Tuncel, Jie M. Zhang 외 arxiv

Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Existing code summarization solutions often rely on a single language mod…

Code-Craft: Hierarchical Graph-Based Code Summarization for Enhanced Context Retrieval

2025-04-11 · David Sounthiraraj, Jared Hancock, Yassin Kortam, Ashok Javvaji 외

Understanding and navigating large-scale codebases remains a significant challenge in software engineering. Existing methods often treat code as flat text or focus primarily on local structural relationships, limiting th…

Code SummarizationInformation RetrievalRetrieval

XMainframe: A Large Language Model for Mainframe Modernization

2024-08-05 · Anh T. V. Dau, Hieu Trung Dao, Anh Tuan Nguyen, Hieu Trung Tran 외

Mainframe operating systems, despite their inception in the 1940s, continue to support critical sectors like finance and government. However, these systems are often viewed as outdated, requiring extensive maintenance an…

Code SummarizationLanguage ModelingLanguage ModellingLarge Language Model+3

Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs

2025-01-14 · Nilesh Dhulshette, Sapan Shah, Vinay Kulkarni

In large-scale software development, understanding the functionality and intent behind complex codebases is critical for effective development and maintenance. While code summarization has been widely studied, existing m…

Code Summarization

Assertion-Aware Test Code Summarization with Large Language Models

2025-11-09 · Anamul Haque Mollah, Ahmed Aljohani, Hyunsook Do arxiv

Unit tests often lack concise summaries that convey test intent, especially in auto-generated or poorly documented codebases. Large Language Models (LLMs) offer a promising solution, but their effectiveness depends heavi…

Semantic Similarity