paper-with-me

홈 › Papers

Optimizing Code Runtime Performance through Context-Aware Retrieval-Augmented Generation

2025-01-28 · Manish Acharya, Yifan Zhang, Kevin Leach, Yu Huang

Optimizing software performance through automated code refinement offers a promising avenue for enhancing execution speed and efficiency. Despite recent advancements in LLMs, a significant gap remains in their ability to perform in-depth program analysis. This study introduces AUTOPATCH, an in-context learning approach designed to bridge this gap by enabling LLMs to automatically generate optimized code. Inspired by how programmers learn and apply knowledge to optimize software, AUTOPATCH incorporates three key components: (1) an analogy-driven framework to align LLM optimization with human cognitive processes, (2) a unified approach that integrates historical code examples and CFG analysis for context-aware learning, and (3) an automated pipeline for generating optimized code through in-context prompting. Experimental results demonstrate that AUTOPATCH achieves a 7.3% improvement in execution efficiency over GPT-4o across common generated executable code, highlighting its potential to advance automated program runtime optimization.

📄 PDF Abstract BibTeX arXiv:2501.16692

Code (1)

manishacharya60/rag-optimization 공식 구현

Tasks

In-Context LearningRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Streaming-capable High-performance Architecture of Learned Image Compression Codecs

2022-08-02 · Fangzheng Lin, Heming Sun, Jiro Katto

Learned image compression allows achieving state-of-the-art accuracy and compression ratios, but their relatively slow runtime performance limits their usage. While previous attempts on optimizing learned image codecs fo…

CPUDecoderGPUImage Compression+1

AdaSpring: Context-adaptive and Runtime-evolutionary Deep Model Compression for Mobile Applications

2021-01-28 · Sicong Liu, Bin Guo, Ke Ma, Zhiwen Yu 외

There are many deep learning (e.g., DNN) powered mobile and wearable applications today continuously and unobtrusively sensing the ambient surroundings to enhance all aspects of human lives. To enable robust and private …

Model Compression

Transformer Multivariate Forecasting: Less is More?

2023-12-30 · Jingjing Xu, Caesar Wu, Yuan-Fang Li, Pascal Bouvry

In the domain of multivariate forecasting, transformer models stand out as powerful apparatus, displaying exceptional capabilities in handling messy datasets from real-world contexts. However, the inherent complexity of …

Temporal SequencesTime SeriesTime Series Forecasting

Semantic-Aware Scheduling for GPU Clusters with Large Language Models

2025-10-02 · Zerui Wang, Qinghao Hu, Ana Klimovic, Tianwei Zhang 외 arxiv

Deep learning (DL) schedulers are pivotal in optimizing resource allocation in GPU clusters, but operate with a critical limitation: they are largely blind to the semantic context of the jobs they manage. This forces the…

Reinforcement Learning based Interconnection Routing for Adaptive Traffic Optimization

2019-08-13 · Sheng-Chun Kao, Chao-Han Huck Yang, Pin-Yu Chen, Xiaoli Ma 외

Applying Machine Learning (ML) techniques to design and optimize computer architectures is a promising research direction. Optimizing the runtime performance of a Network-on-Chip (NoC) necessitates a continuous learning …

BIG-bench Machine Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)