paper-with-me

Papers

LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis

2026-03-06 · Tao Zhang, Rui Ma, Shuotao Xu, Yongqiang Xiong, Peng Cheng arxiv

GPU design space exploration (DSE) for modern AI workloads, such as Large-Language Model (LLM) inference, is challenging because of GPUs' vast, multi-modal design spaces, high simulation costs, and complex design optimization objectives (e.g. performance, power and area trade-offs). Existing automated DSE methods are often prohibitively expensive, either requiring an excessive number of exploration samples or depending on intricate, manually crafted analyses of interdependent critical paths guided by human heuristics. We present LUMINA, an LLM-driven GPU architecture exploration framework that leverage AI to enhance the DSE efficiency and efficacy for GPUs. LUMINA extracts architectural knowledge from simulator code and performs sensitivity studies to automatically compose DSE rules,which are auto-corrected during exploration. A core component of LUMINA is a DSE Benchmark that comprehensively evaluates and enhances LLMs' capabilities across three fundamental skills required for architecture optimization, which provides a principled and reproducible basis for model selection and ensuring consistent architectural reasoning. In the design space with 4.7 million possible samples, LUMINA identifies 6 designs of better performance and area than an A100 GPU efficiently, using only 20 steps via LLM-assisted bottleneck analysis. In comparison, LUMINA achieves 17.5x higher than design space exploration efficiency, and 32.9% better designs (i.e. Pareto Hypervolume) than Machine-Learning baselines, showcasing its ability to deliver high-quality design guidance with minimal search cost.

📄 PDF Abstract BibTeX arXiv:2603.05904

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration

2025-08-25 · Aditri Paul, Archan Paul arxiv

Autonomous planetary exploration demands real-time, high-fidelity environmental perception. Standard deep learning models require massive computational resources. Conversely, space-qualified onboard computers operate und…

LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts

2026-01-26 · Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour 외 arxiv

Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architect…

Adversarially Guided Actor-Critic

2021-02-08 · ICLR 2021 1 · Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux 외

Despite definite success in deep reinforcement learning problems, actor-critic algorithms are still confronted with sample inefficiency in complex environments, particularly in tasks where efficient exploration is a bott…

Deep Reinforcement LearningEfficient Exploration

Data-Local Autonomous LLM-Guided Neural Architecture Search for Multiclass Multimodal Time-Series Classification

2026-03-16 · Emil Hardarson, Luka Biedebach, Ómar Bessi Ómarsson, Teitur Hrólfsson 외 arxiv

Applying machine learning to sensitive time-series data is often bottlenecked by the iteration loop: Performance depends strongly on preprocessing and architecture, yet training often has to run on-premise under strict d…

Neural Architecture Search

Agent-Guided Gaze Estimation Network by Two-Eye Asymmetry Exploration

2024-10-30 · IEEE International Conference on Image Processing (ICIP) 2024 10 · Yichen Shi, Feifei Zhang, Wenming Yang, Guijin Wang 외

Gaze estimation is an important task in understanding human visual attention. Despite the performance gain brought by recent algorithm development, the task remains challenging due to two-eye appearance asymmetry resulti…

Gaze Estimationregression