paper-with-me

홈 › Papers

Meta-Black-Box-Optimization through Offline Q-function Learning

2025-05-04 · Zeyuan Ma, Zhiguang Cao, Zhou Jiang, Hongshu Guo, Yue-Jiao Gong

Recent progress in Meta-Black-Box-Optimization (MetaBBO) has demonstrated that using RL to learn a meta-level policy for dynamic algorithm configuration (DAC) over an optimization task distribution could significantly enhance the performance of the low-level BBO algorithm. However, the online learning paradigms in existing works makes the efficiency of MetaBBO problematic. To address this, we propose an offline learning-based MetaBBO framework in this paper, termed Q-Mamba, to attain both effectiveness and efficiency in MetaBBO. Specifically, we first transform DAC task into long-sequence decision process. This allows us further introduce an effective Q-function decomposition mechanism to reduce the learning difficulty within the intricate algorithm configuration space. Under this setting, we propose three novel designs to meta-learn DAC policy from offline data: we first propose a novel collection strategy for constructing offline DAC experiences dataset with balanced exploration and exploitation. We then establish a decomposition-based Q-loss that incorporates conservative Q-learning to promote stable offline learning from the offline dataset. To further improve the offline learning efficiency, we equip our work with a Mamba architecture which helps long-sequence learning effectiveness and efficiency by selective state model and hardware-aware parallel scan respectively. Through extensive benchmarking, we observe that Q-Mamba achieves competitive or even superior performance to prior online/offline baselines, while significantly improving the training efficiency of existing online baselines. We provide sourcecodes of Q-Mamba at https://github.com/MetaEvo/Q-Mamba.

📄 PDF Abstract BibTeX arXiv:2505.02010

Code (1)

metaevo/q-mamba 공식 구현 pytorch

Tasks

BenchmarkingMambaQ-Learning

Methods 이 논문이 사용한 방법론

DAC 설명 없음
Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…
Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Diffusion Models for Black-Box Optimization

2023-06-12 · Siddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya Grover

The goal of offline black-box optimization (BBO) is to optimize an expensive black-box function using a fixed dataset of function evaluations. Prior works consider forward approaches that learn surrogates to the black-bo…

Denoising

Black-Box Optimization From Small Offline Datasets via Meta Learning with Synthetic Tasks

2026-04-14 · Azza Fadhel, The Hung Tran, Trong Nghia Hoang, Jana Doppa arxiv

We consider the problem of offline black-box optimization, where the goal is to discover optimal designs (e.g., molecules or materials) from past experimental data. A key challenge in this setting is data scarcity: in ma…

Offline Stochastic Optimization of Black-Box Objective Functions

2024-12-03 · Juncheng Dong, Zihao Wu, Hamid Jafarkhani, Ali Pezeshki 외

Many challenges in science and engineering, such as drug discovery and communication network design, involve optimizing complex and expensive black-box functions across vast search spaces. Thus, it is essential to levera…

Drug DiscoveryStochastic Optimization

Generative Pretraining for Black-Box Optimization

2022-06-22 · Siddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya Grover

Many problems in science and engineering involve optimizing an expensive black-box function over a high-dimensional space. For such black-box optimization (BBO) problems, we typically assume a small budget for online fun…

An Offline Meta Black-box Optimization Framework for Adaptive Design of Urban Traffic Light Management Systems

2024-08-14 · Taeyoung Yun, Kanghoon Lee, Sujin Yun, Ilmyung Kim 외

Complex urban road networks with high vehicle occupancy frequently face severe traffic congestion. Designing an effective strategy for managing multiple traffic lights plays a crucial role in managing congestion. However…

Bayesian OptimizationManagement