Domain-adapted Learning and Imitation: DRL for Power Arbitrage
In this paper, we discuss the Dutch power market, which is comprised of a day-ahead market and an intraday balancing market that operates like an auction. Due to fluctuations in power supply and demand, there is often an imbalance that leads to different prices in the two markets, providing an opportunity for arbitrage. To address this issue, we restructure the problem and propose a collaborative dual-agent reinforcement learning approach for this bi-level simulation and optimization of European power arbitrage trading. We also introduce two new implementations designed to incorporate domain-specific knowledge by imitating the trading behaviours of power traders. By utilizing reward engineering to imitate domain expertise, we are able to reform the reward system for the RL agent, which improves convergence during training and enhances overall performance. Additionally, the tranching of orders increases bidding success rates and significantly boosts profit and loss (P&L). Our study demonstrates that by leveraging domain expertise in a general learning problem, the performance can be improved substantially, and the final integrated approach leads to a three-fold improvement in cumulative P&L compared to the original agent. Furthermore, our methodology outperforms the highest benchmark policy by around 50% while maintaining efficient computational performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Imitation Learningreinforcement-learningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Computational Arbitrage in AI Model Markets
Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a budget for a verifiable solution. An ar…
Quantitative Fundamental Theorem of Asset Pricing
In this paper we provide a quantitative analysis to the concept of arbitrage, that allows to deal with model uncertainty without imposing the no-arbitrage condition. In markets that admit ``small arbitrage", we can still…
An Improved Algorithm to Identify More Arbitrage Opportunities on Decentralized Exchanges
In decentralized exchanges (DEXs), the arbitrage paths exist abundantly in the form of both arbitrage loops (e.g. the arbitrage path starts from token A and back to token A again in the end, A, B,..., A) and non-loops (e…
On the quasi-sure superhedging duality with frictions
We prove the superhedging duality for a discrete-time financial market with proportional transaction costs under model uncertainty. Frictions are modeled through solvency cones as in the original model of [Kabanov, Y., H…
MathDeep Learning Statistical Arbitrage
Statistical arbitrage exploits temporal price differences between similar assets. We develop a unifying conceptual framework for statistical arbitrage and a novel data driven solution. First, we construct arbitrage portf…
Deep LearningTime SeriesTime Series Analysis