Source-linked AI summary
Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities
Pablo Benalcazar, Maciej Kalka, Wilian Guamán, Jacek Kamiński
TL;DR
The paper addresses limited comparative evidence on financial outcomes from rule-based versus learning-based P2P electricity pricing. It compares these mechanisms across PV-only and selected PV-BES settings, finding that rule-based benchmarks lead in PV-only while SDR-shaped RL policies perform best within the RL family and improve with storage.
Problem
Comparative evidence on the financial outcomes of rule-based and learning-based P2P pricing mechanisms remains limited, while their relative advantage is unclear.
Method
The paper compares bill-sharing, mid-market rate, and supply-demand-ratio benchmarks with RL pricing modes across PV-only and selected PV-BES configurations using financial and operational indicators.
Results
Rule-based benchmarks outperform the best RL policy in PV-only; within RL, SDR-shaped policies outperform multiplier-based pricing, and RL-SDR-L savings rise from €734.23 to €978.52 with storage.
Takeaways & Limitations
The findings support adaptive SDR-based pricing for local electricity markets, particularly when PV communities include storage, while rule-based pricing remains highly competitive in direct comparisons.
Takeaways & Limitations
Rule-based benchmarks were evaluated only in PV-only, and the study does not directly compare them with RL policies under storage; future work should test seasonal, network-constrained, and inter-community settings.
Abstract
from arXiv · showhide
This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities. The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing. The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with a fixed-parameter SDR variant as a non-learning control. Performance is assessed through community savings together with complementary financial and operational indicators. In the base PV-only configuration, the rule-based benchmarks outperform the best RL policy. With battery energy storage, evaluated for the RL policies only, community savings under the best RL policy increase from EUR 734.23 to EUR 978.52. Across the learning-based modes and in both configurations, SDR-shaped pricing outperforms the multiplier-based parameterization considered. The results indicate that rule-based pricing remains highly competitive wherever the two families are compared directly, and that storage substantially improves the learning-based outcomes under this accounting, while the distribution of benefits remains heterogeneous across households.
Nomenclature
The nomenclature defines the paper’s energy, pricing, cost, trading, and reinforcement-learning variables for residential P2P electricity settlement.
- Energy quantities: Energy variables describe community PV generation, household load, local surplus, residual demand, traded energy, and grid imports or exports.The notation distinguishes interval quantities from annual totals and separates energy sold internally from energy purchased internally.
- Prices: Price variables distinguish grid import and export prices, benchmark internal prices, RL-quoted buy and sell prices, and realized trading prices.The notation also includes the realized internal trading price and its average.
- Costs and indicators: Economic variables represent reference and post-trading community costs, bill-sharing allocations, community benefit, household cost, prosumer revenue, and community savings.The community benefit is defined as the difference between reference cost and pooled cost.
- Reinforcement learning: The RL notation covers the state vector, discretized net balance, mean battery state of charge, actions, Q-networks, greedy policy, discounting, optimization, target synchronization, episodes, and exploration.The state includes hourly context and learning parameters such as the exploration rate and decay rate.
1. Introduction
The paper studies whether learning-based P2P pricing improves financial outcomes over simpler benchmarks in residential PV communities, while considering the role of storage.
- Introduction: P2P electricity trading lets prosumers and consumers exchange electricity locally rather than relying exclusively on the external grid.The setting belongs to local energy markets whose participation, governance, and settlement mechanisms continue to evolve.
- Introduction: Pricing matters because it shapes the value of internal transactions and how economic gains are shared among participants.Prior work identifies SDR and MMR as widely discussed mechanisms, with pricing effects depending on technologies, tariffs, and solar conditions.
- Introduction: RL has been applied to P2P price adjustment and trading coordination, but its advantage over simpler rule-based benchmarks remains unclear.This uncertainty motivates a direct comparison of the two pricing families.
- Introduction: The paper compares bill-sharing, MMR, SDR, and three RL operating modes using financial, trading, and self-sufficiency indicators in PV-only and PV-BES settings.Rule-based mechanisms are assessed only in the base PV-only framework, whereas RL policies are evaluated in both configurations.
2.1. P2P Market Framework
The market framework matches community PV surplus with residual internal demand, then applies different settlement conditions for rule-based and RL pricing mechanisms.
- 2.1. P2P Market Framework: The community consists of PV-equipped prosumers and consumers without on-site generation.Internal exchange is limited by aggregate local surplus and aggregate residual demand after self-consumption.
- 2.1. P2P Market Framework: The traded volume is the internally matchable quantity determined by aggregate surplus and residual internal demand.Both quantities are measured net of self-consumption.
- 2.1. P2P Market Framework: Rule-based mechanisms settle the matchable volume fully, with internal prices bounded by grid import and export prices.This price corridor ensures neither side is worse off than grid settlement when the export price is below the import price.
- 2.1. P2P Market Framework: RL mechanisms quote prices ex ante and execute exchange only when prices lie strictly within the grid corridor and the internal spread is non-negative.Otherwise, no internal exchange occurs in that interval.
2.2. Benchmark Pricing and Settlement Mechanisms
The benchmark mechanisms use distinct settlement designs: bill-sharing redistributes community benefit, MMR uses a grid-corridor midpoint, and SDR adapts pricing to local balance.
- 2.2.1. Bill-Sharing Mechanism: Bill-sharing allocates the community benefit from local trading across participants according to their internally matched energy shares.It is an ex post collective settlement rather than an hourly internal price mechanism.
- 2.2.2. Mid-Market Rate Pricing: The mid-market rate sets one internal price between the grid import and export prices.This benchmark therefore maintains a direct relation to the external grid price corridor.
- 2.2.3. Supply–Demand Ratio Pricing: SDR pricing makes the internal price depend on local market balance through the supply–demand ratio.Its price responds to scarcity and surplus conditions within the grid price bounds.
- 2.2.3. Supply–Demand Ratio Pricing: The SDR mechanism uses a single internal settlement price while buyers cover any residual shortage at the grid import price.The reported buy price is an average across internal and residual purchases, preserving budget balance and leaving no external spread.
2.3. RL-Based Pricing
The RL formulation uses a centralized DQN to select either multiplier-based or SDR-shaped pricing parameters from community-state information. Prices are posted ex ante, and settlement depends on remaining inside the grid-price corridor with a non-negative internal spread.
- RL pricing formulation: The RL formulation uses a centralized agent that observes community state sequentially and selects pricing parameters from candidate actions or a fixed control.The three modes comprise multiplier-based pricing, learnable SDR-shaped pricing, and fixed-parameter SDR-shaped pricing.
- RL-Multiplier: RL-M applies action-specific buy and sell multipliers directly to the prevailing grid prices, using eight paired actions.Because the multipliers scale different grid prices, the resulting internal spread can invert and prevent internal settlement.
- SDR-shaped pricing: SDR-shaped pricing computes P2P prices from the lagged supply–demand ratio and bounds them within the grid-price corridor.Prices are posted ex ante from the state at the start of the interval; the lagged ratio is initialized to one at each episode start.
- SDR-shaped pricing: RL-SDR-Fixed uses one sensitivity pair throughout the horizon, so its single-action control performs no learning and cannot adapt prices to state.The fixed mode isolates the contribution of adapting the SDR sensitivity parameter to operating conditions.
- State and reward: The state compactly represents discretized community balance, mean battery state of charge, lagged SDR, and cyclic hour information.For PV-only operation, SOC is set to zero; PV generation is represented through community balance and the lagged SDR rather than as a separate feature.
- State and reward: The reward is hourly cost saving relative to a no-P2P reference, so the DQN maximizes aggregate community savings while accounting for settlement-spread costs.In PV-BES operation, battery energy offered for internal exchange has no grid-export alternative and may be curtailed when internal settlement fails.
- Learning procedure: The DQN approximates the action-value function with a feedforward network, target network, and experience replay, while evaluation uses greedy policies on the annual training dataset.Results are based on a single fixed-seed training run per mode and configuration.
2.4. Evaluation Metrics
The evaluation treats community savings as the main economic outcome and uses financial and operational metrics to explain how pricing mechanisms produce it. The comparison defines distinct treatments for rule-based and RL prices, energy flows, and self-sufficiency.
- Evaluation framework: Community savings are the main economic outcome because the RL reward is defined from aggregate community savings.Average trading price, user cost, prosumer revenue, traded energy, self-sufficiency, and grid imports provide complementary interpretation.
- Financial metrics: The metrics distinguish internally bought and sold energy, community grid imports and exports, and the internal settlement price.For rule-based mechanisms, one price enters both user cost and prosumer revenue; RL modes use separate internal buy and sell prices.
- Financial metrics: For RL modes, the reported average trading price uses the midpoint of the internal buy and sell prices over intervals with non-zero internal exchange.Rule-based benchmarks use their single settlement price for both user cost and prosumer revenue.
- Operational metrics: The self-sufficiency index uses total household consumption, including the share met by on-site generation, rather than residual demand.Community savings coincide with the community benefit in the metric definitions.
- Evaluation procedure: The evaluation procedure samples days with hourly PV, load, and grid-price profiles, steps through 24-hour episodes, and trains with replayed transitions.The algorithm updates the target network periodically and evaluates the resulting greedy policy.
3. Case Study
The case study evaluates P2P pricing in a 20-household residential community using annual hourly settlement, Polish demand and PV data, and two technical configurations. The comparison combines financial and operational metrics, with PV-BES storage modeled using a fixed upstream dispatch.
- Community and data: The community contains 20 heterogeneous households, including eight prosumers with PV and 12 consumers without local generation.Demand profiles are synthetic and evaluated over an annual horizon with hourly settlement.
- Community and data: PV generation profiles are derived from Polish irradiance and ambient-temperature data using a capacity-based model with temperature derating.
- Technical configurations: The study evaluates PV-only and PV-BES configurations with the same community composition, settlement interval, and PV portfolio.The PV-BES case adds aggregate storage capacity of 68.15 kWh and aggregate charge/discharge power of 32.07 kW.
- Technical configurations: Battery sizing and hourly dispatch are fixed upstream, so the settlement mechanism and RL policy do not alter battery operation.The dispatch is obtained through a genetic algorithm and includes degradation-aware, conservative cycling.
- Evaluation metrics: Hourly household-level P2P participation under RL-SDR-L is illustrated for June 18 in both PV-only and PV-BES configurations.
4. Results and Discussion
Rule-based benchmarks achieve the strongest community savings in PV-only operation, while storage improves learning-based outcomes and SDR-shaped RL policies outperform multiplier-based pricing. These aggregate gains remain unevenly distributed across households.
- Community-level comparison: €829.98: Bill-sharing, MMR, and SDR each achieve the highest community savings in the PV-only configuration.RL-SDR-L reaches €734.23, followed by RL-SDR-F at €616.26 and RL-M at €418.49.
- Community-level comparison: SDR-shaped RL formulations outperform multiplier-based pricing but do not surpass rule-based benchmarks in the PV-only case.The difference is attributed to settlement design: rule-based mechanisms are budget-balanced, whereas RL modes retain an internal buy–sell spread.
- Community-level comparison: €418.49 to €606.94, €616.26 to €860.09, and €734.23 to €978.52: storage raises savings for RL-M, RL-SDR-F, and RL-SDR-L, respectively.Under RL-SDR-L, €105.76 comes from storage alone and the internal settlement contributes €872.76.
- Operational outcomes: PV-BES extends trading beyond the solar window through evening battery discharge exchanged locally.The operational pattern is consistent with higher traded energy and self-sufficiency for RL-SDR-L in the storage configuration.
- Community-level comparison: RL-SDR-F and RL-SDR-L achieve the same traded energy and self-sufficiency index, so RL-SDR-L’s advantage comes from more effective economic settlement.In PV-only operation, its annual internal spread is €92.49 versus €210.46 for RL-SDR-F, a reduction of €117.97 matching the savings difference.
- Household-level distribution: PV-only mechanisms deliver identical aggregate savings but distribute benefits differently, with SDR more even between consumer and prosumer groups than BS and MMR.SDR has consumer and prosumer medians of 19.3% and 28.6%, while BS and MMR have consumer medians near 10% and prosumer medians above 100%.
- Household-level distribution: Storage increases RL community savings but produces a more extended upper tail, indicating stronger gains for only part of the household population.RL-SDR-L remains the best-performing learning-based mode under battery storage.
5. Conclusions
Rule-based pricing outperformed the best RL policy in the PV-only comparison, while RL community savings increased with battery storage. Within the RL family, SDR-shaped pricing outperformed the multiplier-based parameterization, although network charges, levies, and taxes were excluded.
- Rule-based benchmarks outperformed the best RL policy in the base PV-only configuration.
- €978.52 in community savings was achieved by the best RL policy with battery storage, up from €734.23 without storage.These values were reported before accounting for energy used to charge the batteries.
- SDR-shaped RL policies outperformed the multiplier-based parameterization in both configurations.RL-SDR-L’s advantage reflected more effective economic settlement of the same internal exchange, rather than larger traded volume or higher self-sufficiency.
- Reported savings are upper bounds because the settlement excludes network charges, levies, and taxes.
- Seasonal conditions, rule-based benchmarks with storage, and network-constrained or inter-community trading remain directions for future work.