Source-linked AI summary
Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach
Iias Faiud, Jonaid Shianifar, Michael Schukat, Karl Mason
TL;DR
Solar PV policy design must balance adoption gains and public expenditure amid uncertainty and heterogeneous decision-making. The study combines reinforcement learning with a stochastic agent-based model so a policymaker can learn annual incentives and explore adoption–cost trade-offs. Results show a consistent trade-off across algorithms, while dynamic RL policies cover a broader configuration range than static baselines.
Problem
PV policy design must balance adoption gains against public expenditure under uncertainty and heterogeneous decision-making.
Method
The study integrates reinforcement learning with a stochastic agent-based model, using a scalarised reward and continuous-control algorithms to learn annual policy incentives.
Results
A clear adoption–expenditure trade-off appears across PPO, SAC, and TD3, with stronger incentives producing higher adoption and substantially greater public cost.
Takeaways & Limitations
The framework enables dynamic policy adaptation and more flexible exploration of adoption–cost trade-offs than static policy configurations.
Takeaways & Limitations
The study covers one context over 2025–2040, uses a limited policy set, and simplifies behavioural, financial, and implementation constraints.
Abstract
from arXiv · showhide
Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent selects annual incentives, including capital grants, subsidised loan rates, and feed-in tariffs, over a 16-year horizon. Adoption--cost trade-offs are explored by varying policy preferences within a scalarised reward framework. Policies are learned using PPO, SAC, and TD3 and evaluated under stochastic simulation. The results show that this approach produces a clear trade-off structure: the highest-adoption policy (TD3, $w_{\text{cost}}=0.5$) achieves approximately 4,145 adopters at a cost of EUR 41.73 million, while the lowest-cost policy (PPO, $w_{\text{cost}}=2.0$) reduces expenditure to EUR 7.27 million with 2,682 adopters. The balanced policy (PPO, $w_{\text{cost}}=1.6$) achieves 3,495 adopters at a cost of EUR 22.47 million. Across algorithms, consistent trade-off patterns are observed, indicating robustness of the adoption--cost relationship. Compared with static baseline policies, the RL framework explores a broader range of policy configurations. These findings demonstrate the potential of RL as a flexible tool for adaptive policy design under uncertainty.
1 Introduction
PV policy design must account for heterogeneous, socially influenced adoption and changing conditions over time. The study addresses this by combining reinforcement learning with a stochastic agent-based model to learn dynamic incentive policies and explore adoption–cost trade-offs.
- PV uptake depends on financial returns, actor heterogeneity, behavioural factors, and social influences.
- Scenario-based ABMs capture heterogeneous adopters and interactions, while optimisation methods identify competing-objective trade-offs; both commonly treat policy as static.
- Sequential reinforcement learning adapts incentives to evolving costs, prices, and deployment outcomes, extending beyond its predominant use in energy-system operational control.
- The study combines RL with a stochastic ABM of yearly PV adoption among Irish dairy farms under uncertainty.
- A scalarised reward balances adoption and public cost, while varying the cost weight generates policies approximating an adoption–cost trade-off frontier.
- Across PPO, SAC, and TD3, stronger incentives increase adoption and expenditure, while dynamic policies explore a broader configuration range than static baselines.
2 Literature Review
Prior work uses ABMs to represent heterogeneous adoption and optimisation to study competing objectives, but these approaches generally do not model policy as sequential. This study positions scalarised RL within a stochastic ABM to learn adaptive policy trajectories and systematically explore adoption–cost trade-offs.
- ABMs represent solar PV uptake as a behavioural process involving heterogeneous, interacting decision-makers and multiple demographic, social, environmental, and economic factors.
- Feed-in-tariff effectiveness depends on policy design and market context, including tariff structure, differentiation, degression, deployment, and public cost.
- Multi-objective evolutionary algorithms such as NSGA-II recover non-dominated solutions for trade-offs across cost, adoption, sustainability, and related dimensions.
- RL has mainly addressed operational energy problems, though related work shows its feasibility for agent-based simulation and sequential public-policy interventions.
- Scalarisation combines competing objectives into one reward, and varying preference weights produces a family of policies approximating the adoption–cost trade-off.
- Relative to prior approaches, the study learns adaptive policy trajectories within a stochastic ABM rather than relying on static optimisation.
3 Methodology
The study formulates solar PV policy design as a finite-horizon sequential decision problem in which an RL policymaker interacts with a stochastic ABM under uncertainty. Continuous annual incentives are evaluated through adoption, cost, and policy-adjustment outcomes to learn preference-dependent policies and compare them with static baselines.
- Problem formulation: 16 years: The policymaker selects one continuous annual action containing grant share, subsidised loan rate, and feed-in tariff.Each action is applied for one year before the environment transitions.
- Learning and evaluation: Policy outcomes are evaluated under uncertainty to construct an adoption–cost trade-off frontier and examine dynamic configurations against static baselines.The framework uses annual RL actions, stochastic ABM outcomes, and offline evaluation of cumulative adoption and public cost.
- State representation: The observed state includes time, cumulative and new adoption, PV costs, electricity prices, cumulative public cost, and adoption rates across representative farm segments.The agent does not observe scenario-specific uncertainty draws, export shares, or segment-level latent parameters.
- Problem formulation: The stochastic ABM updates PV costs, electricity prices, population characteristics, and heterogeneous adoption dynamics after each policy action.Adoption decisions are generated probabilistically at the segment level using utilities based on techno-economic factors and policy incentives.
- Reward and objective: The scalar reward balances normalised adoption, public expenditure, and policy adjustment, with weights controlling the adoption–cost–stability trade-off.The policymaker maximises expected discounted return, while varying w_cost generates policies with different adoption and cost preferences.
- Learning and evaluation: PPO, SAC, and TD3 learn continuous-control policies, each trained for 1,000,000 timesteps across five random seeds and evaluated on 300 stochastic episodes.Evaluation also compares low-, moderate-, and high-support static baselines using cumulative adoption, total public cost, and cost per adopter.
4 Results
Across algorithms, training converges consistently and learned policies reveal a clear adoption–cost trade-off. Representative policies span low-cost, balanced, and high-adoption outcomes, while policy actions intensify over time in response to simulated adoption dynamics.
- Training Behaviour: Rewards increase rapidly and then stabilise across PPO, SAC, and TD3, with variability from stochastic simulation but no observed divergence.The trajectories indicate consistent convergence across seeds, suggesting suitability for the continuous stochastic optimisation setting.
- Adoption–Cost Trade-off: Higher adoption requires substantially greater public expenditure across PPO, SAC, and TD3, with the trade-off primarily driven by system dynamics.The consistency of the frontier indicates that the relationship is not mainly algorithm-specific.
- Adoption–Cost Trade-off: 2,682 adopters at about EUR 7.27 million defines the low-cost PPO policy, while 4,145 adopters at about EUR 41.73 million defines the highest-adoption TD3 policy.These representative configurations use wcost = 2.0 and wcost = 0.5, respectively.
- Comparison with Baselines: RL-derived policies cover distinct outcome regions relative to fixed low, moderate, and high-support baselines, including lower-cost, balanced, and high-adoption configurations.The representative policies are selected from the pooled frontier for outcome-based comparison rather than algorithm ranking.
- Adoption–Cost Trade-off: 3,495 adopters at about EUR 22.47 million characterises the balanced PPO policy.This representative policy uses wcost = 1.6 and lies between the low-cost and high-adoption regimes.
- Policy Dynamics and RL–ABM Interaction: The balanced policy increases grant share sharply after the early years, raises the feed-in tariff gradually to its maximum, and reduces the loan rate toward zero.These trajectories indicate adjustment of policy intensity as the ABM-simulated adoption dynamics evolve.
5 Discussion and Conclusion
The study finds that sequential RL enables dynamic policy adaptation and exposes adoption–cost trade-offs more flexibly than static configurations. Results are consistent across algorithms, but conclusions remain bounded by the model’s horizon, context, policy space, and behavioural assumptions.
- Dynamic policy adaptation enables more flexible exploration of adoption–cost trade-offs than static policy configurations.
- Stronger incentives increase adoption but require higher public expenditure, with the highest-adoption policies also having higher cost per adopter.
- Consistent trade-off structures across PPO, SAC, and TD3 suggest that system dynamics, rather than algorithm-specific behaviour, mainly drive the results.
- Evaluation outcomes are averaged over independent stochastic episodes to provide more stable estimates of expected adoption and public cost under uncertainty.
- The study is limited to a 16-year horizon, one context, a limited policy set, and an ABM that simplifies adoption behaviour and omits some constraints.
- The approach is presented as a practical framework for systematic trade-off evaluation, while resulting policies should be interpreted as model-based scenarios rather than definitive recommendations.