Source-linked AI summary

Dimension-Reduced ADP for Real-Time Microgrid Operation with Massive Air-Conditioning Loads under Multiple Uncertainties

Jingguan Liu, Xiaomeng Ai, Shichang Cui, Jiakun Fang, Jinyu Wen

arXiv:2609.01972v1eess.SY

TL;DR

Microgrid operation with massive air-conditioning loads must manage sequential decisions under uncertain wind, demand, and temperature while keeping computation tractable. The paper combines post-decision ADP with consistency-based state projection and piecewise linear approximation; tests report near-optimal performance, low computational cost, and scalability, with day-ahead participation as a stated limitation.

  • Problem

    Massive air-conditioning loads make ADP value-function approximation and offline training computationally intractable for real-time microgrid operation under multiple uncertainties.

  • Method

    The method combines a multi-stage ADP formulation, post-decision value functions, consistency-based aggregation of node-level air-conditioning states, and piecewise linear approximation trained from historical data.

  • Results

    The method achieves near-optimal operation with low computational cost and good scalability across deterministic and stochastic 33-bus and 123-bus case studies.

  • Takeaways & Limitations

    The approach provides an efficient and scalable framework for real-time utilization of regulation potential from massive air-conditioning loads.

  • Takeaways & Limitations

    The participating air-conditioning-load set is assumed to be determined day-ahead, leaving time-varying real-time participation for future work.

Abstract

from arXiv · show

This paper proposes a dimension-reduced approximate dynamic programming (ADP) method for real-time microgrid operation with massive air-conditioning loads under multiple uncertainties. The operation problem is formulated as a multi-stage Markov decision process, and a post-decision value function is introduced to characterize the impact of current decisions on future operating costs. To address the curse of dimensionality caused by massive air-conditioning loads, a consistency-based value function projection is developed to map the high-dimensional state space at each node into a tractable aggregated state space. Based on the reduced states, piecewise linear approximation is further employed for efficient value function training. Case studies on 33-bus and 123-bus systems show that the proposed method achieves near-optimal operation performance with low computational cost and good scalability under both deterministic and stochastic conditions.

A. Indices and Sets

The paper models real-time microgrid operation with renewable generation, flexible air-conditioning loads, and uncertainties revealed sequentially over time. It uses ADP to address the resulting high-dimensional decision problem.

  • Air-conditioning loads provide thermal-inertia-based flexibility for short-term power adjustment while preserving user comfort.
  • Wind power, load demand, and outdoor temperature are uncertain and progressively revealed during real-time operation.
  • Existing robust optimization, stochastic programming, and model predictive control approaches may fail to capture sequential decisions explicitly or become computationally burdensome under multiple uncertainties.
  • Massive air-conditioning populations make conventional ADP value-function approximation and offline training computationally intractable.
  • The proposed method projects node-level air-conditioning value functions into an aggregated state space while retaining sequential structure and supporting efficient approximation.

A. Thermal Dynamics of Air-Conditioning Loads

The formulation represents air-conditioned-room thermal behavior with a first-order equivalent thermal parameter model and embeds these dynamics in a constrained microgrid operation problem. The objective accounts for energy exchange, gas generation, air-conditioning response, and wind-curtailment costs.

  • Indoor temperature evolves according to a first-order equivalent thermal parameter model involving outdoor temperature, cooling power, thermal resistance, and room heat dissipation.
  • The continuous thermal model is discretized over short intervals assuming cooling power and outdoor temperature remain constant.
  • The operation model includes distributed gas turbines, the external grid, wind turbines, air-conditioning buildings, and rigid loads.
  • The expected operating-cost objective combines power-exchange, gas-turbine generation, air-conditioning response, and wind-curtailment penalty costs.
  • Air-conditioning power and temperature limits, linearized power-flow constraints, and equipment operating limits define feasibility.

III. DIMENSION-REDUCED ADP METHOD

The real-time problem is reformulated as a multi-stage Markov decision process with states, decisions, exogenous information, and transitions. A post-decision value function avoids explicit enumeration of uncertainty realizations in Bellman optimization.

  • Real-time control is determined from currently observed uncertainty realizations and system states in a rolling decision process.
  • The Markov decision process comprises state variables, decision variables, exogenous information, and transition functions.
  • Indoor air-conditioning temperatures form the state variables, while control actions include grid exchange, generation, demand response, curtailment, and network quantities.
  • Wind power, outdoor temperature, and rigid active and reactive load demand constitute the exogenous information.
  • The post-decision value function approximates expected future cost after decisions but before uncertainty realization, eliminating explicit enumeration of many uncertainty outcomes.

B. Dimension-Reduced Reformulation

The reformulation enforces consistent states and state derivatives for air-conditioning loads at each node, replacing many individual states with one aggregated state. This preserves thermal dynamics while reducing computational burden and supporting practical flexibility allocation.

  • Massive air-conditioning populations make the value-function space extremely high-dimensional, motivating a consistency-based dimension reduction.
  • The normalized state of comfort represents indoor temperature and is bounded by the comfort limits derived from the thermal model.
  • Air-conditioning loads connected to the same node share identical state-of-comfort values and derivatives under the consistency constraint.
  • The reformulation aggregates node-level air-conditioning states, projecting individual-load value functions into a single aggregated-state value-function space.
  • The projection preserves main thermal dynamics, reduces computational burden, and provides a favorable trade-off between solution quality and computational tractability.
  • A consistent SOC signal enables automatic and fair power-response allocation while simplifying operator-load information exchange.

C. Value Function Approximation Around Reduced States

After dimension reduction, the method approximates the post-decision value function with convex piecewise linear functions whose slopes are learned from historical data through temporal-difference learning.

  • Value-function representation: Convex piecewise linear functions approximate the post-decision value function after dimension reduction.The formulation represents the value function as a sum over piecewise linear segments.
  • Value-function representation: Each PLF segment has a slope parameter and an associated state-of-charge quantity allocated to that segment.The slopes characterize marginal values, while the allocated quantities represent SOC assigned across segments.
  • Offline training: Temporal-difference learning trains PLF slopes from historical uncertainty realizations and resulting operating decisions.The offline procedure observes uncertainties, solves the operating problem, estimates marginal contributions through perturbations, and updates slopes.
  • Offline training: The offline algorithm iterates through forward observations and backward slope updates before outputting the trained PLFs.Iterations initialize slopes, process historical samples, update slopes in a backward pass, and retain convexity before returning the functions.
  • Offline training: Slope updates combine current approximations with sampled marginal values using a step size while preserving PLF convexity.A concave adaptive value estimation algorithm is applied after slope updates to retain convexity.

D. Overall Application Process

The proposed application separates offline value-function training from real-time operation, where the trained approximation supports decisions through efficiently solved linear programs.

  • Overall application process: The method consists of offline training followed by real-time operation.Historical data train PLF slopes until convergence, and the resulting approximation is then used during operation.
  • Overall application process: During real-time operation, the operator observes current exogenous information and solves the operating formulation using the pre-trained value function.The decision problem is formulated as a linear program and can be solved with commercial optimization solvers.

IV. CASE STUDIES

The case studies validate the proposed method’s effectiveness and scalability using GUROBI on a computer with specified hardware.

  • Case studies: Three case studies evaluate the effectiveness and scalability of the proposed method.Simulations use GUROBI 12.0.1 on a 2.30 GHz CPU with 16 GB RAM.

A. Deterministic Case

In the deterministic 33-bus case, the proposed dimension-reduced ADP coordinates air-conditioning flexibility across the operating horizon and achieves near-optimal operation. It anticipates price and wind conditions to schedule cooling-energy storage and release strategically.

  • Case setup: The 33-bus test uses a 24 h horizon with 15 min resolution, 240 air-conditioning units, wind turbines, and gas turbines.The system includes two 2 MW wind turbines and four 0.5 MW gas turbines.
  • Performance: The proposed method reduces the solution gap to 1.50%, compared with 10.77% for myopic control and 7.52% for MPC.The Oracle case provides a hindsight-based lower bound and is infeasible in practice.
  • State trajectory: Its bus-13 SOC trajectory closely follows the Oracle trajectory, whereas myopic and MPC methods fail to prepare for wind accommodation and incur curtailment.The results indicate that the trained value function captures the impact of current decisions on future operating costs.
  • Energy coordination: Air-conditioning loads store cooling energy during 3:00–5:00 to absorb excess wind power, producing no wind curtailment in Case 2.The schedule also uses cooling-energy release to increase upward regulation capability.
  • Conclusion: Overall, the method effectively coordinates air-conditioning loads and achieves near-optimal economic operation.The comparison covers the full operation horizon rather than only local decisions.

B. Stochastic Case

In stochastic testing, the proposed ADP method delivers a small average solution gap and more concentrated performance across scenarios, supporting real-time operation under multiple uncertainties.

  • 2.10% average solution gap over 200 testing scenarios, significantly outperforming the other methods.
  • The solution gap distribution is more concentrated, indicating stronger performance under multiple uncertainties.
  • The method determines real-time control actions using only currently observed information because its value function captures prediction-error distributions from historical data.
  • The method achieves a much smaller solution gap than the myopic method.

C. Scalability Test

The scalability test evaluates the method on a larger 123-bus system with 500 air-conditioning units, where it achieves strong testing performance and computational efficiency.

  • The modified IEEE 123-bus system includes ten nodes with 50 air-conditioning loads each, totaling 500 units.
  • The scalability comparison includes traditional ADP without dimension reduction alongside the proposed method.
  • The proposed method achieves the smallest real-time testing gap among all practical methods.
  • The proposed method requires the shortest computation time, benefiting from dimension reduction.
  • The results demonstrate computational tractability, solution quality, and scalability in larger-scale systems.

V. CONCLUSION

The paper concludes that combining post-decision value functions with consistency-based state projection makes large-scale sequential microgrid operation tractable, while balancing quality, efficiency, and scalability.

  • The proposed method combines post-decision value functions with consistency-based state projection for microgrid operation with massive air-conditioning loads under multiple uncertainties.
  • The method makes sequential decision-making with large numbers of flexible loads computationally tractable.
  • The case studies show a good balance among solution quality, computational efficiency, and scalability under stochastic conditions.
  • Air-conditioning loads participating in demand response are assumed to be determined day-ahead.
  • Handling time-varying participation of air-conditioning loads in real-time operation is left for future work.
Loading 2609.01972v1…