Source-linked AI summary
Optimal control of end-user energy storage
Peter M. van de Ven, Nidhi Hegde, Laurent Massoulie, Theodoros Salonidis
TL;DR
The paper studies how end users can use batteries to exploit fluctuating energy prices without shifting demand. It models storage control as a stochastic optimization problem, proves a threshold-based optimal policy, and reports substantial numerical cost savings.
Problem
Users face fluctuating energy prices but generally show only minor demand shifts, motivating storage-based cost minimization under stochastic prices and demands.
Method
The paper models battery purchasing and discharge decisions as a Markov decision process and derives a stationary two-threshold policy.
Results
38% relative energy savings are achieved at most as battery size increases, with savings saturating at Bmax = 16 kWh.
Takeaways & Limitations
Energy storage can produce significant cost savings, while the optimal policy has a simple threshold structure.
Takeaways & Limitations
The study focuses on a single user minimizing its own costs; multi-user peak-load effects remain future research.
Abstract
from arXiv · showhide
An increasing number of retail energy markets show price fluctuations, providing users with the opportunity to buy energy at lower than average prices. We propose to temporarily store this inexpensive energy in a battery, and use it to satisfy demand when energy prices are high, thus allowing users to exploit the price variations without having to shift their demand to the low-price periods. We study the battery control policy that yields the best performance, i.e., minimizes the total discounted costs. The optimal policy is shown to have a threshold structure, and we derive these thresholds in a few special cases. The cost savings obtained from energy storage are demonstrated through extensive numerical experiments, and we offer various directions for future research.
I. INTRODUCTION
Dynamic pricing exposes users to time-varying energy costs, but limited demand shifting motivates battery storage as a way to exploit low-price energy without changing consumption. The paper models this stochastic storage-control problem and establishes a cost-minimizing threshold policy.
- Motivation: Dynamic pricing varies by time of day or wholesale-market conditions, creating opportunities to reduce energy costs.Time-of-use pricing uses scheduled price levels, while real-time pricing changes hourly or half-hourly.
- Motivation: Battery storage lets users charge when energy is inexpensive and discharge when prices are high, avoiding demand shifts.Storage may come from dedicated batteries, electric-vehicle battery packs, or data-center backup supplies.
- Model and approach: The paper formulates energy-storage purchasing as a Markov decision process with stochastic prices, demands, and modulating states.Demand and price may be correlated, and the model includes battery charging and discharging inefficiencies.
- Model and approach: A stationary two-threshold policy minimizes discounted costs: charge below the lower threshold and discharge above the upper threshold.The decisions include direct grid purchases, battery charging, and battery discharge used to satisfy demand.
- Relation to prior work: Unlike related storage models, this work incorporates battery inefficiencies and investigates optimal scheduling rather than a sub-optimal heuristic.The model also uses periodicity to represent daily price and demand fluctuations.
III. THE STRUCTURE OF THE OPTIMAL POLICY
The paper formulates battery control as a discounted-cost Markov decision process and proves that the optimal stationary policy uses two state-dependent battery thresholds. Battery levels below the lower threshold trigger charging, levels above the upper threshold trigger discharging, and intermediate levels are left unchanged; with perfect efficiency, the thresholds coincide.
- Model and optimality: The discounted-cost control problem is expressed through a Bellman equation over feasible next battery levels, with a stationary optimal policy attained.The action is the next battery level β, constrained by the control set Ux(b).
- Threshold policy: When the battery level lies between the thresholds, the battery is neither charged nor discharged and demand is supplied from the grid.At sufficiently low levels the battery is charged, while at high levels demand is partially met from the battery.
- Proof structure: Convexity and monotonicity of the cost function support the threshold structure of the optimal policy.Lemma 1 states that Jx(b) is convex and non-increasing in battery level; the theorem uses the resulting convex optimization problem and its subdifferential.
- Threshold policy: For each state, the cost-minimizing policy has two battery thresholds that determine the optimal next battery level.The policy charges toward the lower threshold, holds the battery between thresholds, and discharges toward the upper threshold.
- Special case: With a fully efficient battery, the charging and discharging thresholds are identical.The figure description likewise states that the two horizontal threshold lines coincide when ηdηc = 1.
IV. BATTERY LEVEL THRESHOLDS
The optimal policy has a threshold-based battery structure, with analytical conditions identifying when thresholds are empty or full and special cases determining how thresholds vary with prices. Price monotonicity holds for i.i.d. prices but can fail under Markovian price dynamics.
- Threshold structure: The paper establishes a threshold-based structure for the optimal battery policy and presents structural results and special cases.The general problem is difficult; the analysis first characterizes threshold behavior before treating specific conditions.
- Boundary thresholds: Thresholds equal 0 or ¯B under sufficient conditions comparing battery states across possible successor states.These conditions provide general criteria for always maintaining an empty or full target battery level.
- Boundary thresholds: βx = 0 when the current price is ¯P, or when pmin > 0 and α < pmin/¯P.At the highest price, or under the stated discount and price condition, the optimal policy does not charge the battery.
- Price-dependent thresholds: When transitions are i.i.d. or price-dependent, βx depends only on price; for i.i.d. prices, thresholds decrease as price increases.This monotonicity is stated formally for i.i.d. prices and demands.
- Price-dependent thresholds: Markovian price dynamics can break monotonicity: with four price levels, β1 = β3 = 1 and β2 = β4 = 0.The example uses α ≥ 3/4, ¯B = 1, D ≡ 1, and specified price-transition probabilities.
V. NUMERICAL EVALUATION
The numerical evaluation examines the optimal energy-storage policy in residential real-time-pricing settings. It measures practical feasibility and cost savings for individual home storage and shared energy storage under demand and price fluctuations.
- Evaluation goals: The evaluation studies the optimal policy in residential environments and real-time pricing scenarios.Its goals are to assess practical feasibility and quantify cost savings under realistic demand fluctuations.
- Evaluation goals: The experiments evaluate cost savings for both individual home storage units and shared energy storage.The evaluation uses scenarios involving residential storage configurations.
A. Price and demand datasets
The evaluation uses historical Ontario hourly spot prices as a residential real-time-pricing proxy and synthetically generated household demand. Thresholds are learned from January 2011 hourly empirical distributions and evaluated on February 2011 data.
- Datasets: Historical hourly Ontario spot prices approximate residential real-time pricing, while synthetic demand models occupancy, appliances, weather, and household characteristics.Ontario did not use real-time pricing, so its spot prices serve as an estimate for the experiments.
- Datasets: The representative scenario uses a four-occupant home with a battery and Ontario market prices from January and February 2011.The reported setup focuses on one residential scenario for concise presentation.
- Price patterns: January Ontario prices show multiple daytime peaks between 9 a.m. and 10 p.m. and lower prices at night; February follows a similar trend.The price series therefore contains repeated daily variation for storage operation.
- Threshold estimation: The method computes separate empirical price and demand distributions for each hour using 31 January observations.These hourly distributions determine the optimal-policy thresholds.
- Threshold estimation: February 2011 prices and demands are then used to emulate policy operation and compute resulting electricity cost.January data determine thresholds, while February data provide the evaluation period.
B. Implementation
The implementation computes optimal thresholds numerically and evaluates storage across battery sizes, prices, demand timing, and pooled-user configurations. Storage reduces peak-period purchases and achieves up to 38% savings, with savings saturating at 16 kWh.
- Implementation: Policy iteration numerically computes thresholds after discretizing demand, storage, and price states.The implementation uses 0.5 kWh state increments, 5 ct price increments, and hourly slots.
- Energy savings: 38% maximum relative energy savings are achieved over no storage, with savings saturating at Bmax = 16 kWh.The evaluation assumes a fully efficient battery without charging constraints, so the savings provide an upper bound.
- Threshold behavior: Lower prices produce lower optimal thresholds throughout the day, while the 15 ct threshold reaches zero at 11 p.m. before rising again.Thresholds peak early in the morning when prices are low.
- Purchase timing: Storage shifts energy purchases toward early morning, whereas users without storage purchase more during high-price peak hours.Figure 6 compares average energy bought by hour with and without optimal storage.
- Resource pooling: Pooling compares separate 16 kWh batteries per user with one shared 16n kWh battery using aggregate demand and shared thresholds.The comparison plots aggregate monthly costs without storage, without pooling, and with pooling against the number of users.
- Resource pooling: Negatively time-shifting users' demands produced similar performance because the common price signal primarily drives optimal-policy behavior.The authors report that this eliminates potential pooling benefits in the tested setting.
VI. EXTENSIONS & OUTLOOK
The paper outlines extensions for population-scale objectives, battery replacement, self-discharge, varying efficiency, and user generation. These extensions broaden the model while preserving or modestly modifying its threshold-based analysis, subject to stated scope boundaries.
- Scope and future applications: The model focuses on one user minimizing individual energy costs, while future work could study many users minimizing grid peak load.A fraction of battery-equipped users could apply the policy to shift demand from peak to off-peak hours.
- Battery replacement: Battery replacement costs require jointly optimizing battery size and storage management, and smaller batteries may exploit fewer price fluctuations.The preferred battery size may depend on the spread and volatility of energy prices.
- Battery replacement: Assuming breakdown after a geometric number of operations preserves most of the analysis, with modifications to the immediate cost function and optimality results.The modified cost function is nonconvex at zero action, requiring changes to Lemma 1 and Theorem 1.
- Battery dynamics: Self-discharge and time-varying efficiency can be incorporated by modifying storage evolution to dissipate a fraction ξ of stored energy each slot.The authors state that the remaining derivations and results continue to hold after adjustment.
- Energy generation: User generation extends the model to joint decisions about buying, selling, and storing energy, including renewable-energy applications.The demand variable is reinterpreted as demand minus generation, and nonnegativity constraints on energy actions are removed.
- Conclusion: The paper concludes that the cost-minimizing policy is threshold-based, produces significant numerical savings, and supports extensions without fundamentally altering the results.The conclusion summarizes the policy structure, numerical study, and model generalizations.
APPENDIX
The appendix establishes structural properties of the finite-horizon value functions and characterizes optimal battery targets through convex subgradient conditions. It also shows that a continuum of threshold choices can be optimal.
- Value-function structure: The finite-horizon costs converge to the infinite-horizon value function, so convexity and monotonicity can be established by induction.The base case is Jx,0 ≡ 0, and the induction uses the infimal convolution operator.
- Value-function structure: Convexity of γx and the induction hypothesis make the Bellman operator preserve convexity of Jx,n.The proof identifies the operator as an infimal convolution operator.
- Value-function structure: For b1 ≤ b2, the proof establishes Jx,n(b1) ≥ Jx,n(b2), showing that costs are non-increasing in the battery level.The argument distinguishes cases according to the minimizing target battery level β*x.
- Optimality conditions: Because γx and Jx are convex, selecting the next battery level β is a convex optimization problem characterized by a suitable subgradient of Hx.A minimizer is identified through subgradient inequalities over the feasible target set.
- Optimality conditions: The threshold sets are defined through subgradient sign conditions, and any choice in [β−1,x, β+2,x] yields an optimal policy.The proof therefore permits a continuum of optimal policies rather than a unique threshold choice.
C. Proof of Corollary 1
The corollary proof specializes the threshold characterization to unit charging and discharging efficiencies. Under this condition, the relevant threshold intervals coincide and the stated result follows.
- Proof of Corollary 1: When ηc = ηd = 1, the threshold intervals coincide, so the corollary follows directly from the preceding characterization.The proof explicitly invokes the interval [β−1,x, β+2,x].
D. Proof of Proposition 1
The proof of Proposition 1 verifies the boundary conditions for the optimal target battery level and uses finite-horizon induction to establish the required inequality. It treats both empty and full battery cases.
- Proof of Proposition 1: At an empty battery, the proof checks that βx = 0 by verifying the relevant inequality for every feasible target β.This establishes the lower-boundary condition used in the proposition.
- Proof of Proposition 1: At a full battery, the proof similarly verifies βx = B by checking the corresponding inequality over all feasible targets.This establishes the upper-boundary condition.
- Proof of Proposition 1: The proof seeks to establish the proposition's inequality for all b1 < b2 and every state x.It then applies induction on the finite horizon to prove the required relation.
- Proof of Proposition 1: Once the two cases satisfy the first condition of Proposition 1, the desired conclusion follows for both cases.The proof states that the condition is satisfied in cases (i) and (ii).
F. Proof of Proposition 3
For i.i.d. prices, the future-cost function is state-independent, so the threshold sets depend on the current state only through price. Consequently, the thresholds are non-increasing in price.
- Proof of Proposition 3: With i.i.d. prices, future prices do not depend on the current state, making G and its subdifferential independent of x.The subdifferential is written as ∂G(β) = [σ−(β), σ+(β)].
- Proof of Proposition 3: The threshold sets depend on the state only through p(x), and they are non-increasing in p(x).This follows from the non-decreasing behavior of σ− and σ+.
- Proof of Proposition 3: In the four-price example, states xi are defined by p(xi) = i for i = 1, …, 4, while the control set does not depend on battery level.The example uses this state indexing to verify the threshold values.
- Proof of Proposition 3: The example verifies β4 = 0 together with β1 = 1, β2 = 0, and β3 = 1 using the proposition's conditions.The proof checks the required conditions for these four threshold values.