Source-linked AI summary
q-Learning in Continuous Time
Yanwei Jia, Xun Yu Zhou
TL;DR
Conventional Q-learning loses action information in continuous time, creating a need for a suitable continuous-time analogue. The paper introduces a first-order q-function and martingale-based learning theory for entropy-regularized diffusion RL, yielding on-policy and off-policy actor-critic algorithms. The framework recovers SARSA, connects policy improvement to Hamiltonian-based Gibbs sampling, and supports continuous-time learning without time discretization.
Problem
The conventional Q-function collapses to the action-independent value function in continuous time, so it cannot rank current actions without time discretization or action restrictions.
Method
The paper defines a first-order q-function and jointly characterizes it with the value function through martingale conditions under enlarged filtrations for on-policy and off-policy learning.
Results
The resulting algorithms recover and interpret SARSA, support Gibbs-based policy improvement, and include actor-critic procedures that can operate without time discretization.
Takeaways & Limitations
The q-function supplies an action-sensitive continuous-time object whose Hamiltonian component supports policy updating and links continuous-time RL with classical Boltzmann exploration.
Takeaways & Limitations
Convergence-rate analysis remains open because continuous state spaces, nonlinear function approximation, and differing Q-function behavior prevent direct extension of existing discrete-time results.
Abstract
from arXiv · showhide
We study the continuous-time counterpart of Q-learning for reinforcement learning (RL) under the entropy-regularized, exploratory diffusion process formulation introduced by Wang et al. (2020). As the conventional (big) Q-function collapses in continuous time, we consider its first-order approximation and coin the term ``(little) q-function". This function is related to the instantaneous advantage rate function as well as the Hamiltonian. We develop a ``q-learning" theory around the q-function that is independent of time discretization. Given a stochastic policy, we jointly characterize the associated q-function and value function by martingale conditions of certain stochastic processes, in both on-policy and off-policy settings. We then apply the theory to devise different actor-critic algorithms for solving underlying RL problems, depending on whether or not the density function of the Gibbs measure generated from the q-function can be computed explicitly. One of our algorithms interprets the well-known Q-learning algorithm SARSA, and another recovers a policy gradient (PG) based continuous-time algorithm proposed in Jia and Zhou (2022b). Finally, we conduct simulation experiments to compare the performance of our algorithms with those of PG-based algorithms in Jia and Zhou (2022b) and time-discretized conventional Q-learning algorithms.
1 Introduction
The paper addresses the collapse of conventional Q-functions in continuous time by developing a first-order q-function theory for entropy-regularized diffusion RL, with martingale characterizations and resulting actor-critic algorithms.
- Motivation: Continuous-time Q-learning lacks an action-sensitive conventional Q-function because the discrete-time state-action value collapses to the action-independent value function.Time discretization is possible but can materially affect continuous-time Q-learning, motivating a formulation that remains continuous in time.
- q-function: The paper defines the little q-function as the first-order approximation of the conventional Q-function and relates it to the Hamiltonian and instantaneous advantage rate.In the entropy-regularized diffusion setting, the q-function combines the Hamiltonian with temporal dispersion terms from value-function changes and discounting.
- Theory: A given stochastic policy’s q-function and value function are characterized through martingale conditions in both on-policy and off-policy settings.The characterization uses an enlarged filtration containing environmental noise and policy randomization, and the functions can also be determined jointly through martingality.
- Algorithms: The martingale-based theory yields actor-critic algorithms whose forms depend on whether the Gibbs measure’s normalizing constant or density can be computed explicitly.The framework applies to online and offline model-free RL and learns value functions and stochastic policies simultaneously and alternatingly.
- Connections: The resulting temporal-difference algorithm recovers and interprets the classical SARSA algorithm, while the framework also provides new martingale conditions for discrete-time Q-learning.The paper presents this as evidence that continuous-time RL can offer new perspectives on discrete-time methods.
- Policy improvement: A Gibbs sampler using a properly scaled current q-function as the density exponent improves the current policy, linking continuous-time policy improvement to Boltzmann exploration.The q-function’s policy-relevant component is the Hamiltonian, which connects the method to classical stochastic control.
2 Problem Formulation and Preliminaries
The paper formulates entropy-regularized reinforcement learning for diffusion processes with stochastic policies, continuous-time state dynamics, and rewards observed through trial-and-error interaction.
- 2.1 Classical model-based formulation: The underlying control problem models a diffusion state driven by drift b and volatility σ, with actions chosen over a continuous time horizon.The objective maximizes expected discounted instantaneous rewards and a terminal lump-sum reward under standard stochastic-control assumptions.
- 2.1 Classical model-based formulation: The Hamiltonian and HJB equation characterize optimal control by combining instantaneous reward with the action’s risk-adjusted impact on system dynamics.The optimal feedback policy maximizes the Hamiltonian at each time-state pair under the model-based formulation.
- 2.1 Classical model-based formulation: Well-posedness is supported by continuity, Lipschitz, linear-growth, and polynomial-growth conditions on the dynamics and reward functions.These assumptions are adopted from Jia and Zhou (2022b) for the stochastic control problem.
- 2.1 Classical model-based formulation: The classical formulation assumes known dynamics and reward primitives, whereas reinforcement learning removes that knowledge and relies on trial-and-error observations.The agent observes state processes and payoffs while updating actions and policies from collected data.
- 2.2 Exploratory formulation in reinforcement learning: In the RL formulation, a stochastic policy samples actions from a time-state-dependent distribution using randomization independent of the environmental Brownian motion.The filtration is enlarged to include policy randomization, and the resulting action process generates a random-coefficient SDE.
- 2.2 Exploratory formulation in reinforcement learning: The framework supports both offline learning from repeated full-horizon trajectories and online learning that updates actions from historical observations.This scope covers the two principal data-collection settings considered for the exploratory RL problem.
- 2.2 Exploratory formulation in reinforcement learning: Entropy regularization adds an exploration term to the reward, with γ serving as the weighting or temperature parameter.The expectation accounts for both Brownian noise and action randomization.
- 2.2 Exploratory formulation in reinforcement learning: The exploratory formulation has both theoretical and computational roles: an averaged SDE supports analysis, while observable sample trajectories support learning algorithms.The two formulations are mathematically equivalent in value but serve different purposes for theory and algorithm design.
3 q-Function in Continuous Time: The Theory
The paper replaces the collapsing continuous-time Q-function with a time-discretization-independent q-function, then characterizes q-functions and value functions through martingale conditions for learning and policy improvement.
- 3.1 Q-function: The Δt-based Q-function equals the policy value plus a first-order term containing the Hamiltonian and temporal value change, with an o(Δt) residual.Only the Hamiltonian in the first-order term depends on the current action; temporal change includes value-time variation and discount depreciation.
- 3 q-Function in Continuous Time: The Theory: The continuous-time Q-function collapses to the value function as the action-dependent interval shrinks, motivating a first-order q-function instead of upfront time discretization.The paper distinguishes its continuous-time parameter Δt from discrete-time discretization and reaches the same limiting object through a direct continuous-time analysis.
- 3.2 q-function: The q-function is the instantaneous advantage rate and, in the exploratory diffusion setting, refines the conjecture that the continuous-time Q-function is the Hamiltonian.It is defined independently of time discretization and provides action-sensitive information for continuous-time policy improvement.
- 3.2 q-function: Martingale conditions characterize the q-function and value function jointly in on-policy and off-policy settings, forming the theoretical basis for learning algorithms.The conditions use an enlarged filtration containing environmental and action-generation randomness, and terminal conditions identify the correct value and q-functions.
- 3.3 Optimal q-function: The exploratory HJB relation alone is insufficient to identify the optimal q-function when model primitives are unknown, so martingale conditions are required.The HJB relation can determine the optimal value function through its PDE relationship, but q* requires additional characterization in the model-free setting.
- 3.3 Optimal q-function: A Gibbs policy generated from q-values improves a policy, while policy-independent martingale constraints enable off-policy learning of the optimal value function and q-function.When the Gibbs measure is generated by the optimal q-function, the resulting pair is optimal; the off-policy foundation avoids iterative policy improvement.
4 q-Learning Algorithms When Normalizing Constant Is Available
With an explicitly computable Gibbs normalizing constant, the paper designs on-policy and off-policy actor–critic q-learning algorithms by enforcing martingale conditions for value and q-function approximators. The resulting framework includes continuous-time methods related to SARSA and time-discretization-independent learning.
- Algorithm design: The algorithms jointly learn value and q-function approximators, update actor and critic through martingale conditions, and support both on-policy and off-policy settings.The construction is based on Theorems 7 and 9 and uses parameterized approximators of J and q.
- Computable normalizer: When the Gibbs normalizing constant is computable, the q-function directly determines a policy through the Gibbs density, enabling actor–critic updates from learned q-values.The policy constraint is automatically satisfied for the specified q-induced policy form.
- q-learning interpretation: The q-function has a dual actor–critic role: it is determined by the value function as a critic and derives an improved policy as an actor.This motivates calling the methods actor–critic despite the actor not being purely exogenous.
- Algorithm design: The framework offers martingale-loss, stochastic-approximation, and GMM-style procedures, including offline full-trajectory methods and online or offline test-function algorithms.Full-trajectory gradient updates are analogous to gradient Monte Carlo or TD(1), while test-function choices yield TD-style q-learning variants.
- Connections with SARSA: The continuous-time method learns the zeroth- and first-order terms of the discretized Q-function separately, making both learned quantities independent of the time step.This contrasts with discrete-time advantage-based approaches whose behavior depends on time discretization.
- Connections with SARSA: One q-learning algorithm recovers a modification of SARSA, while its comparison with modified SARSA differs by a mean-zero action-randomization term.The paper states that this makes the modified SARSA algorithm noisier and potentially slower to converge.
5 q-Learning Algorithms When Normalizing Constant Is Unavailable
When the Gibbs normalizing constant is unavailable, the paper replaces the intractable policy family with tractable densities and derives policy-improvement and actor–critic procedures using trajectory-level q-values. This construction also recovers the continuous-time policy-gradient updates of Jia and Zhou (2022b).
- Motivation: High-dimensional Gibbs normalizing constants can be daunting or impossible to compute, so the paper introduces tractable policy-density families to address this limitation.The proposed family includes distributions such as multivariate normals whose densities and normalizing constants are computable.
- A stronger policy improvement theorem: The stronger policy-improvement result guarantees that an updated tractable policy improves the current tractable policy under the stated update condition, even if it does not reach the desired target policy.The theorem compares policies generally, while its algorithmic implication restricts both policies to a tractable family.
- A stronger policy improvement theorem: The method need not learn the full q-function associated with a tractable policy; temporal-difference learning can estimate only q-values along sampled trajectories.Trajectory-level estimation is presented as easier than recovering the full functional form.
- Connection with policy gradient: Using a learned value-function approximator, the policy update becomes the entropy-regularized continuous-time policy-gradient rule of Jia and Zhou (2022b).Value learning may use martingale conditions, whereas policy learning does not require them in this construction.
- Connection with policy gradient: The paper provides a theoretically justified continuous-time counterpart to the soft Q-learning and policy-gradient equivalence known in discrete time.The resulting actor–critic algorithms recover the policy-gradient-based algorithms of Jia and Zhou (2022b) when combined with suitable policy-evaluation methods.
6 Extension to Ergodic Tasks
The paper extends q-learning to ergodic diffusion tasks with infinite-horizon average-reward objectives. It characterizes ergodic value, value-function, and q-function quantities and derives corresponding learning algorithms connected to SARSA and policy-gradient methods.
- Ergodic setting: Ergodic tasks maximize a regularized long-run average over an infinite horizon, with time-homogeneous dynamics and no terminal payoff.The objective is formulated for stationary admissible policies.
- Ergodic setting: In the ergodic formulation, the value is a scalar independent of initial state, while the value function is defined only up to an additive constant.The long-run average does not depend on initial state or time, whereas adding a constant to J leaves the solution valid.
- Ergodic q-learning theory: An ergodic theorem characterizes the value, value function, and q-function through martingale conditions for a policy, including corresponding on-policy and off-policy cases.The theorem gives conditions for identifying these quantities and includes an optimality implication under an additional policy relation.
- Ergodic q-learning theory: Under the theorem’s additional condition, the policy is optimal and the associated scalar value is the optimal value.The optimality conclusion is stated for the policy satisfying the displayed condition.
- Ergodic algorithms: The resulting ergodic q-learning algorithms learn the value, value function, and q-function simultaneously, with connections to SARSA and the policy-gradient algorithms of Jia and Zhou (2022b).An online algorithm is presented as an example, while further algorithmic details are related to the episodic case.
7 Applications
The applications evaluate q-learning against policy-gradient and time-discretized Q-learning methods in mean–variance portfolio selection and ergodic control. Across online and off-policy experiments, q-learning is competitive or more stable, while conventional Q-learning is often inferior or sensitive to discretization.
- 7.1 Mean–variance portfolio selection: The algorithms learn portfolio policies by parameterizing the value and q-functions and jointly updating their parameters with the Lagrange multiplier for the expected-return constraint.The portfolio objective minimizes terminal wealth variance while maintaining a target expected return, and the full stochastic-approximation procedure is summarized in Algorithm 5.
- 7.1 Mean–variance portfolio selection: The portfolio experiments use 20 years of training data, 20,000 episodes with batch size 32, and 100 out-of-sample evaluations per market configuration.Market configurations vary µ over {0, ±0.1, ±0.3, ±0.5} and σ over {0.1, 0.2, 0.3, 0.4}.
- 7.1 Mean–variance portfolio selection: Q-learning is almost always the worst performer across mean, variance, and Sharpe-ratio metrics, and can diverge in high-volatility, low-return environments.The q-learning and policy-gradient methods have similar overall performance, although q-learning tends to outperform in high-volatility markets; q-learning can have higher terminal variance when |µ| is small or σ is large.
- 7.2 Ergodic linear–quadratic control: The ergodic-control comparison evaluates q-learning against policy gradient, conventional Q-learning, and omniscient reward benchmarks using repeated online trajectories.The benchmarks represent optimal average reward with perfect environmental knowledge and the corresponding level after entropy-regularized exploration costs.
- 7.2 Ergodic linear–quadratic control: In online ergodic control, policy gradient and q-learning converge to the optimal reward level much faster than Q-learning, with q-learning slightly outperforming policy gradient.Q-learning has the slowest convergence; q-learning's parameter learning is smoother, whereas policy gradient initially learns faster but overshoots before correcting.
- 7.2 Ergodic linear–quadratic control: Reducing the time step severely degrades conventional Q-learning, which shows almost no improvement at ∆t = 0.01, while policy gradient and q-learning remain robust.The experiment compares ∆t ∈ {1, 0.1, 0.01} using fixed learning rates and initializations across 100 repetitions.
- 7.3 Off-policy ergodic linear–quadratic control: The q-learning algorithm is the most stable off-policy method and converges in all tested scenarios, whereas policy gradient diverges quickly and Q-learning remains time-discretization-sensitive.With coarse ∆t = 1, Q-learning and q-learning converge to the same limits, but these differ from the continuous-time optimal parameters because finite sums poorly approximate the theoretical integrals.
8 Conclusion
The paper develops continuous-time q-learning as a missing foundation for policy improvement, characterizing q-functions and value functions through martingale conditions in on- and off-policy settings. It connects this theory to actor–critic algorithms, SARSA, policy-gradient methods, and Hamiltonian-based model-free optimization.
- 8 Conclusion: Continuous-time q-learning fills a theoretical gap by addressing general policy improvement in both on-policy and off-policy settings.
- 8 Conclusion: The continuous-time q-function’s policy-relevant component is the Hamiltonian, connecting the theory to entropy-regularized stochastic control and Boltzmann exploration.The q-function is the first-order approximation that retains action information when the conventional Q-function collapses to the value function.
- 8 Conclusion: The q-function is characterized as the compensator that preserves martingality for a process combining value and cumulative reward, enabling simultaneous q-function and value-function learning.The resulting temporal-difference algorithm links to SARSA in classical Q-learning.
- 8 Conclusion: Convergence-rate analysis remains an outstanding problem because continuous state spaces, nonlinear function approximations, and differing Q- versus q-function behavior prevent direct extension of discrete-time results.The paper identifies stochastic approximation as a possible direction for future analysis.
- 8 Conclusion: The theory supports model-free policy learning by estimating the Hamiltonian rather than separately estimating every model coefficient, potentially reducing over-parameterization and error sensitivity.The conclusion contrasts this with model-based approaches that first estimate a full model and then optimize over it.
Appendix A1. Q-function associated with an arbitrary policy
This appendix characterizes the q-function and value function associated with an arbitrary policy through martingale conditions, then derives learning updates and links them to established RL algorithms.
- Relation to discrete-time quantities: The q-function is the continuous-time analogue of an advantage function, while the appendix relates it to the conventional Q-function and value function through entropy regularization.The q-function is described as an advantage rate function in continuous time, and the entropy term links Q and the value function.
- Algorithmic consequences: The resulting parameter updates recover classical SARSA when suitable stochastic-approximation methods are applied to the martingale equation.The appendix also distinguishes little q-learning from conventional Q-learning based on the time-discretized Q-function Q_∆t, which is used for simulation comparisons.
- Martingale characterization: The q-function and value function can be jointly characterized by martingality under an enlarged filtration that includes environmental noise and action randomization.The filtration is larger than the usual historical information set because it includes the current randomized action.
- On-policy and off-policy learning: Using the advantage function as the compensator makes the martingale condition valid for both target-policy and behavior-policy sampling.This provides the stated reason that the resulting Q-learning construction supports on-policy and off-policy learning.
- Comparison with policy gradients: The same martingale perspective does not generally support policy-gradient learning off policy, because its corresponding process is guaranteed only under on-policy sampling.The analysis also shows that choosing the correct filtration and test functions is essential for martingality.
- Algorithmic consequences: Algorithm design branches according to whether the Gibbs-policy normalizing constant is explicitly computable, with approximating policy families used otherwise.The same distinction is used for parameterized policies and q-functions in the appendix’s algorithmic constructions.
Proof of Theorem 2
The proof establishes that entropy-regularized policy improvement is optimality-preserving: the Gibbs policy generated from a q-function improves value, and a fixed point satisfies the optimality equations.
- Entropy maximization: The entropy-maximizing density is the unique optimizer of the action-distribution problem induced by a measurable q-function.The proof obtains the optimizer by relaxing the density constraint and then verifies feasibility and uniqueness.
- Optimality: The proof connects the continuous-time improvement argument to the q-function framework through the Hamiltonian and the associated martingale construction.The appendix explicitly situates the result within the q-function proof strategy rather than a time-discretized argument.
- Policy improvement: For any admissible policy, the Gibbs-improvement policy has value at least as large as the original policy’s value.The argument uses Itô’s lemma, localization, and the entropy-regularized Hamiltonian inequalities.
- Optimality: If the improvement map leaves a policy unchanged, its value function satisfies the HJB equation and therefore equals the optimal value function.The fixed-point condition supplies both the PDE and HJB characterizations needed for optimality.
Proof of Theorem 6
Theorem 6 is proved by showing that a candidate q-function makes a discounted stochastic process a martingale, with converse arguments recovering the policy’s value and q-function from that condition.
- Martingale equivalence: If the candidate q-function matches the policy’s q-function, Itô’s formula makes the associated discounted process an enlarged-filtration martingale.The converse uses the fact that a finite-variation local martingale must vanish, yielding the pointwise q-function identity.
- Pointwise identification: Continuity and full-support admissibility extend the martingale-derived identity from almost-everywhere statements to every time, state, and action.The contradiction argument uses continuity and the positive probability of neighborhoods under full-support policies.
- Joint characterization: A jointly martingale pair of candidate value and q-functions satisfies the policy-evaluation PDE and therefore coincides with the policy’s actual value and q-functions.Uniqueness of the Feynman–Kac PDE supplies the value-function identification before the q-function identity is recovered.
- Optimality: When the candidate q-function induces the policy through the Gibbs map, the martingale characterization implies that the induced policy is optimal.The proof combines the policy-evaluation result with the optimality result for the improvement map.
- Policy improvement: The proof also uses KL-divergence comparisons to show that the induced policy improves value relative to the original policy.This supplies an alternative value-improvement step within the theorem’s broader policy-improvement argument.
1 Introduction
The introduction identifies a measurability gap in continuum action sampling for prior continuous-time q-function theory and adopts discretely sampled actions to obtain well-defined controlled diffusions.
- Prior framework: Jia and Zhou (2023) develop martingale characterizations for continuous-time q-learning but implicitly assume continuum independent action sampling from an admissible policy.At each time–state pair, the construction requires independent draws from the policy’s non-degenerate action distribution.
- Measure-theoretical issue: The continuum-sampling construction does not generally ensure that the action process is progressively measurable, so the drift and stochastic integrals may be undefined.Szpruch et al. (2024) and Bender and Thuan (2024) are cited for this technical issue.
- Motivation: The authors characterize this as a delicate technical gap while maintaining that the underlying theoretical results are important enough to warrant correction.The introduction frames the issue as motivating an erratum rather than rejecting the prior theory wholesale.
- Resolution: Several works propose time-discretely sampled action processes as a remedy, and this treatment follows the framework of Jia et al. (2025).The adopted construction replaces continuum action draws with samples taken on a time grid.
- Resolution: Independent random variables generate grid-point actions on a product probability space, while the resulting piecewise-constant action process is adapted to a suitable filtration.The construction separates Brownian environmental noise from the random variables used to generate actions.
- Resolution: The resulting state process is a well-posed SDE with continuous trajectories under the discretely sampled action policy.The dynamics can be written using the grid projection δ(s), which holds the sampled action constant between grid points.
2 Martingale Characterizations for q-Learning with Discretely Sampled Processes
The section revises q-learning’s martingale characterizations for discretely sampled state–action processes while keeping the q-function defined independently of discrete sampling. The results cover on-policy and off-policy identification and characterize optimal policies through the q-function.
- Setup: The revised theory defines the q-function from the exploratory problem, independently of any discrete sampling scheme.The value function is likewise based on the exploratory problem, while discretely sampled processes are used in the martingale characterizations.
- Theorem 6: If the candidate q-function equals the policy-associated q-function, the martingale conditions hold for discretely sampled state–action processes.Conversely, the martingale conditions recover the q-function under the stated initial-condition and time-grid requirements.
- Theorem 6: Martingale conditions characterize a policy’s value function and q-function jointly, including both on-policy and off-policy cases.The theorem’s three cases identify the functions under the target policy, under another policy, or from a suitable off-policy martingale condition.
- Optimality: When the policy density is proportional to the exponential of the q-function, the policy is optimal and the associated value function is optimal.This conclusion is obtained in the theorem’s additional policy-improvement condition.
Proof
The proofs derive the martingale characterizations using Itô’s lemma, moment estimates, zero quadratic variation, and uniqueness of the Feynman–Kac PDE. The optimality results then follow by applying the preceding characterization and Gibbs-policy arguments.
- Theorem 7 proof: Uniqueness of the Feynman–Kac PDE identifies the candidate value function with the policy value function and then identifies the corresponding q-function.The PDE is obtained from the martingale constraints and the terminal condition.
- Scope of revision: The revised theoretical results do not change the algorithms or numerical experiments because those procedures already use discretely sampled state processes.The continuous-time state process remains continuous; discretely sampled actions are used from the policy, and convergence to the exploratory-problem value function is established as grid size decreases.