Source-linked AI summary
A Multiresolution Stochastic Process Model for Predicting Basketball Possession Outcomes
Daniel Cervone, Alex D'Amour, Luke Bornn, Kirk Goldsberry
TL;DR
The paper tackles the limits of discretized basketball statistics by estimating expected points throughout a possession from optical tracking data. It combines continuous movement and discrete-event models in a multiresolution stochastic process, with hierarchical models sharing information across players. The resulting EPV estimates track how player locations and decisions change expected possession outcomes, including sharp changes around passes and shots.
Problem
Existing basketball evaluation models rely on reductive box-score summaries that omit high-resolution player movements and decisions.
Method
The paper combines continuous and coarsened possession models with hierarchical transition models to estimate EPV from tracking data.
Results
EPV changes with player locations and decisions; in the illustrated possession, a layup raises EPV from 1.52 to 1.62.
Takeaways & Limitations
EPV curves and their aggregations provide a way to analyze how movements and decisions contribute value across possessions, games, or seasons.
Takeaways & Limitations
The model uses a constant rebound-allocation probability, does not distinguish turnover types, and lacks some physical information such as players’ hand and foot positions.
Abstract
from arXiv · showhide
Basketball games evolve continuously in space and time as players constantly interact with their teammates, the opposing team, and the ball. However, current analyses of basketball outcomes rely on discretized summaries of the game that reduce such interactions to tallies of points, assists, and similar events. In this paper, we propose a framework for using optical player tracking data to estimate, in real time, the expected number of points obtained by the end of a possession. This quantity, called \textit{expected possession value} (EPV), derives from a stochastic process model for the evolution of a basketball possession; we model this process at multiple levels of resolution, differentiating between continuous, infinitesimal movements of players, and discrete events such as shot attempts and turnovers. Transition kernels are estimated using hierarchical spatiotemporal models that share information across players while remaining computationally tractable on very large data sets. In addition to estimating EPV, these models reveal novel insights on players' decision-making tendencies as a function of their spatial strategy.
1. INTRODUCTION
The paper addresses the limits of box-score evaluation by using optical tracking data to model basketball possessions continuously and estimate expected points in real time. Its multiresolution stochastic-process framework captures both player movement and discrete events, producing EPV curves that reflect how decisions change a possession’s value.
- Motivation: Box-score models omit high-resolution strategic actions such as deceiving defenders or passing instead of taking an open shot.Optical tracking data provide the basis for analyzing these continuous interactions and decisions.
- Expected Possession Value: EPV is the expected number of points the offense will score on a possession conditional on its evolution up to time t.The quantity is defined from the possession’s observed history rather than only its final box-score outcome.
- Illustration: In the example possession, EPV rises to about 1.34 when Norris Cole reaches the basket, falls to 0.77 after he dribbles past it, and resets to 1.00 after his pass to Rashard Lewis.The curve therefore tracks how locations, movements, and macrotransitions alter expected point yield over time.
- Illustration: After the ball reaches LeBron James, EPV rises to 1.03 and his layup increases it from 1.52 to 1.62.The example illustrates sensitivity to exact player locations and movement toward the basket.
- Method: The framework models possessions at continuous and highly coarsened resolutions, combining player movements with discrete events such as passes and shots.This multiresolution design is intended to make EPV estimation reliable, stochastically consistent, and computationally feasible.
2. MULTIRESOLUTION MODELING
The paper defines EPV as the expected points from a possession given tracking information through time t, then combines full-resolution and coarsened process models to estimate it tractably. The coarsened process represents meaningful possession events with semi-Markov states while retaining selected spatial and defensive summaries.
- Full-resolution process: The tracking process includes player and ball locations, game context, and real-time event annotations such as passes, turnovers, and shot attempts.These variables define the high-dimensional information used to condition EPV.
- EPV definition: EPV is the expected possession point total conditional on all optical-tracking information available through time t.The modeled possessions exclude fouls and therefore end with 0, 2, or 3 points.
- Multiresolution framework: The framework combines a fully continuous player-movement model with a Markov chain over a highly coarsened possession process.This combines fine-grained information with the computational simplicity of discrete transitions.
- Coarsened process: The coarsened state space summarizes ballhandler identity, seven court regions, and whether the ballhandler has a defender within 5 feet.End states encode made 2-point shots, made 3-point shots, or possession endings worth 0, 2, or 3 points.
- Coarsened process: Transition states represent constrained actions in progress, including shots, passes, turnovers, and rebounds, and preserve information such as the recent ballcarrier.The figure separates possession states, transition states, and end states.
- Model assumptions: The semi-Markov assumption makes the embedded sequence of disjoint coarsened states a Markov chain, enabling EPV calculation from a transition matrix.The model also assumes that transition states decouple later possession evolution from earlier history after the transition ends.
3. TRANSITION MODEL SPECIFICATION
The transition specification predicts how possessions evolve by modeling continuous player movement, macrotransition hazards, and the outcomes of passes, shots, and turnovers. Hierarchical spatiotemporal models estimate these components, while a coarsened transition matrix summarizes later possession evolution.
- 3. TRANSITION MODEL SPECIFICATION: The full-resolution models predict the next decoupling state, after which EPV depends only on the coarsened state.This switches from high-resolution conditioning to coarsened-state transitions at the endpoint of a pass, shot, or turnover.
- 3. TRANSITION MODEL SPECIFICATION: The framework alternates infinitesimal movement and macrotransition-entry models until a pass, shot, or turnover occurs, then predicts its outcome state.The movement model assumes no major ball movement during the infinitesimal step.
- 3.1 Microtransition Model: The microtransition model predicts player movement using velocity, acceleration, and spatial effects while holding the ballcarrier constant.Offensive dynamics include basket-directed acceleration beyond the three-point line and deceleration near the basket; defensive predictions combine dynamics with the guarded offensive player's induced path.
- 3.2 Macrotransition Entry Model: Macrotransition entry models use competing risks for four pass options, shot attempts, and turnovers.Each macrotransition type has a hazard that depends on the full-resolution possession history and ballcarrier location.
- 3.4 Transition Probability Matrix for Coarsened Process: Replacing raw transition counts with hazard-based conditional expectations yields a Rao-Blackwellized estimator that is unbiased and lower variance.Hierarchical hazard parameterization shares information across players while supporting estimation on large datasets.
4. HIERARCHICAL MODELING AND INFERENCE
The paper estimates rich player-, action-, and space-specific transition models using hierarchical Bayesian methods. Functional-basis representations and between-player priors share information while keeping spatial effects computationally tractable.
- Motivation: Hierarchical models are needed because EPV averages over future possession paths, including player-location situations with little or no observed data.For example, estimating DeAndre Jordan’s shooting ability across the court requires information beyond his observed three-point attempts.
- Player-specific modeling: Standard exchangeability assumptions are too strong for NBA players, so the models share information while preserving meaningful differences between players.The paper contrasts LeBron James and Steve Novak despite their shared listed position.
- Spatial effects: Spatial effects are represented with low-dimensional functional bases, improving computational efficiency and allowing non-stationary court-specific dependence.The basis representation replaces computationally costly Gaussian-process inference while retaining interpretable spatial patterns.
- Spatial effects: Basis functions are shared across players within each macrotransition type, while player-specific weights vary across both players and macrotransition types.These weights encode interpretable spatial motifs associated with players’ decision-making tendencies.
- Parameter estimation: The full likelihood is factored into partial likelihoods for microtransitions, macrotransition entries, and macrotransition exits, with the remaining term ignored during inference.This factorization separates the data informing movement, entry, and exit parameters.
- Parameter estimation: Shot-probability parameters are estimated with hierarchical logistic regression using INLA, while microtransition models are fit separately for each player because movement data are sufficiently informative.The shot model uses only time points at which shots are attempted, requiring fewer computational resources than the movement models.
5. RESULTS
The results show that the multiresolution EPV framework produces calibrated, interpretable possession and player analyses, with hierarchical shrinkage improving predictive performance and EPV-derived metrics revealing context-dependent value.
- 5.1 Predictive Performance of EPV: 90% of the data were used for parameter inference and 10% for out-of-sample evaluation of macrotransition models.
- 5.1 Predictive Performance of EPV: With shrinkage, the full hierarchical model consistently achieved the highest out-of-sample log-likelihood among the compared configurations.Without shrinkage, the full model sometimes performed worse than a model without spatial effects.
- 5.2 Possession Inference from Multiresolution Transitions: EPV represents a weighted average of the values of possible next macrotransitions, using multiresolution transition probabilities.Figure 7 combines short-horizon microtransitions with macrotransition arrows whose color and thickness encode value and probability.
- 5.2 Possession Inference from Multiresolution Transitions: Estimated transition probabilities and values align with basketball intuition: shots become more likely near the basket, while player-specific shooting skill affects shot value.LeBron James’s shot attempt is valued more highly than Norris Cole’s despite occurring farther from the basket.
- 5. RESULTS: EPV plots and diagrams can support offensive and defensive strategy by showing how movements and decisions contribute value across possessions.The analysis can be repeated across hundreds of thousands of tracked possessions in a season.
- 5.3 EPV-Added: EPVA evaluates offensive value relative to a hypothetical league-average player in the same situations, but excludes defense and off-ball contributions.The authors caution that EPVA is not a best/worst-player ranking and relies on extrapolation to a hypothetical baseline.
- 5.4 Shot Satisfaction: Shot satisfaction is expressed per shot and is highest for efficient three-point or near-basket shooters, while remaining mostly positive across the league.The metric analyzes shot selection and efficiency in their spatiotemporal contexts.
6. DISCUSSION
The paper argues that EPV captures offensive value omitted by box scores and can reveal how spatial actions and alternative possession paths affect expected scoring. It also identifies modeling assumptions, omitted information, and computational demands that limit interpretation and use.
- 6. DISCUSSION: EPV captures offensive schemes and player actions omitted by box-score summaries, including attacks, passes, and separation from defenders.It can also decompose value into weighted transition values, exposing probable alternative possession paths that were not realized.
- 6. DISCUSSION: Alternative possession paths can influence EPV even when they never occur, such as an open teammate positioned for a good shot.
- 6. DISCUSSION: The model’s marginal semi-Markov assumption is a first-order approximation that can be violated by preset plays linking sequences of runs and passes.The authors suggest wider macrotransition and coarsened-state sets to encode such playbook structure.
- 6. DISCUSSION: EPV omits defensive contributions, off-ball effects, player attributes absent from tracking, and distinctions among several rebound and turnover situations.The authors recommend accompanying EPV analyses with game film and basketball expertise.
- 6. DISCUSSION: Estimating EPV curves requires computational resources that may restrict its use to academic circles and professional basketball teams with appropriate infrastructure.
APPENDIX A. ADDITIONAL TECHNICAL DETAILS
The appendix provides additional technical details on fitting the paper’s multiresolution models and deriving basketball metrics from EPV estimates.
- APPENDIX A. ADDITIONAL TECHNICAL DETAILS: The appendix details model-fitting steps and the derivation of basketball metrics from EPV estimates.
A.1 Time-Varying Covariates in Macrotransition Entry Model
The macrotransition entry models use covariates tailored to event types, with some variables shared across players within each macrotransition type. These covariates describe dribbling, defender distance, ball movement, teammate closeness, and receiver openness.
- A.1 Time-Varying Covariates in Macrotransition Entry Model: The macrotransition entry formulation includes coefficients for pass events and situation covariates, with covariate selection varying by transition type.
- A.1 Time-Varying Covariates in Macrotransition Entry Model: The covariates may differ by macrotransition, but each transition type uses the same covariates across players.
- A.1 Time-Varying Covariates in Macrotransition Entry Model: The model defines dribble, nearest-defender distance, recent ball travel, teammate closeness, and receiver openness as situation covariates.
- A.1 Time-Varying Covariates in Macrotransition Entry Model: Pass transitions use dribble, defender distance, teammate closeness, and openness, while shot and turnover transitions additionally use recent ball travel.The shot probability model uses only dribble and defender distance; all models include an intercept.
A.2 Player Similarity Matrix H for CAR Prior
The CAR prior encodes player similarity through low-dimensional summaries of court occupancy distributions. Players with similar spatial roles are connected as neighbors so their macrotransition parameters can share information.
- A.2 Player Similarity Matrix H for CAR Prior: Player position is represented by a rank-r approximation of a 461 × 575 court-occupancy matrix, using r = 5.The matrix counts player locations across 575 four-square-foot offensive-court bins.
- A.2 Player Similarity Matrix H for CAR Prior: The factorization represents court occupancy distributions with nonnegative basis vectors and gives each player a low-dimensional position summary.The rows of V are basis vectors, while each row of U summarizes where a player spends time on the court.
- A.2 Player Similarity Matrix H for CAR Prior: Players with smaller Euclidean distances between their position summaries are treated as having more similar team roles and expected parameters.
- A.2 Player Similarity Matrix H for CAR Prior: The similarity matrix connects each player to the eight closest players, then symmetrizes the connections for the CAR prior.The cutoff of eight neighbors is arbitrary.
- A.2 Player Similarity Matrix H for CAR Prior: The resulting player-specific macrotransition parameters are indexed by player and transition type within the hierarchical modeling framework.
A.3 Basis Functions for Spatial Effects
The model represents spatial effects with piecewise-linear basis functions induced by a triangular court mesh, then uses preprocessing and factorization to construct player- and transition-specific functions. This representation reduces spatial fields to finite-dimensional coefficients and supports computationally tractable inference.
- Macrotransition effects: For each macrotransition type, player-specific spatial effects are represented as linear combinations of shared basis functions φ_ji(z), whose coefficients are estimated during preprocessing.The basis functions φ_ji are treated as fixed and known during later modeling and inference.
- Basis construction: Each spatial effect is expanded in shared basis functions ψ_k induced by a triangular mesh with d0 = 383 court vertices.The basis functions are zero at all other mesh vertices, equal one at their indexed vertex, and linearly interpolated within each triangle.
- Basis construction: Because the basis interpolates linearly within each mesh triangle, the resulting spatial effects are piecewise linear over the court.
- Computational representation: The same mesh basis supports Gaussian-process approximation through a finite-dimensional representation with sparse precision, providing computational savings through a Gaussian Markov random field.
- Model fitting: Microtransition models and simplified macrotransition models are fit independently for players using R-INLA, with Poisson regression used for the latter’s parameter estimation.The macrotransition fitting procedure scales across L = 461 processors, with each fit requiring at most 32GB RAM and no more than 16 hours.
- Macrotransition effects: Nonnegative matrix factorization extracts low-dimensional spatial structure from simplified macrotransition coefficient estimates across players.The resulting factor loadings serve as coefficients for the functional basis representation and capture structured variation across court locations.
A.4 Calculating EPVA: Baseline EPV for League-Average Player
The baseline EPV replaces an individual player’s transition behavior with a league-average profile while preserving a consistent interpretation of transition probabilities across player roles. The resulting value is evaluated from the modified transition model and depends on the current coarsened state.
- League-average transition profile: Player-specific transition distributions are aligned by removing passes to opposing-team players and reordering recipient columns by positional role.This makes transition probabilities across players as consistent as possible despite player identities being encoded in the coarsened states.
- League-average transition profile: The baseline transition profile is constructed by averaging the aligned player-specific transition matrices across the L players.
- League-average transition profile: Baseline EPV is defined for a hypothetical league-average player by replacing the focal player’s transition probabilities with league-average probabilities.The league-average profile is formed by averaging transition probabilities across all players.
- EPV calculation: The league-average baseline EPV at time t is obtained by evaluating the value under the modified transition matrix that substitutes the league-average profile for player ℓ.
- EPV calculation: The baseline EPV depends only on the current coarsened state C_t rather than the full possession history F(Z)_t.When averaged over a season, this coarsened-state baseline has the same expected results as the corresponding full-resolution baseline for the hypothetical league-average player.
APPENDIX B. DATA AND CODE
The paper provides a Git repository containing sample optical tracking data, visualization and EPV code, precomputed computational outputs, and a reproducible tutorial.
- Reproducibility: The repository includes one game of optical tracking data in CSV format and R code for visualizing model results and reproducing EPV calculations.
- Reproducibility: Precomputed Rdata files allow computationally intensive steps to be loaded without repeating the full calculations.
- Reproducibility: A reproducible knitr tutorial introduces the data and demonstrates core code functionality.