Source-linked AI summary
PhaseShift: Topology-Aware Data Harmonization and Model Consolidation Across Signalized Intersections
Yash Ranjan, Artur Kumik, Rahul Sengupta, Anand Rangarajan, Sanjay Ranka
TL;DR
Site-specific traffic models limit evidence reuse across heterogeneous intersections. PhaseShift harmonizes roadside trajectories into a shared actor-centric representation and trains a reusable backbone for pooled, zero-shot, and low-data adapted deployment. In the five-site evaluation, pooled training improved both 10-s displacement metrics over local training at every site, while the protocol remained conditional generation under replayed context rather than closed-loop simulation.
Problem
Traffic-behavior models are commonly trained separately for each intersection, limiting principled reuse across heterogeneous physical, control, demand, and sensing regimes.
Method
PhaseShift harmonizes trajectories, geometry, signal state, and actor context into a shared actor-centric space, then pretrains one actor-conditional backbone for pooled use, zero-shot transfer, or fine-tuning.
Results
At 10 s, POOLED improves both displacement metrics over LOCAL at all five sites, with median site-level reductions of 36.8% for minADE and 22.0% for minFDE.
Takeaways & Limitations
Within this five-site evaluation, one shared checkpoint supports consolidation across heterogeneous intersections, while residual performance identifies sites where adaptation remains useful.
Takeaways & Limitations
The evaluation uses one checkpoint per learned configuration, limited temporal evidence, and a replay-conditioned best-of-six protocol rather than closed-loop simulation.
Abstract
from arXiv · showhide
Learned traffic-behavior models are commonly trained separately for each intersection, creating model portfolios that cannot share evidence across sites. We present PhaseShift, a topology-aware framework that harmonizes heterogeneous roadside trajectories into a shared actor-centric representation and trains one reusable backbone. Ego-relative coordinates, trajectory-induced movement paths, normalized signal context, and variable-cardinality interaction tokens remove site conventions while preserving behaviorally relevant topology. The backbone supports pooled operation, zero-shot at a held-out intersection, and low-data adaptation. We evaluate five intersections in two Florida regions on balanced field data, 100k training windows and equal-sized test sets per site under a replay-conditioned, best-of-sampled-trajectory protocol. At 10s, one pooled model lowers both minADE and minFDE relative to trained local models at all five sites, with median reductions of 36.8% and 22.0%. Leave-one-intersection-out deployment, including one cross-region fold, beats local training on both 10-s metrics at four of five sites, although short-horizon performance is less uniform. Fine-tuning with 1,000 target update windows improves on zero-shot at three sites and is the strongest regime at one. At site 7, every cross-site mixture sharply lowers long-horizon error under a fixed 100k-window budget; test-likelihood gains argue against a best-of-sample dispersion-only explanation. Local models fall behind calibrated IDM at the two highest-flow sites after long autoregressive rollouts; pretrained-backbone regimes do not. Within this five-site evaluation, PhaseShift demonstrates consolidation across heterogeneous physical control settings while identifying sites that still require adaptation. The protocol measures conditional single-vehicle generation under replayed context, not closed-loop traffic simulation.
I. INTRODUCTION
PhaseShift addresses the cost and limited reuse of site-specific traffic models by harmonizing heterogeneous intersection observations into a shared actor-centric representation. It trains one reusable backbone for pooled operation, held-out deployment, and selective adaptation while preserving local physical and control context.
- Motivation: Up to 35 models can be required when five intersections across seven day-of-week domains are treated independently.Separate training, calibration, validation, and maintenance prevent observations from one domain from strengthening another.
- Motivation: Cross-site learning must remove coordinate and controller conventions while retaining topology, signal control, demand, interactions, and sensing context.These factors jointly generate observed behavior at signalized intersections.
- Contributions: PhaseShift maps roadside trajectories, neighboring actors, pedestrians, movement geometry, stop lines, and signal states into a shared actor-centric token space.The representation is designed for conditional trajectory generation without erasing local physical-control structure.
- Contributions: One shared backbone supports pooled operation, leave-one-intersection-out zero-shot deployment, and low-data adaptation without a learned site identifier.Residuals can indicate where specialization or investigation is warranted.
- Evaluation: The evaluation covers balanced field trajectories from five intersections in two Florida regions, comparing pooled, local, held-out, adapted, and fixed-window mixture regimes.The composition sweep controls nominal window count at one site as an initial check on the diversity–volume confound.
III. MULTI-INTERSECTION PRETRAINING FORMULATION
PhaseShift defines domains as joint physical-control regimes and applies a non-learned harmonization operator before shared actor-conditional pretraining. The resulting parameters support pooled use, zero-shot transfer with target context, or fine-tuning for residual target behavior.
- Domain formulation: A domain may represent an intersection or an intersection under a recurring operating condition, with each domain treated as a joint physical-control regime.Raw scenes differ in coordinates, movement paths, signal identifiers, and visible actor sets.
- Domain formulation: The non-learned harmonization operator hm converts each scene into a common actor-centric context using local geometry and normalized signal state.This establishes a shared input space before behavioral transfer.
- Pretraining: Shared parameters are pretrained over harmonized domains, with wm controlling domain sampling and fθ denoting the shared actor-conditional model.No learned site identifier is required.
- Deployment: For an unseen target, its geometry and signal context are supplied to the pretrained model for zero-shot use, or the model is initialized from θ⋆ and fine-tuned with target data.Zero-shot transfer therefore uses target context even without target behavior-model parameter optimization.
- Conditional generation: The transition model conditions on a target actor’s harmonized history, valid neighboring observations, movement geometry, stop lines, and signal state.Demand is represented through evolving actor and movement occupancy rather than a learned site label.
- Conditional generation: Generated target states are fed back autoregressively while neighboring actors, pedestrians, and signal state are replayed from the logged scene.The formulation separates harmonization, shared behavior learning, and optional residual adaptation.
IV. CROSS-DOMAIN DATA HARMONIZATION
Cross-domain harmonization expresses roadside observations in a shared token space while preserving geometry, control state, and actor context. It uses trajectory-induced movement geometry and variable-cardinality observations without claiming robustness beyond the reported field evaluation.
- Representation: The harmonizer removes global coordinate and orientation conventions while preserving local geometry and structured behavioral variation.It also avoids fixed actor ordering and learned site identifiers.
- Representation: Ego-relative tokens encode each entity’s pose and type-specific dynamics or extent, making globally translated or rotated scenes equivalent.This invariance does not extend to local geometry itself.
- Representation: Token descriptors separately encode neighbor actors, lane geometry, stop lines, and normalized signal state.Signal state is attached to the governed geometric element rather than represented by controller-specific identifiers.
- Movement geometry: Trajectory-induced movement polylines replace authored lane-connectivity graphs and are built from sufficiently complete tracks using movement grouping and arc-length normalization.Representative paths are aligned and discretized, while under-supported movement classes are filtered out.
- Data separation: The geometry-building interval is disjoint from behavioral training, validation, and test data, while behavioral inference can consume arbitrary partial tracks.The study does not quantify the per-site deployment cost of constructing geometry.
- Observation handling: Variable-cardinality encoding removes tokens for currently invalid observations instead of imputing a dense fixed-size actor tensor.This accommodates changing roadside tracks without asserting robustness beyond the field evaluation.
V. SHARED BACKBONE AND DEPLOYMENT LIFECYCLE
PhaseShift combines per-timestep spatial encoding with autoregressive temporal decoding inside a reusable multi-intersection lifecycle. The shared backbone is trained once across harmonized domains and then pooled, deployed zero-shot, or fine-tuned as needed.
- Backbone: PhaseShift encodes each harmonized scene spatially at every timestep, then decodes a multimodal next-state distribution autoregressively.The spatial encoder aggregates a variable-cardinality scene into an ego embedding.
- Stage 1 — Spatial encoding: Stage 1 uses weights shared across the sequence to produce timestep-level scene representations.These representations feed the temporal decoder.
- Stage 1 — Spatial encoding: Each modality is projected into a common dmodel = 256 space, with shared actor-state and map-speed tokenizers and type-specific relative-pose encoders.The tokenization blocks use Linear, LeakyReLU, and Dropout components.
- Stage 1 — Spatial encoding: Additive token composition combines modality encodings and adds a signal term for signal tokens.The factorization separates agent behavior from relative position.
- Deployment lifecycle: The lifecycle reuses the shared backbone as a pooled model, deploys it zero-shot at a new intersection, or fine-tunes it when residual error indicates specialization may be needed.Residual error is a prioritization signal, not proof of a specific behavioral cause.
- Stage 1 — Spatial encoding: Neighbor attention queries with the ego embedding and attends over neighbor, lane-polyline, and signal tokens to form one scene-aware ego embedding per timestep.The output has shape [batch, seq_len, dmodel].
B. Stage 2 — Temporal decoding
Stage 2 decodes future motion autoregressively by attending to scene-aware history and producing multimodal displacement distributions.
- The autoregressive decoder queries the latest embedding against history embeddings at each future step.
- Relative position encoding adds pose displacement and elapsed time to temporal attention without exposing a global coordinate frame.
- During training, teacher forcing uses ground-truth state history; during inference, each sampled state is reencoded and appended to the sliding window.
- A K = 25 Gaussian mixture head predicts mode weights, means, and diagonal log-scales to represent multimodal uncertainty.
- Masked negative log-likelihood trains on valid future targets while allowing partial tracks to contribute without discarding surrounding scene context.
VI. DATA AND MULTI-INTERSECTION PROTOCOL
PhaseShift evaluates five heterogeneous signalized intersections using balanced, chronologically separated field data and site-level or held-out-domain protocols.
- Five intersections span four Gainesville sites and one high-volume South Florida arterial intersection used as the cross-region target.
- Each intersection is treated as a pretraining domain, with roadside tracks generated from 10-Hz fisheye video and mapped into a rectangular ground frame.
- The domains vary in recorded duration, demand, movement count, and signal operation, while Gainesville sites share regional driving norms absent at site 22.
- The aerial layouts differ in footprint, approach width, lane arrangement, turn movements, and crosswalk placement, so the views are illustrative rather than metrically aligned.
- Each site contributes 100,000 training samples, and chronological partitioning prevents tracks or overlapping windows from crossing evaluation boundaries.
- POOLED trains on all five domains, whereas ZERO-SHOT leaves one intersection out and applies the resulting model without target-site parameter optimization.
- Checkpoint selection uses validation NLL without the target test interval, although zero-shot checkpoint selection still uses target validation data.
VII. PRETRAINING AND DEPLOYMENT REGIMES
PhaseShift compares local, pooled, zero-shot, and fine-tuned shared architectures under conditional single-vehicle rollout evaluation, emphasizing long-horizon consolidation and held-out deployment.
- PhaseShift seeks one common backbone plus specialization only for domains whose validation results support it, rather than one parameter set per intersection.
- LOCAL trains from scratch, POOLED trains across known domains, ZERO-SHOT excludes the target during pretraining, and FINE-TUNED updates zero-shot initialization with 1,000 target windows.
- The protocol evaluates conditional single-vehicle rollouts with replayed neighbors, pedestrians, and signal state rather than fully interactive traffic simulation.
- Figure 3 reports shared-backbone error reductions relative to LOCAL, where positive values indicate lower shared-model error and negative values indicate a specialization gap.
- At 10 s, POOLED lowers minADE and minFDE relative to LOCAL at all five sites, with median reductions of 36.8% and 22.0%.
- At 10 s, ZERO-SHOT beats LOCAL on both metrics at four of five sites, while one cross-region fold reduces errors by 38.2%/27.7%.
- At 2 s, ZERO-SHOT trails LOCAL everywhere except site 8, with median changes of −36.0%/−41.9%.
D. Target fine-tuning across intersections
Fine-tuning uses 1,000 target windows to adapt zero-shot checkpoints, improving long-horizon performance at several sites but remaining horizon- and site-dependent. A fixed-volume site-7 sweep shows that cross-site composition, rather than simply more windows, drives large gains, although temporal coverage and sampling effects remain unresolved.
- Target fine-tuning: 1,000 target windows improve both zero-shot metrics at sites 8, 9, and 22, leave site 7 essentially unchanged, and degrade site 10 at 10 s.Site 10 minADE changes from 0.8805 to 1.1256.
- Target fine-tuning: 38.2%/30.8% are the median 10-s minADE/minFDE reductions of FINE-TUNED relative to LOCAL across the five sites.The site-10 minFDE comparison is the exception, with a −11.1% change.
- Fixed-volume composition sweep: At fixed site-7 training volume, replacing 20% of local windows with equal shares from the other four sites cuts 10-s error by 66.2%/58.4%.The sweep uses 100,000 total windows, with 100/0/0/0/0 representing LOCAL.
- Fixed-volume composition sweep: The 80/5/5/5/5 mixture is best at 10 s and comes within 2.8%/4.0% of POOLED despite using one fifth as many total windows.The decisive contrast is between pure-local and mixed-site training; more non-local data is not uniformly better.
- Fixed-volume composition sweep: The composition sweep fixes window count but not independent temporal evidence, and best-of-N scoring may reward wider output dispersion.At 10 s, improved test NLL argues against a dispersion-only explanation, while additional controls are still needed before generalization.
F. Classical baselines and long-horizon behavior
Learned models outperform classical baselines at short horizons, but local models can lose their advantage during long autoregressive rollouts at the two highest-flow sites. Cross-domain pretrained regimes avoid that observed reversal, while likelihood results support genuine long-horizon gains and site-specific specialization remains evident.
- Classical baselines: At 2 and 5 s, every learned regime has substantially lower best-of-sample error than CONST.-VEL. and IDM at every site.The comparison is subject to the unequal hypothesis budget noted in the evaluation.
- Long-horizon behavior: At 10 s, LOCAL falls behind both classical baselines at site 7 and behind IDM at site 22, the two highest-flow sites.At site 7, LOCAL records 2.1595/5.1121 versus IDM’s 1.7216/4.0604.
- Likelihood: At 10 s, a pretrained-backbone regime has the best NLL at every site, including −5.94 versus −4.02 for LOCAL at site 7.The result argues against long-horizon displacement gains being solely an artifact of wider output dispersion under best-of-N scoring.
- Per-site synthesis: 67.1%/60.0% are site 7’s POOLED reductions in 10-s minADE/minFDE, while every cross-domain regime has much lower 10-s error than LOCAL.Site 7 is the busiest Gainesville domain and shows the sharpest long-horizon failure.
- Per-site synthesis: Site 22’s local 10-s error of 1.80/4.56 improves to 1.11/3.29 with ZERO-SHOT and 0.96/2.69 with POOLED in the cross-region case.FINE-TUNED and POOLED split the best 10-s metrics, and both decisively outperform local training.
- Per-site synthesis: Site 9 is the clearest specialization signal: ZERO-SHOT degrades both 10-s metrics, POOLED concedes short-horizon accuracy, yet POOLED remains best at 10 s.The site has atypical signal timing and the lowest flow among the five intersections.
- Per-site synthesis: At site 10, POOLED wins at 2 and 10 s, whereas 10-s fine-tuning increases minADE from 0.88 to 1.13 versus ZERO-SHOT.Across sites, the backbone supports consolidation while site 9 still signals horizon-specific specialization.
IX. DISCUSSION
PhaseShift supports consolidation across five heterogeneous intersections while retaining selective adaptation, but the evidence remains bounded by the replay-conditioned evaluation and unresolved attribution and selection questions.
- Results: At site 7, every fixed-volume cross-site mixture sharply outperforms pure-local training, indicating that mixture composition matters beyond nominal window count.The sweep fixes window count but not independent temporal evidence, and wider dispersion remains a possible influence under best-of-N scoring.
- Operational consolidation: At 10 s, POOLED is the best learned regime at sites 7, 9, and 10, while FINE-TUNED wins both metrics at site 8.POOLED is within 5% of the per-metric best at site 22, but this assignment was chosen from test results rather than a validated deployment policy.
- Scope and limitations: The evaluation does not execute the proposed validation-driven policy for consolidating known domains and specializing only when residuals warrant it.Operational selection must use validation under one fixed objective rather than test results.
- Physical-AI perspective: PhaseShift frames harmonization as an interface between learned models and infrastructure because layout, control, demand, behavior, and sensing jointly shape observations.The study bridges five realized physical-control regimes but does not separate their contributions or establish robustness to any one factor.
- Limitations: The evidence is limited by one checkpoint per learned configuration, restricted temporal coverage, no harmonization ablation, no contemporary learned baseline, and replayed exogenous context.Best-of-six metrics favor probabilistic models over deterministic baselines, and fine-tuning lacks a matched 1,000-window scratch control.
- Results: One pooled checkpoint improves both 10-s displacement metrics over local training at all five sites, while leave-one-site-out reuse improves both metrics at four sites.Leave-one-site-out reuse usually sacrifices 2-s accuracy, making the benefit horizon-dependent.
APPENDIX A ARTIFACTS
The appendix documents data preparation, shared training settings, fine-tuning, and non-learned longitudinal baselines used in the comparisons.
- Data preparation: Data is split 80/10/10 by collected-data time, preprocessed separately, and filtered to reduce leakage from trajectories crossing splits.Noisy and incomplete trajectories are removed, and bounding-box endpoints mark movement classes.
- Training and fine-tuning: The same architecture and hyperparameters are used across all five sites, while fine-tuning samples 1,000 windows from each 100,000-window training split.Fine-tuning updates all model weights; freezing different components is not analyzed.
- Least-squares and IDM baselines: The least-squares baseline fits a first-order polynomial to historical ego speed and extrapolates the fitted slope as constant acceleration.Both non-learned predictors are rolled out for the same H steps from the ego’s last observed ground-truth state.
- Least-squares and IDM baselines: IDM sets acceleration from the bumper gap and approach rate, using the nearer forward-cone vehicle or a virtual stop-line leader when the signal is not green.Parameters are calibrated per intersection on the training split using Bayesian posterior means.
- Baseline implementation: Both baselines project each predicted position onto the ego movement polyline and use the lane tangent as the new heading.This applies the same map constraint to both longitudinal predictors.