Source-linked AI summary
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
TL;DR
Dense, chaotic Global South urban scenes remain underrepresented in world-model benchmarks, despite demanding prediction under density, occlusion, heterogeneity, and partial observability. FactorJEPA factorizes future prediction into visibility-aware layout, agent, and interaction channels, separating from competitors across predictive diagnostics, including 43.3× Mask-ratio slope separation at 1B with full training.
Problem
Dense Global South urban scenes are underrepresented in world-model benchmarks despite requiring representations that preserve layout, agent state, and interactions under uncertainty.
Method
FactorJEPA decomposes future prediction into visibility-aware layout, agent, and interaction channels instead of a monolithic latent.
Results
FactorJEPA separates from the strongest competitor across all four primary diagnostics at 1B with full 115k-clip training, reaching 43.3× Mask-ratio slope separation.
Takeaways & Limitations
Structured predictive channels improve world modeling in dense, heterogeneous, and partially observed urban scenes, while revealing a trade-off with linearly accessible motion.
Takeaways & Limitations
The evaluation emphasizes short-horizon prediction from prerecorded video and measures intervention sensitivity rather than causal identification or unrestricted counterfactual reasoning.
Abstract
from arXiv · showhide
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).
Crowded, and Chaotic Global South
DENSEWORLD defines a measurable, underrepresented regime of crowded and chaotic Global South urban scenes characterized by density, heterogeneity, occlusion, and rapid interaction. Its 1,000-hour, 22-city dataset and quantitative comparison with established benchmarks establish this regime as a distinct stress test for predictive world models.
- Regime definition: DENSEWORLD combines high agent density and occupancy, persistent occlusion, heterogeneous actors, soft spatial boundaries, and rapid interaction pressure.These conditions require memory, object permanence, relational inference, and representations that preserve layout, agent state, and interactions under uncertainty.
- Dataset: The dataset contains approximately 1,000 hours of drive-through, walk-through, and aerial footage collected across 22 Tier-1 and Tier-2 Indian cities.Long-form recordings are segmented into 6–13-second clips and cover commercial, residential, transit, heritage, and other urban scenes.
- Interaction structure: Mixed traffic places cars, buses, trucks, auto-rickshaws, two-wheelers, bicycles, carts, and pedestrians in shared right of way under weak lane discipline.The resulting fluid spatial support, unstable visibility, and dense local negotiation create strong lateral and longitudinal coupling.
- Measurement: Five axes quantify the regime: agent count density, agent occupancy, occlusion pressure, interaction pressure, and agent heterogeneity.Together, these metrics measure multi-agent load, visual congestion, visibility degradation, spatially proximate agent pairs, and actor-category diversity.
- Benchmark comparison: DENSEWORLD shows consistently higher density and occupancy than BDD100K and nuScenes across matched scene categories, with the largest gaps in market, commercial, and flyover/underpass scenes.Statistics use a shared dynamic-agent taxonomy and comparable scene categories, supporting DENSEWORLD as a quantitatively distinct benchmark regime rather than merely a geographic extension.
Can Fine-Tuning Close the DENSEWORLD · Gap?
The section tests whether parameter-efficient adaptation can recover DENSEWORLD’s predictive structure without reorganizing the JEPA predictor. LoRA, DoRA, and Auto-RGN are compared under matched data, optimization, masking, and evaluation protocols.
- Gap?: Three adaptation strategies—LoRA, DoRA, and Auto-RGN—are evaluated to test whether conventional fine-tuning can recover DENSEWORLD predictive structure.The predictor architecture is not reorganized.
- Gap?: All methods use identical raw clips, optimization steps, optimizer, masking policy, and evaluation protocol; only the parameter-update mechanism changes.This isolates adaptation capacity while holding the training and evaluation setup fixed.
- Gap?: The baselines retain the executed V-JEPA objective, with context and target masks, online and momentum encoders, and a predictor.The notation defines x, Mc, Mt, fθ, ¯f¯θ, and gϕ.
- Gap?: The data- and step-matched protocol isolates adaptation capacity without introducing a different prediction target or training signal.Stop-gradient is denoted by sg.
- Gap?: LoRA adapts weight matrices with a low-rank update whose rank is much smaller than the input and output dimensions.The passage specifies r ≪ min(din, dout).
- Gap?: DoRA separates weight magnitude from direction, enabling directional adaptation without coupling it to weight magnitude.Its formulation combines a magnitude term with a normalized direction and an update ΔW.
- Gap?: Auto-RGN scores transformer blocks using relative gradient norms and updates only the selected block set.Normalization by parameter magnitude prevents larger blocks from being favored solely because of scale.
- Gap?: LoRA and DoRA target the same projection families, while Auto-RGN receives a matched trainable-parameter budget.This keeps the adaptation comparison controlled across update mechanisms.
FactorJEPA: Explicitly Factorized Predictive · Channels
FactorJEPA replaces monolithic future prediction with semantically anchored layout, agent, and interaction channels designed to resist shortcut mixtures in dense urban scenes. Its construction combines visibility-aware targets, block-structured separation, surgical adaptation, and depth-wise diagnostics.
- Channels: FactorJEPA replaces the monolithic predictor with explicit layout, agent, and interaction channels, using visibility gating and block-structured regularization to limit cross-factor leakage.The design factorizes future embeddings into semantically anchored coordinates and factor-specific predictive subspaces.
- DINOv2-Based Segmentations Agent-Layout-Interactions: DINOv2-derived region masks, boxes, descriptors, and confidences are temporally associated into tracklets and converted into structural layout, agent, visibility, and interaction targets.These targets supervise training only; they are not reused as evaluation labels, and reliability audits and association rules are provided in the appendix.
- Structured Factor Coordinates: A soft visibility gate suppresses uncertain or occluded agents without removing them, while localized interaction regularization encourages sparse pairwise structure.FactorJEPA composes target embeddings from layout, agents, and interactions, applying the visibility gate only to entity terms.
- Block-Structured Matrix Factorization: Distinct pathways, factor supervision, and learned dictionaries make the decomposition architectural rather than post hoc while preserving complete future-token geometry.The full scorecard spans predictive, motion, semantic, and temporal diagnostics at both backbone scales.
- Semantic Anchoring and Channel Separation: Factor-specific heads and semantic anchoring identify layout, agent, and interaction blocks, but the method claims semantic block separation rather than coordinate-level identifiability.Cross-channel penalties suppress linear and nonlinear leakage without constraining within-channel variation or implying statistical or causal independence.
- Predictor Surgery: Predictor surgery retains the pretrained V-JEPA encoder, target encoder, target construction, and masking policy while replacing only the monolithic predictor.Training uses the factorized predictor and top-K encoder blocks, with a staged curriculum progressing from layout to agents to interactions.
- Depth-Wise Factor Realization: Distinct onset and growth profiles across encoder depth indicate that layout, agent, and interaction information becomes linearly accessible at different depths.This is a representation diagnostic, not a neuron-level or causal decomposition.
Experiments & Evaluation
Experiments evaluate FactorJEPA with four diagnostics spanning future-latent fidelity, intervention sensitivity, masking robustness, and accessible motion information. Across data regimes and backbone scales, FactorJEPA improves structured forecasting, while its motion-readout trade-off depends on training data.
- Evaluation metrics: Four diagnostics measure future-latent accuracy, intervention sensitivity, partial-observability robustness, and linearly accessible motion information.The metrics are Future-frame L1, Causal L1, Mask-ratio slope, and Motion cosine.
- Matched 10k evaluation: Under matched stratified 10k training, FactorJEPA separates from the strongest competitor on Future-frame L1 and Causal L1, while robustness separates only at 1B.Confidence-interval separations are 6.3×/4.8× for Future-frame L1, 2.3×/2.7× for Causal L1, and 1.9× for Mask-ratio slope at 1B versus 0.9× at 2B.
- Full-corpus evaluation: Full 115k-clip training makes FactorJEPA separate on all four primary diagnostics at 1B.Separation reaches 43.3× for Mask-ratio slope, 33.2× for Future-frame L1, 20.0× for Motion cosine, and 13.9× for Causal L1; these are paired confidence-interval units, not multiplicative gains.
- Backbone-scale replication: Cross-scale rankings are strongest for Causal L1 at ρ = 0.979 and remain substantial for Motion cosine at ρ = 0.952, Future-frame L1 at ρ = 0.938, and Mask-ratio slope at ρ = 0.895.Twelve of fifteen diagnostics retain broadly consistent rankings, supporting the 1B backbone as a lower-cost screening proxy for adaptation choices.
Conclusion
The conclusion presents DENSEWORLD as a 1,000-hour, 22-city benchmark for dense urban world modeling and FactorJEPA as a visibility-aware factorization of future prediction. Across 2B and 1B backbones, FactorJEPA improves future-latent error, intervention sensitivity, and robustness to missing evidence, with stable cross-scale rankings.
- Benchmark: DENSEWORLD provides 1,000 hours of benchmark data from 22 cities for world modeling under dense traffic, heterogeneity, occlusion, and partial observability.The benchmark targets urban conditions characterized by dense traffic and incomplete visual evidence.
- Method: FactorJEPA decomposes future prediction into visibility-aware layout, agent, and interaction channels.The factorization explicitly organizes prediction around visibility-aware structural components.
- Results: ρ = 0.895–0.978: Method rankings remain stable across 2B and 1B backbones while improving future-latent error, intervention sensitivity, and robustness to missing evidence.The reported gains correspond to lower future-latent error, stronger intervention sensitivity, and greater robustness to missing evidence.
Main-Paper Figures and Tables, Enlarged
This section reproduces the main-paper figures and ablation table at full width, linking them to FactorJEPA’s decomposition of future structure into three predictive channels. The channels are automatically constructed and recomposed in the JEPA embedding to discourage cross-factor shortcuts.
- Method: FactorJEPA decomposes the future into layout, visibility-gated agents, and sparse interactions, then recomposes these channels in the JEPA embedding.The three views are constructed automatically by a fixed detection-and-segmentation pipeline.
Limitations
The evidence leaves open questions about interaction grounding, factor semantics, evaluation scope, geographic transfer, and full-scale training. These limitations motivate richer relational supervision, broader validation, and stronger causal and long-horizon analyses.
- Interaction grounding: Interaction grounding remains the principal open challenge because scalable targets capture only a tractable subset of temporally extended urban relations.The targets combine frozen DINOv2 regions, temporal association, relative motion, visibility, proximity, and reliability weighting across 1,000 hours of video.
- Factor semantics: Factor targets may contain missed agents, fragmented tracks, uncertain boundaries, and reduced reliability under occlusion, blur, poor illumination, or camera motion.The fixed DINOv2-based pipeline enables dataset-scale supervision but does not provide exhaustive human annotation.
- Factor semantics: Channel separation does not establish coordinate-level identifiability, complete statistical independence, or a unique decomposition of the future latent.Alternative teachers, taxonomies, or factor ranks may yield different yet comparably predictive partitions.
- Evaluation scope: The four diagnostics have bounded interpretations: Future-frame L1 measures frozen target-encoder fidelity, Mask-ratio slope controlled evidence robustness, Motion cosine linearly accessible motion, and Causal L1 target-encoder change consistency.Visual plausibility should be interpreted alongside latent-space and oracle-controlled evaluation rather than as standalone evidence of predictive correctness.
- Geographic scope: DENSEWORLD covers 22 Indian cities but is one large-scale realization rather than an exhaustive characterization of Global South urban environments.The dataset spans three capture modes and variation in density, infrastructure, illumination, weather, road structure, and traffic composition.
- Scale and transfer: Full 115k-clip training is reported for the 1B backbone, while the 2B model uses a matched stratified 10k regime, leaving full-data 2B behavior open.Cross-scale rank correlations indicate stable relative method behavior but do not replace a full-data 2B experiment.
Appendix
The appendix provides the complete evidence and implementation record for DENSEWORLD and FactorJEPA, progressing from provenance and reproducibility to methodological, statistical, temporal, and decoder analyses. It distinguishes confirmatory evidence from diagnostic and robustness analyses while defining the operational interpretation of factor separation and intervention-based evaluation.
- Appendix organization: Appendices A–C document DENSEWORLD provenance, distinctiveness, factor targets, and executed training protocols.These appendices establish the study’s reproducibility record.
- Appendix organization: Appendices D–G test component contributions, statistical reliability, methodological claims, and temporal or downstream consequences of the prediction–motion trade-off.Appendix G specifically evaluates the temporal and downstream consequences of that trade-off.
- Appendix organization: Appendix H documents the latent-to-RGB decoder and separates forecast error from reconstruction limitations using quantitative and qualitative future predictions.The decoder analysis addresses how latent forecasts translate into RGB outputs.
- Evaluation principles: Comparisons generally use identical evaluation clips, paired resampling units, and independently defined audit signals, while factor separation is interpreted as predictive-channel separation rather than coordinate identifiability or complete independence.Intervention-based evaluation measures consistency with controlled edits.
Appendix index.
The appendices document DENSEWORLD’s construction and validation, FactorJEPA’s targets and training protocol, attribution and robustness analyses, and latent-to-RGB evaluation. Together, they provide dataset provenance, methodological controls, metric definitions, reproducibility details, and qualitative diagnostics.
- Appendix A: DENSEWORLD Construction, Splits, and Responsible Use: Appendix A covers DENSEWORLD construction, source processing, source-grouped partitioning, corpus composition, privacy protection, responsible use, and research artifacts.It includes drive-through, walk-through, and aerial capture; an 80/10/10 train–validation–test allocation; automated privacy redaction; and governed research access.
- Appendix B: Five-Axis Validation of the DENSEWORLD Regime: Appendix B validates the DENSEWORLD regime using five axes, matched cross-dataset comparisons, uncertainty estimates, and sensitivity analyses.The axes are agent count density, agent occupancy, occlusion pressure, interaction pressure, and agent heterogeneity; results are interpreted as evidence for an Indian urban realization rather than exhaustive Global-South characterization.
- Appendix C: Segmentation and Training: Appendix C specifies FactorJEPA’s frozen-teacher factor targets, architecture, optimization, baseline matching, attribution controls, evaluation provenance, and resource accounting.It defines layout, agent, visibility, and interaction targets; compares full fine-tuning, LoRA, DoRA, Auto-RGN, FactorJEPA-RAW, and FactorJEPA configurations; and records parameters, GPU-hours, memory, throughput, latency, and reproducibility data.
- Appendix E: Metrics, Clustered Inference, and Robustness: Appendix E defines the evaluation contract and predictive, semantic, temporal, clustered-inference, multiplicity, sensitivity, and complete numerical result procedures.Primary diagnostics include Future-frame MSE, intervention-consistency L1, Mask-ratio slope, and Motion cosine, with horizon and masking curves and cross-scale ranking analyses.
- Appendix H: Latent-to-RGB Decoding and Qualitative Results: Appendix H evaluates latent-to-RGB decoding through decoder configuration, transport architecture, training objectives, neutrality controls, oracle and forecast-gap decomposition, quantitative metrics, and qualitative diagnostics.Evaluation includes PSNR, SSIM, LPIPS, Flow EPE, Agent F1, dynamic-region metrics, density- and occlusion-stratified rendering, factor-removal influence maps, and failure modes such as missed agents and decoder hallucination.
Responsible Use
DENSEWORLD provides approximately 1,000 hours of video from 22 Indian cities for predictive modeling in dense, heterogeneous, and partially observed urban traffic. It contains no manually annotated semantic labels; targets are generated automatically with a fixed DINOv2-based pipeline.
- Approximately 1,000 hours of drive-through, walk-through, and aerial video were acquired across 22 Indian cities.
- The dataset targets predictive modeling under high agent density, heterogeneous traffic, persistent occlusion, soft spatial boundaries, and frequent local interaction.
- DENSEWORLD contains no manually annotated semantic labels, with layout, agent, visibility, and interaction targets produced automatically using a fixed DINOv2-based pipeline.
Acquisition Protocol and Capture Modes
DENSEWORLD was collected by two professional video-collection organizations under a shared specification covering diverse sites and conditions without scripting behavior. Three capture modes—drive-through, walk-through, and aerial—target complementary spatial and interaction contexts, with acquisition metadata retained.
- Acquisition specification: Two professional organizations acquired DENSEWORLD under a shared specification defining coverage targets across sites, scenes, times, weather, visibility, roads, and infrastructure.The protocol controlled where and how videos were acquired.
- Acquisition specification: The acquisition protocol controlled video collection without scripting traffic, pedestrian, or animal behavior.Behavior remained unscripted while collection conditions were specified.
- Capture modes: Three capture modes cover complementary contexts: drive-through emphasizes road-level ego-motion and near-field interactions, walk-through captures pedestrian-scale shared spaces, and aerial provides broader crowd and traffic structure.City, collection session, capture mode, and source timestamps are retained as acquisition metadata.
Source Processing, Clip Construction, and Provenance … Qualitative Comparisons and Failure Analysis
The paper constructs DENSEWORLD through shot-aware, provenance-preserving preprocessing, fixed privacy and teacher-target pipelines, and dependence-aware validation. FactorJEPA factorizes future prediction into structurally supervised pathways, while the reported analyses constrain attribution and interpret qualitative outputs as complementary, decoder-conditioned evidence rather than causal explanations.
- Source Processing, Clip Construction, and Provenance: FFmpeg decoding and PySceneDetect AdaptiveDetector produce timestamped, source-linked clips of approximately 6–13 seconds without crossing detected shot boundaries.Adaptive normalization reduces false boundaries from rapid ego-motion and camera shake.
- Partitioning and DINOv2 Target Provenance: An 80/10/10 city- and capture-mode-stratified split assigns indivisible source-video groups before extraction, while fixed DINOv2 targets support training and teacher-relative held-out diagnostics.Test-derived targets, thresholds, and metrics never influence gradient updates or model selection.
- Corpus Composition and Coverage: DENSEWORLD covers heterogeneous urban environments and agent categories, with automated taxonomy quantities treated as illustrative or DINOv2-estimated rather than manually annotated prevalence.Coverage includes varied illumination, weather, traffic composition, crowd density, road geometry, and pedestrian–vehicle separation.
- Privacy Processing, Responsible Use, and Research Artifacts: Faces and plates are detected, temporally associated, safety-margin expanded, and Gaussian-blurred, with failed or incomplete redaction excluded from downstream use.The corpus is intended for world modeling, representation learning, forecasting, and partial-observability robustness—not identification, persistent tracking, sensitive-attribute inference, or surveillance.
- Regime; Shared Measurement Protocol and Automated Taxonomy; Definition of the Five Regime Axes; Matched Cross-Dataset Comparison: Cross-dataset regime comparisons use a frozen, identical DINOv2-based measurement pipeline and scene-matched, dependence-aware analysis across agent density, occupancy, visibility, interaction, and heterogeneity.These quantities are teacher-relative dataset statistics, not manually annotated ground truth or causal effects of dataset membership.
- Comparative Results, Uncertainty, and Sensitivity; Robustness; Paired Clustered Inference and Multiplicity: Matched estimates use equal-weighted retained source groups, hierarchical clustered BCa intervals, prespecified sensitivity analyses, negative controls, and leave-one-city-out checks.Stable separation requires preserved effect direction, clustered intervals excluding zero, and survival of leave-one-city-out analysis; conclusions remain specific to the measurement system and matched conditions.
- Segmentations, and Training; Frozen Teacher and Factor-Target Construction; FactorJEPA Architecture and Prediction Protocol: FactorJEPA replaces monolithic prediction with layout, agent, interaction, visibility, and interaction-strength branches whose separate computational paths receive factor-specific supervision and are recombined through synthesis dictionaries.DINOv2-derived targets encode layout, agents, visibility, and tracklet-derived interactions, with reliability weighting down-weighting uncertain targets rather than treating them as negative labels.
- Baseline Matching and Attribution Controls; Consolidated Attribution; Metric Sensitivity and Complete Numerical Results; Qualitative Comparisons and Failure Analysis: Attribution identifies the complete factorized structural package and structured teacher targets as supported effects, while qualitative maps and RGB decoding remain perturbation or decoder-conditioned diagnostics complementary to primary latent-space metrics.The executed comparisons do not isolate predictor architecture from separation and sparsity regularization, and no method is declared uniformly superior from isolated secondary wins.