Source-linked AI summary
Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence
Tao Yang, Xuhui Lin, Kunyao Li, Haijiang Li
TL;DR
Existing accounts often treat space as either a common geometric container or disconnected modality-specific representations, leaving heterogeneous spatial processes difficult to unify. The paper defines space as an interventional invariant and develops cross-modal predictive geometry with causal conditions, group actions, and sheaf-valued urban representations. Under joint separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, while experiments test predictive, geometric, and urban-structure claims.
Problem
Existing spatial accounts struggle to express a common structure across heterogeneous sensory and urban processes whose modalities need not share a metric or representation.
Method
The paper models space through local states, modality-specific observations, interventions, conditional future laws, predictive-state quotients, and sheaf-valued representations for stratified urban systems.
Results
Under joint point separation, equivariance, and interventional faithfulness, the latent space is recovered up to the centraliser of the intervention group, with experiments evaluating holonomy and context saturation.
Takeaways & Limitations
Space is unified across mathematical, physical, perceptual, and urban domains by stable, mutually constraining changes under action rather than by erasing their distinctions.
Takeaways & Limitations
Identifiability requires policy excitation, jointly separating modalities, equivariant encoding, and interventional faithfulness; unobserved confounding can reduce the predictive state to a statistical object.
Abstract
from arXiv · showhide
Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of disconnected representations. Such approaches struggle to explain how heterogeneous sensory and urban processes can jointly reveal a common spatial structure, particularly when different modalities do not share the same metric or representation. This paper addresses this gap by defining space as an interventional invariant: the minimal relational structure that preserves local compatibility and the conditional laws of future observations under admissible actions. We develop a cross-modal predictive geometry that integrates local state spaces, modality-specific observation maps, an action groupoid, and a canonical predictive-state quotient, with explicit causal conditions for identifying interventional rather than merely observational structure. The key theoretical result shows that, under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, thereby reducing representational ambiguity to residual coordinate freedom. The framework is further extended to stratified urban systems using sheaf-valued representations, allowing geometric, physical, mobility, social, and economic layers to coexist without being reduced to a single metric. Synthetic experiments under noise evaluate equivariance, predictive sufficiency, holonomy, restriction-map recovery, cross-scale consistency, and context saturation. The resulting framework provides a unified and falsifiable foundation for spatial cognition, urban science, embodied AI, and em-spaced intelligence.
1 Introduction
The paper reframes space as a relational and interventional structure that links heterogeneous modalities through jointly predictable transformations. It formalizes this idea through predictive geometry, identifiability theory, projective limits, sheaf-valued urban models, and testable experiments.
- Motivation: Heterogeneous modalities reveal a common world through transformations that remain jointly predictable under intervention, rather than through a shared signal format.The examples include movement changing visual, auditory, and proprioceptive observations, and road closures changing mobility, economic, and social reachability.
- Problem and proposal: Space is defined relationally and interventionally as the least structure preserving local compatibility and intervention-conditioned future laws, with dimension treated as derived complexity.The proposal excludes mere statistical association and requires locality, reachability, composable change, cross-modal covariance, and stable prediction.
- Contributions: The framework defines cross-modal predictive geometry using probes, local states, interventions, conditional future laws, and causal assumptions distinguishing interventional from observational structure.
- Contributions: Under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group rather than an arbitrary homeomorphism.With a simply transitive action, the centraliser is the group itself, making the space a torsor and coordinates residual gauge freedom.
- Contributions: The framework treats context-indexed predictive quotients as a projective system whose limit defines space, while sheaf-valued urban models combine physical and social geometries without one shared distance function.The projective limit supports experimentally testable context saturation, and the sheaf model captures layered urban structure.
- Validation: Type, limit, and numerical checks under noise, ablation, and refinement translate the construction into an experimentally testable architecture for embodied and em-spaced intelligence.
2 Three senses of space
The paper distinguishes mathematical, physical, and perceptual senses of space by the structures and transformations that make each meaningful. It connects them through an interventional invariant and predictive-state construction without collapsing their distinct equivalence relations.
- Mathematical space: Mathematical space is structure relative to a category and its structure-preserving maps, whether topology, metric, adjacency, sheaf compatibility, or inner-product relations.
- Physical space: Physical space is an empirically constrained model involving fields, laws, symmetries, measurements, and error, with coordinates treated as representational rather than causal or metrical invariants.
- Perceptual space: Perceptual space organizes possible sensorimotor consequences and supports flexible inference beyond immediate sensation through lawful changes under movement.
- Relations among senses: Mathematical, physical, and perceptual spaces are linked by modelling, measurement, and action while retaining distinct equivalence relations.
- Interventional invariant: An interventional invariant is an equivalence class of relational models whose observable conditional laws remain unchanged under representation changes but vary lawfully under interventions.
- Predictive-state connection: The paper extends causal-state ideas to an intervention groupoid acting on typed local states, using locality, cross-modal compatibility, and scale to constrain a common predictive quotient.
3 Cross-modal predictive geometry
The paper defines cross-modal space as a minimal action-conditioned predictive structure and formalizes it through typed local states, modality-specific observations, and an intervention groupoid. Under explicit identification assumptions, this structure supports latent-topology recovery, causal estimation from logged data, and quantitative resolution guarantees.
- 3 Cross-modal predictive geometry: The canonical predictive quotient groups histories that have identical future-observation laws across every policy and horizon.It is minimal among sufficient representations; under suitable Borel conditions, the quotient is a standard Borel space, while otherwise only its predictive σ-algebra is guaranteed.
- 3 Cross-modal predictive geometry: Interventional laws are identifiable from logged data under sequential ignorability and positivity, making the canonical predictive state estimable.If positivity fails for some arrows, identification is restricted to supported policies and yields a strictly coarser space.
- 3 Cross-modal predictive geometry: Joint point separation and equivariance embed heterogeneous observations into a common topological representation, but true observation-map results alone do not establish learned statistical recovery.The embedding argument relies on compactness and Hausdorff conditions, while learned equivariance must be separately enforced or verified.
- 3 Cross-modal predictive geometry: Under E1–E4, the latent space is identifiable up to the centraliser of the intervention group rather than an arbitrary homeomorphism.For simply transitive actions, the residual ambiguity is the choice of origin and frame in a G-torsor; quantitative prediction error ε yields spatial resolution τ_ε.
- 3 Cross-modal predictive geometry: The framework distinguishes directed action geometry from metric geometry and treats dimension as task-, scale-, error-, and model-relative.The induced action cost is only a directed quasi-metric unless mutual reachability, definiteness, and reversal symmetry are added.
4 A sheaf-valued geometry of the city
The urban model uses typed sheaf-valued local states and restriction maps to combine heterogeneous geometric, physical, mobility, social, and economic layers without imposing one metric. Its holonomy-sensitive descent machinery provides compatibility certificates, identifies restriction maps under noise, and models interventions that can alter both states and couplings across scales.
- 4.1 Stratified base and fibres: Typed stalks assign distinct geometric, physical, mobility, social, and economic state spaces over urban strata, while restriction maps connect local descriptions across incidences.Different strata may have different dimensions, and singular junctions are retained rather than smoothed away.
- 4.2 Descent energy, the sheaf Laplacian and holonomy: A sheaf Laplacian certifies global compatibility: zero descent energy is equivalent to membership in the kernel of the sheaf coboundary.The result follows from positive definiteness of the incidence-residual weighting.
- 4.2 Descent energy, the sheaf Laplacian and holonomy: Non-identity restriction maps encode holonomy, allowing a connected sheaf to have no nonzero global section even when an identity-restriction graph Laplacian remains insensitive.The kernel dimension equals the dimension of the subspace fixed by the holonomy group; a single nontrivial cycle can force it to zero.
- 4.2 Descent energy, the sheaf Laplacian and holonomy: Orthogonal Procrustes consistently recovers restriction maps under input measurement error, whereas unconstrained least squares converges to an attenuated, inconsistent map.Projection onto O(n) removes the multiplicative attenuation caused by errors in variables.
- 4.3 Coupled field and flow dynamics: The sheaf-valued evolution law separates compatibility diffusion, nonlinear transport and response, actuation, forcing, and uncertainty across heterogeneous urban layers.Social and economic layers need not obey physical mass conservation; their locality is encoded through stalks, restrictions, reachability, and intervention response.
- 4.3 Coupled field and flow dynamics: Interventions may change coupling architecture as well as current state, while cross-scale residuals quantify when acting and aggregating fail to commute.Large residuals identify resolutions where omitted heterogeneity matters, making transition maps testable rather than assuming one privileged urban scale.
5 Audit of the core formulae
The audit distinguishes identities, hypothesis-dependent claims, and empirically estimated residuals, while noisy synthetic checks test recovery, equivariance, holonomy, restriction maps, scale consistency, and quotient saturation. Correctly aligned actions improve prediction, structural constraints remove asymptotic bias, and saturation is certified only by the noiseless control.
- 5 Audit of the core formulae: The audit explicitly separates identities, hypothesis-dependent statements, and residuals that must be estimated, preventing notation from being mistaken for a theorem.Table 1 records this type and falsifiability audit for the principal formulae.
- 5.1 Reproducible synthetic study under noise: The synthetic study adds isotropic Gaussian noise to inputs and targets and reports means and standard deviations over twenty seeds.Six checks cover cross-modal prediction, equivariance, holonomy, restriction-map recovery, cross-scale residuals, and projective-quotient saturation.
- 5.1 Reproducible synthetic study under noise: Correctly aligned actions achieve fourteen-times-better prediction than action-blind or misaligned conditions and approach the irreducible noise floor σ2 = 2.5 × 10^-3.The action-blind and misaligned conditions are statistically indistinguishable, with Welch statistic −1.94 over twenty seeds.
- 5.1 Reproducible synthetic study under noise: Only the noiseless control supports quotient saturation: the residual falls below 10^-29 at c0 = {v, a} and then stops decreasing.Under noise, later reductions of roughly 1 × 10^-3 are attributed to variance reduction, so noisy saturation alone cannot certify sufficiency.
- 5.1 Reproducible synthetic study under noise: The O(2) constraint removes the unconstrained estimator’s errors-in-variables bias floor, so its benefit persists asymptotically rather than vanishing with sample size.Figure 4 attributes the effect to the structural projection rather than small-sample regularisation.
- 5.1 Reproducible synthetic study under noise: Holonomy reduces the sheaf kernel to zero while the graph Laplacian kernel remains insensitive, and restriction-map recovery has error linear in noise.These checks directly test the sheaf obstruction and Procrustes recovery claims.
- 5.1 Reproducible synthetic study under noise: Cross-scale commutation degrades near the coarse Nyquist wavenumber, identifying a resolution where omitted heterogeneity becomes important.The reported residual measures disagreement between acting-then-aggregating and aggregating-then-acting.
- 5.1 Reproducible synthetic study under noise: The experiments establish internal consistency only; city-scale sheaf validity, known policies, and correct robot latent-state learning require field and embodied experiments.The authors state that these claims must be tested with experiments designed to break the model.
6 From embodied to em-spaced intelligence
Em-spaced intelligence couples a mobile agent with a persistent spatial body through shared prediction, distinct embodiment constraints, and jointly selected interventions. The proposed experiments test this framework across multimodal prediction, urban scale transfer, dual control, context saturation, and holonomy.
- 6.1 Two meanings of ESI: Em-spaced intelligence models a mobile agent and persistent spatial body as a coupled pair sharing predictive state, observations, actions, interventions, and governance constraints.Both bodies can alter the future field of admissible interaction.
- 6.2 Learning and control objective: The training objective combines future multimodal prediction with penalties for action covariance, local-to-global compatibility, scale aggregation, counterfactual interventions, and unsafe acts.The loss is L = Lpred + λeqEeq + λdescEdesc + λscaleEscale + λcfLcf + λsafeCΓ.
- 6.1 Two meanings of ESI: The framework treats intelligence as the coupled law of agent and environment, rather than as a robot merely using a smart building as a sensor.Environmental actions can change illumination, ventilation, access, signage, signalling, and information disclosure, while the robot moves, inspects, and manipulates.
- 6.3 Experimental programme: The experimental programme evaluates held-out intervention prediction, cross-scale transfer, joint robot-environment control, context saturation, and structural holonomy discrimination.Comparisons include modality concatenation, action-blind prediction, graph baselines, robot-only control, environment-only automation, and joint control.
- 6.3 Experimental programme: Holonomy testing distinguishes sheaf models from multilayer networks by testing whether residuals track circuit holonomy under identity versus learned restriction maps.The comparison targets a structural obstruction rather than fitting capacity.
7 Euclidean, non-Euclidean and directional geometry
The framework allows Euclidean, non-Euclidean, graph, and sheaf descriptions to coexist on a stratified urban base. Directional layers require conservative, layer-specific priors because relevant analytic bounds do not themselves determine sensing or reconstruction systems.
- 7 Euclidean, non-Euclidean and directional geometry: Urban layers can use distinct effective geometries on one base, including Riemannian tensors, Finsler costs, graphs, sheaves, and hyperbolic embeddings.These descriptions address directional asymmetry, discontinuities, singular junctions, and hierarchical structure without requiring one universal metric.
- 7 Euclidean, non-Euclidean and directional geometry: Directional acoustic, electromagnetic, and radiative layers concentrate along tubes, making harmonic-analytic obstruction theory relevant but insufficient for a sensing geometry.Kakeya estimates provide inequalities and obstruction mechanisms, not a minimum beam count or reconstruction algorithm.
8 What, then, is space?
The paper defines space as a context-independent projective limit of predictive quotients, with local descriptions linked by compatible restriction maps. Finite context saturation is presented as an empirical, falsifiable condition rather than a theorem about cities.
- 8 What, then, is space?: Space is the minimal local-to-global relational object that makes intervention-conditioned modality transformations jointly representable and predictively sufficient.Spatiality requires localisation, composability, intervention testing, counterfactual stability, and compatibility across overlapping probes.
- 8.1 Context-indexed predictive quotients: An observational context selects probes, modalities, policies, and horizons, producing a context-indexed predictive quotient whose refinements map uniquely onto coarser quotients.The refinement maps form a projective system over componentwise context inclusion.
- 8.2 Space as a projective limit: The canonical representation S∞ is the minimal statistic sufficient for all contexts simultaneously and remains invariant under changes of observational context.Each finite Sc is an interventional invariant relative to its context, while S∞ is the context-independent limit.
- 8.2 Space as a projective limit: Under suitable injectivity, the projective limit is equivalent to a finite-stage quotient, so a sufficient context can provide a complete spatial description.Standard-Borel regularity additionally requires a countable cofinal chain with standard-Borel stages and Borel connecting maps.
- 8.3 Falsifiable saturation: Context saturation is falsifiable: a sufficient context requires Esat(c0, c′) = 0 within tolerance for every refinement, whereas a nonzero residual reveals new spatial information.In the synthetic system, the residual falls by 3.7 × 10^-1 from {v} to {v, a} and later drops below 10^-29 in the noiseless control.
- 8.4 Space and time: The construction separates spatial compatibility and transition from temporal ordering, rate, and irreversibility, which arise through composition coupled to clocks, causal cones, dissipation, or action cost.Space and time are therefore not obtained merely by adding dimensions.
9 Limitations and boundary conditions
The framework’s claims depend on identifiable interventions, jointly separating modalities, stable laws, feasible computation, and governance-aware interpretation. The supplied evidence is synthetic and does not establish field validation.
- Identifiability: Interventional identifiability requires joint separation, equivariance, interventional faithfulness, and policies that excite the relevant degrees of freedom.Without these conditions, representations may not recover the intended latent space under the stated theorem.
- Confounding: Unobserved confounding can turn the estimated predictive state into a statistical rather than interventional object, invalidating the associated corollary.Facility managers, schedules, or occupants may select interventions in response to unrecorded conditions.
- Nonstationarity: Urban nonstationarity can obsolete a fixed sheaf and make the context-indexed projective limit fail to exist as institutions and populations change.Restriction maps and strata may require online change-point tests.
- Normativity: Predictability does not establish legitimacy: governance constraints must be model inputs, and intervention control determines which spatial distinctions the model can represent.This makes the intervention regime a substantive normative boundary, not merely a data-collection detail.
- Computation and evidence: Exact inference over rich sheaves, long horizons, and joint interventions is generally intractable, so approximation changes the effective predictive quotient and must be reported.The evidence base is synthetic and supplies no field validation; publishable empirical claims require real multimodal data, preregistered interventions, ablations, and replication.
10 Conclusion
The paper unifies mathematical, physical, perceptual, and urban senses of space through an interventional invariant while preserving their distinctions. Stratified sheaf constructions make heterogeneous spatial layers jointly representable and yield falsifiable criteria for em-spaced intelligence.
- Space is unified by stable, mutually constraining changes under shared actions, not by similar signals or a single erased representation.
- Stratified bases, typed fibres, and sheaf descent place buildings, fields, mobility, institutions, and exchange within one model.The construction also frames agents and spatial bodies as jointly determining future sensing and action possibilities.
- The framework is falsifiable when modalities fail to align, local models fail to glue, scale maps fail to commute within tolerance, or simpler predictive states perform equally well.
A Notation
This notation section identifies the principal mathematical objects used to describe sheaf structure, cross-scale relations, context-indexed predictive quotients, and residual intervention-group ambiguity.
- Table 3 is the principal notation reference for the paper's formal objects.
- The notation includes sheaf coboundary and weighted Laplacian operators, cross-scale restriction maps, and context-relative predictive quotients.
- S∞ denotes the projective limit of context-indexed quotients, while ZHomeo(X)(Φ(G)) denotes the centraliser representing residual ambiguity.
B Acceptance criteria for an empirical paper
The acceptance criteria require empirical tests to establish intervention coverage, compare competing geometries, isolate structural components, assess calibration under shift, and disclose governance constraints.
- Empirical studies must specify observed and held-out actions and ensure policy support can distinguish the claimed spatial states.
- Studies should compare Euclidean coordinates, graph distance, learned latent distance, action cost, and sheaf geometry, reporting where each succeeds.
- Structural ablations should independently remove action conditioning, cross-modal alignment, restriction maps, cross-scale penalties, and environmental actuation.
- Evaluation should report proper scoring or conditional likelihood, uncertainty calibration, and performance under changes in buildings, populations, weather, and institutions.
- Papers must state intervention control, whose costs enter the objective, and which safety or access constraints remain inviolable.