Source-linked AI summary

Recovering Weighted Tangent Geometry from a Single-Scale Score Field

Ziqi Zhao, Qingjian Ni

arXiv:2608.22334v1stat.MLcs.LG

TL;DR

The paper studies whether a single-noise-level score field determines weighted tangent geometry when the branch center and homogeneity degree are unknown. It calibrates those nuisances from score values, integrates one shell’s tangential score into angular moments, and proves exact recovery results with finite-noise and certification boundaries. Experiments show that improved normalized-score fit can accompany worse angular-moment error.

  • Problem

    The paper asks whether local score queries at one noise level recover a branch point, homogeneity degree, and weighted tangent geometry when the center and degree are unknown.

  • Method

    Weak score identities calibrate the center and degree without score derivatives, after which one-shell tangential scores are integrated and converted into harmonic moments for spectral recovery.

  • Results

    Exact one-shell recovery identifies normalized angular measures in D≥2, while moments through degree 2K−1 recover at most K positive rays in arbitrary dimension.

  • Takeaways & Limitations

    Score fit alone is not a geometry-recovery certificate: controlled experiments find lower validation normalized-score error alongside higher angular-moment error and incomplete recovery.

  • Takeaways & Limitations

    The theory is local and assumes one known noise level, an approximately homogeneous junction window, a branch bound K, and stable calibration or declared certification bounds.

Abstract

from arXiv · show

Near a smooth data manifold, one tangent space summarizes local geometry. At a branch point, the corresponding first-order object is instead a measure over tangent directions, whose normalized masses record the local share of each branch under the chosen data measure. We ask whether a score field at one noise level determines this weighted tangent geometry when the branch center and homogeneity degree $d$ are unknown. In this tangent-measure model, $d$ is the local measure dimension. Gaussian smoothing of a homogeneous tangent measure satisfies an Ornstein--Uhlenbeck eigenfunction equation. Its weak form turns score values---without score derivatives---into a linear system for the center and homogeneity degree, with an explicit rank condition and perturbation bound. After this calibration, the tangential score on one sphere is the spherical log-gradient of a scalar Gaussian--cone transform. Integration recovers that transform up to scale, and all its spherical-harmonic multipliers are positive. Thus one exact shell identifies the normalized angular measure in every ambient dimension $D\geq2$. For at most $K$ positive rays, moments through degree $2K-1$ constructively recover count, directions, and weights in arbitrary dimension. Any fixed observation scheme needs at least $KD-1$ scalar tangential components. In the plane, degree $K$ is both sufficient and necessary, and we give quantitative finite-query certificates. For finite planar $C^{1,β}$ branches with positive $C^{0,β}$ densities, we prove $O(σ^β)$ convergence from the finite-noise score to its tangent model. In controlled experiments, 50k-step training lowers validation normalized-score error across four geometries yet raises angular-moment error, separating ordinary score fit from geometry recovery.

1 Introduction

The paper asks whether one-noise-level score queries can recover weighted branch geometry without a supplied center or homogeneity degree. It develops a score-only calibration and recovery pipeline, with arbitrary-dimensional identifiability, planar certificates, finite-noise rates, and diagnostics separating score fit from geometry recovery.

  • Problem: The inverse problem is to recover a branch point, homogeneity degree, and normalized angular tangent measure from local score queries at one known noise level.The angular measure records branch directions and relative local mass shares.
  • Motivation: A one-point score Hessian cannot identify branch geometry because it depends only on the first two angular moments, whereas shell fields distinguish branch counts away from the vertex.Uniform regular planar junctions with three or more rays share the same vertex Hessian.
  • Theory: Exact homogeneous tangent models are recoverable in every D ≥2 under explicit identifiability conditions, with planar quantitative certificates and arbitrary-dimensional finite-ray reconstruction.The planar theory also proves O(σ^β) tangent-score bias for finite C^1,β branches with positive C^0,β densities.
  • Method: Weak score averages calibrate the unknown center and homogeneity without differentiating the score, then tangential shell queries are integrated into a normalized scalar density for moment recovery.The resulting interface supports spectral estimators such as Prony, matrix-pencil, Toeplitz, and Vandermonde methods.
  • Diagnostics: Lower validation normalized-score error can coexist with higher angular-moment error and incomplete geometry recovery in controlled learned-score experiments.The error chain separates tangent bias, coverage, calibration, harmonic amplification, network, and query-aliasing effects.

2 Self-calibration and one-shell tangent geometry

The method first converts uncentered score values into center and homogeneity estimates, then uses one calibrated sphere to recover angular information. Positive harmonic multipliers enable injectivity, while finite-ray moment methods and query bounds provide constructive recovery and limits.

  • Target: The tangent model identifies normalized angular mass, not total mass; branch weights represent local shares under the reference measure defining µ.Changing the reference measure can change the weights.
  • Finite-noise approximation: Finite planar C^1,β branches with positive C^0,β densities converge uniformly to the tangent model at rate O(σ^β) on compact normalized query sets.The result is stated for the finite planar branch setting.
  • Self-calibration: Weak Ornstein–Uhlenbeck score identities yield a linear system for the unknown center and degree d without differentiating score values.Localized test functions turn the identity into score averages.
  • Self-calibration: Full column rank uniquely recovers the center and homogeneity degree, while perturbation bounds quantify errors and rank failure reflects translational symmetry.Ill-conditioned or symmetric cases can prevent stable calibration.
  • One-shell angular recovery: Integrating the shell tangential score recovers the positive scalar Gaussian–cone transform up to scale, and positive harmonic multipliers make the normalized angular measure injectively identifiable.The argument uses sphere connectivity and Funk–Hecke diagonalization.
  • Finite rays: Moments through degree 2K−1 constructively recover at most K positive rays in arbitrary dimension, using rank, joint eigenvalues, and linear coefficients.The subsequent multivariate Prony step recovers count, directions, and normalized weights.
  • Planar recovery: In the plane, modes |k| ≤K suffice while degree K−1 does not, establishing sharpness for the consecutive-moment interface.The lower bound follows from distinct rotated regular K-ray measures sharing the first K−1 nonconstant moments.
  • Query lower bound: Exact identification through a fixed continuous scalar-query scheme requires at least KD−1 tangential components, even when K is known and observations are noiseless.For D=2, this becomes M ≥2K−1; adaptive observations are not covered.

3 Shell information beyond a one-point Hessian

Vertex information can collapse across different regular ray junctions, while shell densities and tangential scores retain count-specific angular information. The shell therefore resolves distinctions that the one-point Hessian misses.

  • Shell versus vertex: Every uniform regular planar q-ray junction with q≥3 has identical first two angular moments and vertex Hessian, yet its first count-specific shell harmonic is nonzero.This explains why shell observations recover branch-count information absent from the vertex Hessian.
  • Shell versus vertex: Three- and four-ray regular planar measures share the vertex Hessian −I2/2 but have different R=2 shell densities and tangential scores.The figure directly contrasts one-point collision with shell conditioning.
  • Conditioning: Inverse harmonic multipliers grow rapidly with degree at small radii, making shell radius relevant to conditioning.The figure marks tested radii and degrees 1:4 used in planar K=4 experiments.

4 From finite queries to conditional guarantees

The finite-query estimator converts shell score samples into moments and applies staged, conditional certification for count, directions, and weights.

  • Finite-query estimator: M > 2K equally spaced shell queries are integrated spectrally, normalized, deconvolved by known multipliers, and converted into angular moments.Toeplitz rank, unit-circle roots, and a nonnegative solve recover count, directions, and weights.
  • Certification: A numerical candidate uses a relative eigengap, whereas certified output applies sufficient thresholds and abstains when a test fails.The two modes distinguish heuristic reconstruction from conditional guarantees.
  • Certification: ΓK > 2EK guarantees the branch count as the number of estimated Toeplitz eigenvalues above EK for a declared geometry class.The class-level resolution gap is necessary because an arbitrarily small extra branch cannot be excluded from error estimates alone.
  • Certification: Declared angular-separation and Vandermonde-singular-value bounds provide sufficient direction and weight tests; failure causes abstention rather than invalidating a candidate.The required bounds must be fixed before inspecting the tested field.
  • Perturbation margins: Count correctness can precede geometric accuracy because root separation or Vandermonde conditioning may limit later recovery after Toeplitz rank stabilizes.The perturbation margins compose rank, root, and weight stability results.

5 Experiments

Experiments evaluate calibration, shell inversion, finite-query certification, empirical KDE recovery, and learned-score behavior across controlled synthetic settings.

  • Higher-dimensional recovery: The three-dimensional shell experiment reconstructs non-coplanar two- and three-ray measures through degree-five spherical harmonics and a moment pencil.Table 1 defines full geometry using correct count, maximum angular error 0.1°, and maximum weight error 10^-3 for these rows.
  • Finite-query validation: 104,352 finite-query trials vary radius, query count, score noise, separation, weights, clustering, rank selection, and center offset.At R = 3 and M = 32, all exact evaluation cases recover; clustered rays are less stable than isolated close pairs.
  • Certification: 336 bounded-error trials produce positive reports that all respect their count, direction, or weight bounds.Later certification stages abstain before earlier ones as error grows.
  • Population/KDE validation: 6,528 population/KDE trials cover eight named and twelve continuously generated held-out geometries, with population inversion succeeding on all twenty.At coverages 256, 512, and 1024, blind recovery is 98.96%, 98.96%, and 99.74%.
  • Learned-score diagnostics: 50k-step training lowers validation normalized-score error across four geometries but raises angular-moment error.The controlled learned-score diagnostic separates ordinary score fit from downstream geometry recovery.

6 Related work

The paper extends score-based geometry analysis and spectral atom recovery to weighted junction inversion from a single-scale score field.

  • Score geometry: Prior smooth-support work uses scores and Jacobians for normal bundles, dimensions, pullback geometry, or projections, while the paper inverts weighted junction geometry.The proposed recovery targets the center, homogeneity, and tangent measure without a score Jacobian.
  • Spectral and spherical inversion: Finite-rate-of-innovation, MUSIC, ESPRIT, super-resolution, spherical deconvolution, and multivariate Prony provide atom or moment recovery tools.The paper adds score calibration, shell integration, all-frequency injectivity, and the finite-query score-to-moment error chain.
  • Neighboring targets: Global schedule mixture weights and point-cloud methods address different targets: local sector masses and auditing a trained score field remain distinct.The paper’s coverage is KDE sampling rather than the cited memorization mechanism.

7 Limitations and conclusion

The results establish single-scale recovery under explicit assumptions while delimiting the local, approximately homogeneous setting and unresolved finite-noise scope.

  • Limitations: The scope is one known noise level, one approximately homogeneous junction window, a branch bound K, and convergence in a Gaussian-weighted score ratio.The planar theorem covers zero-thickness C1,β branches, not thickness, general strata, or arbitrary-dimensional rates.
  • Limitations: Calibration requires full rank and can be ill-conditioned near translation symmetry, while certification requires external error, weight, and separation bounds.Without a stable tangent measure, the target may vary with scale.
  • Conclusion: The conclusion is exact arbitrary-dimensional identifiability with planar certificates, while broader rates and pretrained-score guarantees remain future work.Experiments expose a gap between score fit and geometry recovery.
  • Calibration: The strong score identity links divergence, score norm, center, dimension, and ambient dimension; its weak form avoids differentiating the score.Full column rank of the stacked system makes center and degree recovery unique.
  • Finite-noise limit: For finite planar C1,β branches with positive C0,β densities, the finite-noise score converges uniformly on compact normalized query sets at rate O(σ^β).C2 branches with C1 positive densities give an O(σ) bias, and the rate check is not used in the proof.
  • Finite-noise limit: Minimum angular separation affects downstream finite-atom conditioning but not the forward tangent approximation.The ratio constant also depends on positive tangent mass.

C Proof of one-shell injectivity

One exact shell identifies the normalized angular measure because score integration recovers a positive scalar transform and every harmonic multiplier is nonzero. Finite positive-ray measures are then reconstructed from finitely many moments.

  • Injectivity: The shell transform is positive, and its spherical-harmonic multipliers are strictly positive, so equality of shell fields forces equality of all angular moments.Density of finite spherical-harmonic combinations in continuous functions then identifies the measure.
  • Finite-ray reconstruction: For at most K positive rays, moments through degree 2K −1 recover the count, directions, and weights through a finite-rank multivariate moment pencil.Commuting multiplication matrices provide directions, followed by a linear solve for weights.
  • Information lower bound: A fixed observation scheme requires at least KD −1 scalar tangential observations to identify arbitrary K-atom measures in dimension D.The lower bound follows from the parameter dimension and invariance of domain.
  • Score-to-moment interface: Integrating the tangential score recovers the scalar shell transform up to a constant, which unit-mass normalization removes.The same score-to-moment interface applies after harmonic deconvolution by the positive multipliers.
  • Planar sharpness: In the plane, degree K moments suffice, while degree K−1 can fail to distinguish rotated regular K-ray measures.The next moment separates the otherwise matching measures.
  • Role of positivity: Positivity is essential: without it, cancellation can hide directions and Toeplitz rank need not equal the atom count.The finite-ray recovery argument relies on nonnegative quadratic forms and positive weights.

H.1 Proof of the stagewise stability proposition

The stability proof is stagewise: moment perturbations first certify count, then root locations certify directions, and finally conditioning certifies weights. Later failures produce abstention rather than invalidate earlier certified stages.

  • Count: The count test uses Weyl’s inequality to separate positive eigenvalues from perturbed null eigenvalues, without solving roots or weights.Its perturbation threshold is based on the Toeplitz rank margin.
  • Directions: Direction recovery requires an additional root-separation margin beyond the count condition.The sufficient angular bound depends on root separation and the annihilating polynomial.
  • Weights: Weight recovery adds a separate Vandermonde conditioning requirement after the direction bound is established.The full certificate is the intersection of rank, root, and weight conditions.
  • Abstention: A later failed inequality causes abstention at that stage, not evidence that an earlier recovered rank is wrong.Thus count may be certified even when full geometry is not.
  • Center error: A center-offset error is added before rank, root, and weight tests through a translated-query bound.The calibrated center error contributes directly to the shell observation error.

I Finite-data and learned-score errors

Finite-data recovery is controlled by tangent bias, local coverage, and score-side errors, while learned-score fit need not track angular geometry. The analysis distinguishes population, KDE, and learned fields through matched recovery stages.

  • KDE concentration: The empirical KDE analysis makes denominator conditions and effective local coverage explicit rather than hiding them in big-O notation.Concentration bounds control scalar weights and vector numerators simultaneously.
  • Error propagation: Score, integration, tangent-model, network, and calibration errors combine into a shell error that propagates to recovered moments.Replacing d by an estimate contributes an additional term proportional to its estimation error on compact intervals.
  • Detector limitations: A one-point Hessian loses posterior-covariance information at very fine noise, whereas the full score retains nearest-sample displacement.This separates a detector limitation from information available in the score field.
  • Tangent approximation: For finite planar C1,β branches with positive C0,β densities, the tangent-score bias is O(σβ) on fixed normalized query sets.The measure choice affects coarse-scale laws, with arc length producing a linear width-to-noise term.

J.4 Why these facts do not prove multiscale necessity

The experiments distinguish fixed-scale score information from detector and training limitations, rather than establishing a need for multiscale observations. Empirical recovery depends on coverage and perturbations, while learned-score fit can diverge from geometry recovery.

  • Score queries: A sufficiently fine raw-score query can recover a nearest sample and distinguish supports with positive separation, unlike the Hessian detector’s bandwidth window.The conclusion applies to absolute-error score queries and does not impose a fixed-scale lower bound.
  • Far field: Far-field density can become extremely small while score magnitude grows, so density decay does not imply a weak log-score.A deterministic check reaches log10 pσ = −435.40 while ∥sσ∥= 800.99.
  • Finite-query evaluation: Exact finite-query cases at R = 3 and M = 32 all recover, while noisy four-ray clusters are less stable than isolated close pairs.The noisy results are presented as an empirical conditioning map rather than a worst-case theorem.
  • Evaluation scope: Oracle-assisted conditional counts are reported separately from blind recovery, and oracle rank does not implement the later root and weight tests.This prevents conditional diagnostics from being read as an end-to-end guarantee.
  • Learned scores: At the standard training budget, three configurations with perfect empirical-KDE success still undercounted in 3/5 learned-score seeds.Longer controls restored count in several settings, but full weighted geometry remained outside tolerance in 80%–100% of seeds.

K.5 Cross-geometry evaluation

The fixed learned-score evaluation compares matched MLPs across four geometries and two training budgets, while the convergence study tests whether improved score fit improves geometry recovery. Longer training reduced normalized-score RMS but did not reliably resolve angular and weight errors.

  • Fixed evaluation: At 5,000 updates, both parameterizations recover count in at least 4/5 seeds but full geometry in at most 2/5 on regular Y, clustered four, and weak four.On weak Y, count is correct in 3/5 plain MLP seeds and 2/5 residual MLP seeds.
  • Fixed evaluation: Count recovery can coexist with angular or weight errors above the declared full-geometry tolerance.Clustered four reaches 5/5 count-correct seeds in both parameterizations while median angular or weight error exceeds tolerance; weak four shows a related weight-error bottleneck.
  • Fixed evaluation: The fixed evaluation uses four geometries, plain and residual parameter-matched MLPs, 1k versus 5k updates, and five seeds.The standard matrix uses N ∈ {512, 8192}, three noise scales, and 1,000 updates; the long matrix uses N = 8192, σ = 0.1, and 5,000 updates.
  • Train-to-convergence evaluation: The train-to-convergence protocol selects checkpoints using validation normalized-score RMS, excluding moment error and geometry recovery from selection.The selected checkpoint minimizes mean validation normalized-score RMS among saved checkpoints at or beyond 20,000 updates.
  • Train-to-convergence evaluation: From 5,000 updates to 50,000 updates, normalized-score RMS improves in every geometry while maximum moment error increases in every geometry.The tested larger residual control also has higher validation score error than the selected plain model on each of its three geometries.
  • Tolerance analysis: Full-recovery fractions remain sensitive to weight tolerance, so count, angle, and weight measurements are retained separately.Failures remain even at the most permissive displayed angle-weight tolerance pair.

L Scope of the results

The paper separates established theoretical results, synthetic observations, and claims ruled out under its stated score-oracle formulation. It also documents executable-code provenance and hardware-sensitive seeded optimization.

  • Proved: The paper proves weak single-scale recovery of center and homogeneity under a full-rank test matrix, exact one-shell injectivity, and finite-ray recovery from moments through degree 2K −1.The listed results also include arbitrary-dimensional recovery and the KD −1 fixed scalar-query lower bound.
  • Proved: The paper proves sharp degree-K planar recovery, finite-query moment, count, direction, and weight tests, and a finite-data local-coverage bound under stated assumptions.
  • Observed synthetically: Synthetic experiments observe weak calibration in D = 2, 3, 5, non-coplanar three-dimensional score-shell recovery, and held-out weighted-geometry recovery beyond fixed templates.They also report inverse-square-root coverage scaling, complete bounded-error certificate execution, and separation between count and full-geometry fidelity across four planar geometries and two MLP parameterizations.
  • Scope boundaries: Under the stated score-oracle formulation, the paper rules out claims that heterogeneous feature sizes alone imply multiscale necessity, fixed-scale failure is universal, or branch loss is irreversible.It also does not claim universality across diffusion architectures or real data.
  • Reproducibility: All reported numbers, tables, and figures are generated from executable code, while seeded neural optimization may vary across hardware.Released raw rows can be reaggregated without retraining.
Loading 2608.22334v1…