Source-linked AI summary
A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
Ihor Kendiukhov
TL;DR
Astronomical foundation models combine survey pixels with catalogue products that may be incomplete, raising the question of whether learned representations inherit those catalogue failures. The paper audits AION-1 with causal input interventions, finding that detection metadata can drive coherent tomographic redshift biases and that sparse dictionaries are unreliable causal handles.
Problem
Training on raw observations together with derived catalogue products may let models rely on catalogue channels without discounting contradictions, creating an accountability concern for coherent cosmological biases.
Method
The study audits AION-1 by editing its detection maps while holding pixels byte-identical, propagating measured survey misses, and testing internal directions and sparse dictionaries.
Results
Detection gating drives the key reported effect, while the measured miss rate shifts tomographic mean redshifts to a median 0.71 times the LSST DESC requirement and exceeds it in 12 of 40 assignments.
Takeaways & Limitations
AION-1’s vulnerability is a metadata preference whose coherent detection failures can become a cosmological systematic; withholding the detection channel removes it at no measurable cost.
Takeaways & Limitations
The propagated systematic is a bounded estimate under stated assumptions, while the positional-error scenario is reported only as an upper bound.
Abstract
from arXiv · showhide
Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic. We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports -- flux, size, ellipticity, redshift -- by 110-4400 times a matched placebo. The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all. The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times. Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale. Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.
1 Introduction
AION-1 makes catalogue-derived decisions directly editable, enabling causal tests of whether it uses survey metadata instead of raw pixels. The audit frames accountability around coherent biases, especially in tomographic mean redshifts.
- The audit targets accountability because coherent shifts in tomographic mean redshift matter more for cosmology than large random errors.
- AION-1 uniquely ingests the survey segmentation image, making pipeline decisions directly editable for causal analysis.
- Holding survey pixels byte-identical while editing the detection map tests whether the model prefers catalogue-derived metadata over raw observations.
- The paper investigates detection gating, contradicted catalogue photometry, downstream redshift cost, tokeniser limits, and sparse dictionaries as causal handles.
- The study uses causal interventions with matched placebos and bootstrap intervals rather than attention inspection.
2 Model, data and methods
The study combines multimodal survey data with controlled input and internal interventions to measure what AION-1 uses. Its sampling and readout choices are designed to preserve physical coordinates and isolate model-supplied information.
- Model: AION-1 uses a 314.3 M-parameter base transformer, with an 859.7 M-parameter large model for scaling tests.
- Data: The dataset cross-matches 115 404 galaxies with DESI spectra, Legacy Survey imaging and segmentation maps, and PROVABGS physical parameters.
- Sampling: Random sampling avoids healpix-order bias: a contiguous field produced redshift-probe R2 = 0.48, versus 0.91 for a random draw.
- Interventions: Input interventions alter one channel while keeping every other input byte-identical and compare the effect with a matched right-ascension placebo.
- Interventions: Internal steering adds a difference-in-means vector to encoder states to test whether that direction restores damage caused by a wrong mask.
- Readouts: Scalar readouts generally use the mean token posterior, while specified redshift-precision and later analyses use the median estimator.
3 The detection gate
Causal interventions show that AION-1 gates its readouts on the survey segmentation map, especially whether a source is present at the target position, rather than integrating the enclosed light. The effect persists across metadata channels and model scale, while real-blend partitioning does not propagate through this channel.
- Existence: 110–4400 times the placebo: altering only the segmentation map changes every reported quantity while image pixels remain byte-identical.The intervention affects flux, size, ellipticity, and redshift.
- Metadata and evidence: 0.312 → 0.003 σ: adding DESI spectra reduces the redshift intervention effect a hundredfold, while ellipticity remains 0.931 → 0.895.Independent evidence selectively suppresses the gate where counter-evidence is available.
- Mechanism: detection gating: r = 0.474 versus r = 0.298: prediction changes track central mask coverage more closely than enclosed light.Eroding the mask collapses estimates, whereas dilation barely changes them, indicating a presence gate rather than aperture behaviour.
- Presence, not partition: R = −0.012 [−0.014, −0.008] across 231 real blends, showing that pipeline deblending partitions do not determine the model readout.For 322 more severe systems, R = −0.0064 [−0.0075, −0.0046], with no trend across severity bins.
- Internal direction: 46 % versus 5 % recovery at α = 4: a mean-difference direction restores part of the flux collapse more than a matched random vector.However, recovery varies with intervention scale and the injected block does not locate where the gate is computed; at α = 8, the gate direction reaches 89 % while the random control reaches 66 %.
- Metadata and evidence: 0.62 versus 0.067 mag: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata.This supports a vulnerability associated with redundant catalogue channels alongside raw observations, not only with the detection map.
- Scaling: The larger model defers more: scaling does not mitigate the gate and appears to aggravate it.The paper reports this failure mode as growing with parameter count at fixed data.
4 Incidence and cosmological consequence
The Legacy Survey’s measured detection failures can bias tomographic mean redshifts, with positional errors producing the largest exceedances; withholding the detection mask removes the vulnerability without measurable cost.
- 3.68% of targets have no segment covering their position, matching the pipeline failure configuration that collapses the readout.
- 0.88× versus 2.7× the requirement results from empirical versus all-zero representations of the same miss rate.The all-zero representation breaches the requirement in two of four bins.
- 8.3× the requirement is reached in the worst bin when masks are displaced by observed-magnitude positional errors, putting three of four bins over.Spectra remove the effect.
- 1.24 is the ratio by which magnitude-dependent miss sampling raises the worst-bin median, but the difference remains inside realisation scatter.The worst-bin median changes from 0.71× to 0.88× the requirement.
- Withholding the detection mask removes the vulnerability, while flux g remains neutral within the reported intervals.A generic centred disc instead worsens flux g by +220%.
5 Tokeniser-limited behaviour
AION-1’s imaging and redshift representations are constrained by tokeniser behaviour: image codes collapse onto brightness, while redshift precision is limited by a coarse grid.
- 93.3% of source patches are carried by 40 image codes lying on a one-dimensional brightness manifold.R2(flux, PC1) = 0.999 identifies the latent scalar as brightness.
- ∼28 versus 934 effective states are resolved by the image and spectrum codecs on source patches, respectively.
- Morphology can only be represented through spatial arrangements across the token grid, not within an individual token.The paper attributes this limit to codebook collapse and says scaling the transformer cannot close the gap.
- ∆z = 0.005865 is the redshift token grid spacing, with 27% of codes unused.The paper identifies the tokeniser as the precision bottleneck.
6 Posterior calibration and the redshift mechanism
The paper finds that AION-1’s redshift posterior mean is unusable for photometry-only outputs, while spectra produce overly thin tails and line-based readouts show limited catastrophic failure.
- 0.236 versus 0.050 is the median |∆z| for the posterior mean versus median with photometry alone.The linear grid drags the mean upward when the posterior is broad.
- 0.513/0.662 coverage is obtained at nominal 0.50/0.68 for photometry-only posteriors.
- 75% coverage is achieved by the nominal 99% interval with spectra, indicating overly thin posterior tails.
- 32.5 is the rest-frame contrast for stacked spectrum-token responses, versus 4.4 observed-frame and 2.95 shuffled-redshift null contrasts.The peak lies at 6643 Å, identifying Hα within one bin.
- 5.7% of induced Hα-window failures land on line-ratio aliases, versus 43.4% for the random null.The pre-registered prediction of catastrophic, predictable failure is falsified.
7 Sparse dictionaries are unreliable causal handles
Sparse dictionaries carry causal signal, but their steering utility varies substantially across configurations and random seeds, and reconstruction quality does not identify effective causal handles.
- The experiment tests whether gate-selective dictionary entries outperform a matched-norm difference-in-means steering vector for restoring collapsed readouts.Fifteen dictionaries vary widths, sparsities, and seeds under identical causal tests and baselines.
- Recovery at α = 4 spans 26.0–74.7 % with median 52.2 % across 15 dictionaries.All 15 beat the matched-norm random control at 13.4 %; only two beat difference-in-means at 64.3 %.
- Thirteen of fifteen dictionaries lose to difference-in-means, while PCA reaches 68.1–73.4 % across k.Only the 32768-feature, k = 16 dictionary exceeds PCA, recovering 74.7 %.
- Seed variation alone moves recovery by up to 18.4 percentage points at fixed width and sparsity.The resulting spread exceeds the gap between the median dictionary and difference-in-means.
8 Discussion
AION-1’s detection vulnerability reflects a broader metadata preference: catalogue-derived inputs can override raw observations, while tokeniser limitations independently constrain imaging and redshift performance.
- AION-1 relies on predictive catalogue-derived metadata over raw signal and does not discount it when the two conflict.The resulting detection failures are coherent across objects, making them a cosmological systematic rather than added noise.
- The detection-miss representation changes the estimated downstream magnitude, with the intuitive representation overstating it fourfold.This bounds how input interventions should be used to price learned-model failure modes.
- Contradictory metadata produces confident wrong answers, whereas removing genuine spectral information appropriately widens the posterior.The effect grows with model scale and is not visible as posterior broadening.
- The image tokeniser has ∼28 effective brightness-ordered states, cannot represent morphology within a token, and does not improve with scale.The imaging–spectroscopy gap should therefore be attributed to the codec before the transformer for morphology applications.
- Sparse dictionaries are beaten by simpler linear methods on the causal task in thirteen of fifteen configurations and vary more with seed than baseline margins.This places dictionary-based interpretability among the paper’s less reliable causal measurements.
- The study’s scope is one model family, one galaxy sample, and Legacy-depth imaging; deeper or more crowded surveys may differ.The miss-shape estimate is least constrained because it uses 184 measured cases, although the 3.68 % miss rate is directly counted.
9 Conclusions
The conclusions identify metadata dependence as a systematic risk, quantify its tomographic cost, and report mitigations and tokeniser limits that persist across model scales.
- AION-1 defers to catalogue-derived metadata over pixels across every readout, including metadata that contradicts the image.
- Detection gating depends on field-centre presence rather than aperture photometry; deblending errors do not propagate, but detection and positional errors do.
- The effect grows with model scale.
- The measured detection-miss rate consumes a median two-thirds of the LSST DESC tomographic mean-redshift budget and exceeds it in 12 of 40 realisations.Observed positional errors raise the worst bin to 8.3× the requirement, and measured magnitude-dependent sampling does not change the result.
- Withholding the detection channel removes the vulnerability at no measurable cost, whereas a generic substitute mask is worse than either option.
- The image tokeniser collapses to ∼28 brightness-ordered states and the redshift readout is quantisation-limited; both persist under scaling.
- Across 15 dictionaries, recovery spans 26–75 %, shifts by up to 18 points on seed alone, and is not predicted by reconstruction quality.Thirteen of fifteen lose to difference-in-means.
- Recommended practice is to withhold unverifiable detection channels and read the posterior median rather than the mean.The authors also recommend one target modality per forward pass and probes attached to the encoder output.
A Engineering notes on the released code
The released code contains implementation and documentation pitfalls affecting batched decoding, spectrum tokenisation, codebook sizing, and preprocessing behaviour.
- All five engineering observations were made on polymathic-aion 0.0.2 with specified model revisions and were not verified against later revisions.
- Batched decoder targets corrupt every readout: median |∆z| rises from 0.0038 with [Z] to 2.350 with [FluxG, FluxZ, Z].Use one target per call.
- Repository HEAD adds spectrum sentinel padding absent from released 0.0.2, and omitting the required trailing λ = 99999 point changes 20 % of 273 tokens.
- The published image-FSQ specification describes the segmentation codec rather than the image codec, which has 4375 codes.Sizing analysis from the documented ∼212-code figure is wrong by −289 codes.
- HSC image decoding raises an attribute-name error on the exercised path, but tokenisation is unaffected and no reported HSC result is touched.
- Preprocessing mutates caller tensors in place, and tok z includes a 1025th sentinel code with no value bucket.
B The measured detection-miss selection
The Legacy Survey detection-miss rate is quantified over real cutouts and varies systematically with r-band magnitude. Missed objects are fainter, while magnitude-dependent smoothing is used for correlated draws.
- The empirical p(miss | mag) uses eight r-band magnitude octiles, each containing roughly 605 objects.Raw bin rates are smoothed toward the global rate with 25 pseudo-counts for correlated draws.
- 3.68% of targets are detection misses in the 5000-cutout scan, comprising 184 missed objects.The rate is measured over the same scan underlying the empirical p(miss | mag).
- Missed objects are a median 0.49 mag fainter than detected ones, with Mann–Whitney p = 5 × 10−24.The miss-rate trend is monotone across seven of eight r-band magnitude bins.
C The dictionary sweep in full
The full dictionary sweep compares widths, sparsities, and seeds under matched training, evaluation, and intervention conditions. Recovery varies across dictionaries, while the repository provides scripts and recorded outputs for verification and figure regeneration.
- The dictionary sweep in full: Every configuration uses the same 576,000-row activation matrix, 192 held-out objects, and intervention norm ∥v∥ = ∥vdiff-in-means∥.Only dictionary width, sparsity, and seed vary down the table, with PCA refit at each k.
- The dictionary sweep in full: 15 dictionaries are evaluated for recovery of the gate-induced flux collapse at α = 4.The first nine rows form a width-by-k factorial at seed 0; the last six vary the seed at k = 32.
- Verification and reproducibility: The repository maps numbered claims to producing scripts and JSON paths, and regenerates every figure from recorded JSON without live model calls.RESULTS.md includes the full experiment log and attached caveats.
- Verification and reproducibility: Reproducing the JSON requires the public data, a GPU-class device, and the pinned polymathic-aion 0.0.2 environment.The data are not redistributed; documentation specifies the on-disk schema and cross-match reconstruction.