Source-linked AI summary
Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
TL;DR
Vision-language evaluations lack a model-independent way to separate image misreading from reasoning errors. This paper introduces and certifies a render ceiling by inverting known rendering geometry, then uses it to measure model deficits and identify extraction failures. On 2,160 rendered crystal structures the ceiling is 1.0000, while exact geometry as text closes under half the gap for 13 of 14 models and a supervised vision baseline reaches 0.8952.
Problem
Existing methods for separating perception from reasoning place a second model in the loop, so the reference itself can contain errors.
Method
The render ceiling inverts frozen cameras and re-solves cross-view correspondence for rendered known objects, with failures characterized by an enumerable set of projection coincidences.
Results
The phantom set is empty on 2,160 certified structures, yielding a ceiling of 1.0000; exact geometry as text closes under half the gap for 13 of 14 models, while supervised vision reaches 0.8952.
Takeaways & Limitations
The instrument separates image-supported information from model competence, exposes extraction-stage fabrication, and provides camera-placement rules for rendering-based benchmarks.
Takeaways & Limitations
The extraction ceiling cannot be measured directly because no detector meeting the required completeness and soundness conditions was evaluated.
Abstract
from arXiv · showhide
Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.
1 Introduction
The paper addresses the unresolved separation between image perception and downstream reasoning by introducing a model-free render ceiling for rendered crystal structures. It certifies the ceiling, uses it to attribute model deficits, and derives implications for model evaluation and benchmark design.
- 1 Introduction: Existing separation methods place a second model in the loop, so their reference can inherit perception or reasoning errors.Misreading an image and misreasoning about it require different remedies, making an exact reference important.
- 1 Introduction: Rendered crystal structures permit a model-free reference because known coordinates and frozen orthographic cameras can be inverted and cross-view correspondence re-solved.The oracle forward-projects known atom positions, triangulates them after correspondence is discarded, verifies matches, and applies the symmetry algorithm.
- 1 Introduction: 2,160 rendered structures have an empty phantom set and render ceiling 1.0000, so every model deficit on those samples is attributable to the model.Withheld cameras instead create projection phantoms and lower the ceiling, showing that the instrument measures image-supported information rather than assuming perfect renders.
- 1 Introduction: Exact geometry as text lifts all 14 models, yet closes under half the gap for 13 of them, while the perception-share relationship with model strength is unsupported.The median perception share is 0.2901, and the raw gain from exact geometry correlates negatively with pixel accuracy at ρ = −0.6439.
- 1 Introduction: A supervised vision model reaches 0.8952 on the same renders, above every vision-language model, demonstrating that the pixels contain readable structural information.The zero-shot best reaches 0.4429 against the oracle’s 1.0000, while the supervised pixel baseline exceeds every vision-language model.
- 1 Introduction: The instrument also yields camera-placement rules and exposes extraction-stage failures that ordinary downstream accuracy would misattribute to reasoning.Its construction transfers to benchmarks whose forward rendering can be written down and inverted; camera separation, rather than simply adding views, sets robustness.
3 Discussion
The render ceiling provides a model-free way to attribute benchmark deficits, showing whether renders contain enough information and separating extraction failures from downstream reasoning. Its certified geometry also supports camera-design guidance, while the method remains bounded by its rendering assumptions and residual upper-bound interpretation.
- A certified ceiling: Inverting known cameras and re-solving cross-view correspondence yields the exact answer supported by each rendered structure, failing only through enumerable projection coincidences.The oracle is model-free and its failure mode can be checked per sample.
- A certified ceiling: 2,160 certified structures have empty phantom sets and a ceiling of 1, so every model deficit on those samples belongs to the model.Withholding cameras instead produces phantoms and a ceiling below 1, showing that exactness depends on the protocol.
- Separating perception from reasoning: The ceiling reframes perception–reasoning attribution as subtraction from a model-free reference rather than comparison with another model's potentially erroneous output.Earlier approaches place a model inside the reference, causing attribution to inherit its errors.
- Separating perception from reasoning: 13 of 14 models have larger post-perception residuals than perception components, while a supervised pixel model exceeds every vision–language model, indicating that the renders contain extractable information.The residual is interpreted as remaining after perception is removed, and model strength does not track perception share.
- Fabrication and auditing: Extraction-stage fabrication can produce well-formed coordinate lists matching almost no real atoms, yet endpoint-only evaluation would score the resulting failure as reasoning.A model-free check for any known forward map supplies an audit primitive for such intermediates.
- Scope and limitations: The instrument requires a known invertible renderer, is demonstrated on one task and domain, certifies exactness at one tolerance, and leaves the residual as an upper bound rather than a direct reasoning measure.The model arms were not rerun on the scale-up sample.
- Benchmark design: Benchmark builders can use the geometry to choose cameras: four separated views empty the phantom set on this protocol, while placement controls tolerance to centroid error.Adding views never destroys identifiability, and a camera aligned with a lattice vector separates no atoms differing by its multiples.
4 Methods
The method fixes crystal labels and camera geometry, renders structures from five orthographic views, and reconstructs atom correspondences through an exact geometric oracle. Its certified ceiling is defined by the absence of projection phantoms under the stated separation and tolerance conditions.
- Benchmark construction: Crystal labels are derived from coordinates with spglib at a fixed tolerance, while unstable space-group assignments are quarantined.The canonical settings are symprec = 0.01 and angle tolerance 5°; all 220 stability-certified cases matched source-database space groups.
- Benchmark construction: Each structure is rendered as a 2 × 2 × 2 conventional-cell tiling through five known orthographic cameras at 768 px.The camera set contains three principal-axis and two oblique directions, with flat depth-sorted circles and no shading or bonds.
- Geometric oracle: The oracle forward-projects known atom positions, discards cross-view correspondence, and recovers candidates by back-projecting same-species image points across view pairs.Accepted points must lie within δ of same-species points in every remaining view, then nearby accepted points are merged using the same δ.
- Evaluation ladder: The render ceiling R1 is the fraction of samples whose oracle output matches the coordinate-derived crystal-system label.R1 is the model-free reference used to measure model deficit, while R3 supplies exact geometry as text and R4 uses pixel renders.
- Identifiability: Oracle soundness and completeness require independent camera directions and distinct same-species atoms separated by more than 2δ.The only remaining failure mode is the phantom set Φδ, consisting of geometric projection coincidences; the certified setting ties δ to the symmetry tolerance τ.
- Scope and conditioning: The ceiling bounds exact extraction but excludes pipelines using shading, occlusion ordering, or bond rendering, and camera alignment can leave lattice-direction differences unresolved.Adding views cannot destroy identifiability; with five views, tiling affects conditioning rather than identifiability on this protocol.
Supplementary Note 1 Tabular classifier: feature specification and read-
A tabular classifier represents crystal geometry with 19 features designed to capture metric, angular, size, and density information while remaining invariant to axis labelling.
- Feature specification: The classifier uses 19 features, including lattice parameters, cell volume, scale-free edge ratios, angle deviations, dispersion terms, site count, and density.Sorted edge-length ratios and dispersion terms are chosen for invariance to axis labelling.
- Classifier result: 0.8952 accuracy is achieved under the canonical feature specification.The reported count is 188/210.
Supplementary Note 2 Secondary axes and supporting figures
Supporting analyses test whether crystal-system prediction depends on cell metrics, occluders, model generation, and cue-sufficiency strata. They identify stable effects while narrowing one contrast to hexagonal or trigonal degeneracy.
- Secondary axes: 140/70 and 141/69 split the samples according to whether conventional-cell metrics alone determine the crystal system.The passage describes this as a stable property of the render convention rather than an accuracy mechanism.
- Cue sufficiency: 60/70 and 58/69 ambiguous structures are hexagonal or trigonal, concentrating the pixel-versus-cell-metric contrast near that degeneracy.After removing it, residual sample sizes are n = 10 and n = 11.
- Render conventions: Removing all occluders drops symmetry recovery by 75 to 78 structures, whereas removing informative occluders leaves recovery unchanged and random removal recovers most of the drop.The conditions are evaluated across both samples and view counts; atom-count matching is reported separately.
- Model progress: Newer models improve more on the ambiguous stratum but close less of their oracle headroom than on the sufficient stratum.The reported raw accuracy changes are 0.0286 to 0.5286 versus 0.6286 to 0.8357, while headroom closure is 54.1% versus 64%.
Supplementary Note 3 Label granularity
The reconstruction-based oracle remains stable as label granularity increases, whereas a cell-metric baseline declines because finer crystallographic labels require atom-position information.
- Label granularity: The cell-metric baseline falls from 0.8952 at crystal system to 0.6810 at space group.Point and space group depend on atom positions that lattice parameters do not encode, while all four oracle labels follow from one reconstruction.
Supplementary Note 4 Samples at a glance
The supplementary samples span training, evaluation, scale-up, controls, and replication sets, with reported oracle, baseline, and model results tied to each sample’s composition and size.
- Samples: 1,610 structures form the training split, while 210 structures form the exactly uniform evaluation sample with 30 structures per crystal system.The released benchmark combines these as 1,820 structures.
- Scale-up: 0.8738 is the 19-feature random-forest accuracy on the 1,933-structure quarantine-clean subset, against a majority-class rate of 0.1474.The corresponding canonical scale-up oracle accuracy is 0.8774.
- Scale-up: 1,995 structures comprise the full e50 scale-up draw before quarantine, with 285 structures per crystal system and conventional cells capped at 80 atoms.The quarantine-clean subset contains 1,933 structures.
- Scale-up: 1.0000 is the certified oracle accuracy on the 1,950-structure scale-up subset, while the released oracle accuracy is 0.8785.The shape-free floor is 0.2056 and the majority-class rate is 0.1462 on this subset.
- Replication: 0.9952 is the certified oracle accuracy on the independent 210-structure expansion sample, where the majority-class rate is 0.1619.The oracle is therefore not exact on this fresh draw, unlike the certified original and scale-up samples.
- Controls: 16 paired convention comparisons produce 0 significant results, with a mean of 6.3 discordant pairs among 70 structures.The largest observed effect is 0.0857, below its comparison’s detection threshold.
Supplementary Note 5 Render conventions: full statement
The render-convention analysis separates identifiability from conditioning: the tested protocols recover the same labels without noise, but camera placement, tiling, tolerances, and centroid noise affect robustness.
- Render geometry: 0.6704 is the mean fraction of atom-copies with at least one overlapping same-species neighbour disc across the five views and 210 structures.Every structure has at least one such overlap in at least one view.
- Centroid noise: At σ=0.02 Å, paired comparisons favor frozen over off-axis cameras by 102 structures to 2, with p=5.4 × 10^-28.The comparison differs only in camera placement.
- Centroid noise: 2.116× is the frozen protocol’s noise-tolerance ratio over off-axis cameras, below the naive conditioning-ratio prediction of 3.084×.The frozen, off-axis, and tiled protocols hold R1 ≥0.95 to σ values of 0.0146 Å, 0.0069 Å, and 0.0048 Å, respectively.
- Identifiability and conditioning: At zero centroid noise, frozen, off-axis, and tiled protocols all return every label, so their differences arise from conditioning rather than identifiability.The oracle-side tiling contrast is 178/210 = 0.8476 versus 200/210 = 0.9524 for the frozen baseline.
- Tiling and extraction: Tiling degrades robustness through eightfold atom-density increase and smaller same-species separations, so its arm cannot validly score camera-conditioning κ.The design requires well-defined centroids and extraction without spurious detections, not high recall alone.
- View count: At δ=τ, mean R1 rises from 0.9214 with two views to 1.0000 with five views, although 195 of 22,050 released-tolerance comparisons violate monotonicity.The theorem’s non-decreasing guarantee is proved at δ=0, not at the released δ>τ setting.
- View count: Five-view reconstruction has zero mean phantom excess, compared with 4.911 at two views, while all 210 structures recover the true atom count.Across 5,460 reconstructions, the oracle never drops an atom; failures are phantom points.
- Tolerance: The oracle’s phantom set is non-monotone in extraction tolerance: δ=0.005 Å causes one under-merging failure, while larger δ values introduce over-merging.At n=1950, the tighter setting costs one structure where δ=0.01 Å costs none.
Supplementary Note 6 Proofs
The proofs characterize when cross-view ray intersections recover crystal geometry and identify the phantom set as the oracle’s sole failure mechanism under the stated separation conditions.
- Corollary 6: Adding a view cannot destroy identifiability at δ=0 when the camera set contains two independent directions.This follows because every independent view pair reconstructs each atom at the unique intersection of its back-projection lines.
- Corollary 6: A view parallel to lattice vector t cannot separate atoms differing by multiples of t, because their projections coincide.This is the lattice-vector consequence of the projection map.
- Phantom set: The phantom set contains species-labelled points formed by cross-view coincidences that pass verification in every camera while remaining farther than δ from every same-species atom.Its definition depends only on the structure, camera set, and extraction tolerance.
- Tolerance: Setting δ=τ ties the oracle’s merge tolerance to the symmetry label map; δ>τ risks over-merging, whereas δ<τ risks failing to merge candidates from one atom.The symmetry tolerance enters only at the final downstream label-assignment step.
- Proof structure: The oracle accepts exactly the points in the phantom set together with genuine atoms, so an empty phantom set yields complete reconstruction after merging.The proof uses pairwise ray-intersection enumeration, remaining-view verification, and δ-distance merging.
- Noise scope: The noise extension is probabilistic: bounded per-view centroid errors imply a κ-scaled reconstruction-error bound with probability at least 1−2α over an enumerated view pair.The Gaussian example is handled through a high-probability bound rather than directly, because Gaussian support is unbounded.
Supplementary Note 7 Extraction share
The extraction-share analysis shows that ideal geometry does not measure real extraction, while one detector’s segmentation and species errors can fabricate downstream structures and obscure attribution.
- Measured extraction: 40/210 = 19.0% of original-sample structures and 47/210 = 22.4% of expansion-sample structures recover zero atoms with the evaluated blob detector.The pre-registered gate therefore treats R2 as a diagnostic rather than a scored rung.
- Measured extraction: R2 is 0.0762 on the original sample and 0.0857 on the expansion sample, below every scored model arm and the 0.5286 shape-free baseline.The evaluated detector has median per-view recall 0.400, precision 0.233, and centroid error 0.717 px on matched centroids.
- Fabrication: Species misassignment contributes to estimated over-triangulation in 39% of original-sample and 64% of expansion-sample affected structures.Replacing detections with oracle species labels raises R2 to 0.1762 and 0.1857, respectively.
- Fabrication: Dropped discs can let rays from different atoms pass cross-view verification, manufacturing phantom sites; extraction errors therefore propagate non-monotonically through triangulation.Matched centroids are sub-pixel accurate, locating the measured failure in segmentation rather than localisation.
- Scope: No detector satisfying completeness and soundness was evaluated, so the achievable extraction ceiling cannot be measured directly.Simulated dropout places R2 near 0.91–0.95 at recall 1.0, but this simulation removes species errors, jitter, and spurious detections.
- Scope: Simulated per-view recall must exceed approximately 0.975–0.995 for R2 to surpass the strongest VLM’s 0.7333, depending on sample.The estimate is a lower bound because the simulation omits several real-detector failure modes.
- Denominator: The fabrication analysis uses a 206-structure denominator because four gate-excluded structures count as unanswered in the full-sample figure.A model-in-the-loop split would lack a model-free reference for checking the excluded intermediate.
Supplementary Note 8 Supervised pixel-only baseline: full results
The supervised pixel-only baseline reads crystal-system labels directly from rendered images, achieving strong original-sample accuracy but substantially lower performance on an independent expansion sample.
- Headline results: ResNet-50 reaches 188/210 = 0.8952, while ViT-small reaches 0.8333 on the original evaluation sample.Each class contributes 30 structures; ResNet-50’s per-class accuracy ranges from 0.7333 for triclinic to 1.0000 for cubic.
- Replication, independent draw: ResNet-50 reaches 0.7524 and ViT-small 0.5810 on the independent expansion sample.These correspond to drops of 0.1428 and 0.2524 from the original sample, respectively.
- Error analysis: 31.8% of ResNet-50’s original-sample errors are hexagonal/trigonal confusions, the largest confusion pair.The pair is consistent with a cue-sufficiency finding that the drawn cell box genuinely underdetermines the distinction in the conventional setting.
- Error analysis: On the expansion sample, orthorhombic accuracy falls to 0.344, with 15 of 32 orthorhombic errors classified as monoclinic.The text attributes this dominant pattern to composition shift rather than lattice-geometry degeneracy.
Supplementary Note 9 Statistical procedures and per-test values
The statistical procedures define attribution quantities, test model-level patterns, and apply correction across the full hypothesis family while documenting confusion-matrix conventions.
- Model-level statistics: The median perception share against exact R1 is 0.2901 with 95% confidence interval [0.1492, 0.3576] across 14 models.A separate rank-correlation analysis reports ρ = −0.1588, p = 0.5877 and declines an ordering claim because n = 14 has little power.
- Model-level statistics: 200/210 models’ paired outcomes favor the certified R1 = 1.0000 ceiling over parity, with p = 1.5 × 10^-12 and a 65:0 result.Residual dominance across models is tested with an exact sign test, p = 0.0018.
- Confusion matrices: Supplementary Fig. 4 places true labels in rows and predicted labels in columns, with cell values giving structure counts.The largest off-diagonal mass in both samples is the hexagonal/trigonal pair.
- Multiple testing: The full family contains 26 hypothesis tests, with Benjamini–Hochberg false-discovery-rate control at α = 0.05.The passage states that every load-bearing claim survives correction.
- Attribution ladder: The attribution ladder defines P = R3 − R4 as the perception component and S = R1 − R3 as the post-perception residual.The share is P/(P+S), measured against the certified R1 = 1.0000 ceiling.
Supplementary Note 10 Reproducibility record
The reproducibility record documents implementation details, identifies render-description mismatches and unpinned dependencies, and reports oracle and extractor constraints.
- Repository record: The released repository documents labeling tolerances, feature specifications, random seeds, render and evaluation entry points, exact arguments, compute, and provenance.These materials extend the Methods and main-text data-and-code disclosures.
- Artifact disclosures: The emitted renders are ball-only discs with dashed unit-cell edges, not ball-and-stick images, because the axis-coloured cell-edge routine is never called.Both the render-module docstring and model prompt inaccurately describe the output as ball-and-stick.
- Artifact disclosures: Unpinned ASE versions can produce sub-pixel render differences because atom radii use 0.5× ASE’s tabulated covalent radii.The analytic oracle does not use rendered pixels, so its numbers are unaffected.
- Runtime and complexity: The oracle scales as O(V^2n^2) for candidate formation and O(V^3n^2) overall, with a median census time of 3.7 ms for 30 atoms on one CPU core.The full 1,950-structure census takes 16.9 s for reconstruction only and 25.5 s with CIF loading and conventional-cell construction.
- Extractor details: Four original-sample structures fail atom-list parsing and scoring gates, count as unanswered in full-sample accuracy, and are excluded from the 206-structure median-recall denominator.The median recall uses a tolerance of 0.15 in fractional-coordinate units, distinct from the Cartesian δ despite the same numeral.