Source-linked AI summary

OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics

Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang

arXiv:2608.17162v1cs.LG

TL;DR

Fine-tuning objectives are typically evaluated after training, leaving unclear what different answer forms teach. OraclePhys controls answer form with exact finite-element grading and byte-identical inputs. It finds that ranking, written, and score-filtered supervision install behavioral physics capabilities, whereas boolean and advantage-weighted objectives do not within the tested recipes.

  • Problem

    What language models internalize from different fine-tuning objectives remains diagnosed retrospectively, limiting controlled evidence about how answer form shapes learning.

  • Method

    OraclePhys combines an exactly graded structural-mechanics benchmark, byte-identical supervision data with varied answer forms, and controlled training across verifier roles.

  • Results

    Ranking installs an out-of-distribution forward model, written or score-filtered answers install capabilities, while boolean and advantage-weighted objectives do not within tested recipes.

  • Takeaways & Limitations

    What the supervision label specifies about the target computation determines what fine-tuning teaches, while advantage-weighted scores suffice only for routing within the tested scope.

  • Takeaways & Limitations

    Results are limited to synthetic physics domains, 8B+LoRA models, GRPO, and short training runs, leaving full-parameter training and prolonged-RL regimes open.

Abstract

from arXiv · show

What a language model internalizes from fine-tuning is usually diagnosed after the fact. We make it an experimental variable. OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging; OraclePhys-30K, a supervision dataset of seven answer forms over byte-identical structure descriptions; and a controlled training study across the seven forms and three verifier roles. The study yields two findings. First, the label's answer form -- not its bit count -- causally determines what fine-tuning teaches: a ranking objective installs an out-of-distribution forward model where the untrained base sits at the guessing prior, a scalar objective at best a partial one, a boolean nothing detectable; the vector-scalar gulf survives a second physics domain, a second model family, and a paraphrased evaluation surface. Second, written or score-filtered answers install this capability, while advantage-weighted scores (GRPO) raise reward yet leave the model statistically equivalent to its start on held-out physics -- within the recipes and budgets tested -- sufficing only for routing. The trained 8B -- the first LLM on spatial structural response -- reaches the task's data-precision frontier: above a frontier LLM at zero- and 32-shot, at a specialist's level. What the label spells out about the target computation is what fine-tuning teaches; what you train on is what you route.

1 Introduction

OraclePhys isolates answer form as a causal experimental variable in LLM fine-tuning using byte-identical structural descriptions and an exact finite-element oracle. Its introduction reports that ranking labels install an out-of-distribution spatial forward model, whereas scalar and boolean objectives do not, while verifier pathways and data precision determine what capabilities emerge.

  • Framework: OraclePhys-Bench uses a finite-element solver to return exact per-story lateral responses for steel frames, including verified intervention-probe answers.The framework is designed as a controlled instrument for studying training, not as an engineering-practice leaderboard.
  • Controlled study: OraclePhys-30K holds models, structures, and input text byte-identical while varying only seven answer forms, including ranking, scalar, and boolean objectives.This controls the answer form’s information content while keeping the underlying training examples fixed.
  • Main finding: 0.86–0.91 OOD localization versus 0.19–0.25 untrained shows that ranking installs a forward model; scalar installs at best a partial model, and boolean installs nothing detectable.The vector–scalar gulf replicates on a second physics domain and a second model family.
  • Verifier roles: Written or score-filtered answers install capabilities, whereas advantage-weighted scores leave the model statistically equivalent to its start while reward climbs.Advantage-weighted scores nevertheless suffice to close a −0.15 action-interface gap, while an explicit reasoning channel adds nothing but installation cost.
  • Ceiling analyses: A frontier model sits near a 30-line formula, a specialist GNN matches the trained 8B, and six learners converge on a band set by data precision rather than task difficulty.The work presents the first LLM trained on per-story structural response.

2 An Oracle-Graded Testbed for Spatial Forward Models

OraclePhys uses a deterministic OpenSees finite-element oracle to grade story-drift rankings from natural-language descriptions, including counterfactual edits, without human or model-based judging. Its held-out, hardened testbed probes non-local spatial reasoning, out-of-distribution structure lengths, and verified post-intervention behavior against a simple non-learning reference.

  • Oracle-graded testbed: OpenSees finite-element solutions provide exact ground truth for per-story drift, eliminating human annotation and model-based judging in training and evaluation.The oracle computes inter-story drift ratios and identifies the governing story for each described frame.
  • Oracle-graded testbed: The model must output a JSON ranking of all stories by drift, evaluated by governing-story top-1 localization and full-ranking Spearman ρ on held-out structures.The evaluation uses n=150 per axis and n=80 for extrapolation.
  • Intervention and OOD tests: Counterfactual oracle grading tests whether the model tracks spatially non-local weakness when a story is strengthened and the governing story moves.The testbed also includes two sequential edits whose errors can compound.
  • Intervention and OOD tests: Tier-2 hardening reduces bag-of-words probe accuracy from 0.63 to 0.26 and structured-feature probe accuracy from 0.76 to 0.57, while the first-order reference scores 0.45/0.47.The hardened structures randomize section, weakening, and load-shape regularities; the reference directly reads numeric section properties and must be exceeded for physics claims.

3 The Objective-Shape Staircase

The objective-shape staircase shows that answer form, rather than bit count alone, determines what fine-tuning installs: ranking learns the strongest out-of-distribution spatial ordering, scalar learning is partial, and boolean learning is negligible. This gate survives calibration, alternate physics and model settings, and paraphrased evaluation surfaces, while control arms implicate exposed computation rather than format alone.

  • Cross-setting replication: P reaches 0.86–0.91 OOD localization across frames/Qwen, heat/Qwen, and frames/Llama, while A and O separate from base only in the primary setting.The vector–scalar gulf replicates on 2-D steady-state heat conduction and a second model family; O is not run on Llama.
  • Calibration: P 0.73 ≫ A 0.47 ≫ O 0.33 ≈ base 0.30 under calibration, with all pairwise reported improvements significant except O−base.P>A is +0.27/+0.29 and A>base is +0.17/+0.26 for forward/post-intervention measures; all p < 10^-4, while O−base is not significant.
  • Surface robustness: The gate remains unchanged under paraphrased held-out templates: calibrated localization is 0.75/0.44/0.31/0.29 versus the original 0.73/0.47/0.33/0.30.The paraphrase reorders fields and rewords phrasing while preserving information, questions, and parsing.
  • Control arms: The calibrated control suite orders base 0.30 ≈ S 0.35 < A 0.47 < K 0.61 = V 0.61 < G 0.68 ≲ P 0.73, separating format, scalar content, and ranking.The placebo remains near base; K recovers over half the gulf, while V—an information superset of ranking—matches K rather than P.
  • Emergent computation: P acquires untrained capabilities, including two-step rollout 0.72, taller-structure extrapolation 0.82±0.02, and post-intervention localization 0.71; A acquires them only attenuated at 0.42 calibrated.The same untrained axes emerge on heat at 0.82/0.73, indicating computation beyond replaying single-structure ranking statistics.

4 One Oracle, Two Uses: Answers Install; Advantage-Weighted Scores Do Not

Within the tested 8B+LoRA regime, dense oracle-written answers install transferable structural-mechanics knowledge, whereas advantage-weighted oracle scores improve reward without producing detectable held-out-physics gains. The asymmetry reflects dense token-level credit specifying the answer versus group-relative advantages reallocating probability among behaviors the policy already samples.

  • Answers install: Dense oracle-written rankings gain +0.09 to +0.14 across axes while tier-1 performance changes only from 0.887 to 0.867.Across three seeds, forward reaches 0.731 ± 0.011 from 0.660, while the post-intervention value is 0.691 ± 0.008.
  • Advantage-weighted scores do not: GRPO leaves forward, extrapolation, and two-step axes within noise despite reward rising from 0.80 to 0.93.On fresh n=600 streams, the dense endpoint leads GRPO round-2 by +0.087 (p=1.7 × 10−14) and the headroom run by +0.105 (p=4 × 10−18).
  • Advantage-weighted scores do not: From the untrained base, GRPO raises reward from 0.20 to 0.45 but reaches 0.433 on tier-1 and random performance of 0.167 on the hardened tier.These results indicate format compliance plus population-prior guessing rather than installed physics knowledge.
  • Controls: Best-practice GRPO reaches 0.751±0.030 against the 0.747 start, while headroom-matched GRPO reaches 0.678 ± 0.014 from 0.660.The corresponding mean gains are +0.004 and +0.013, with p=0.18 and p=0.81 respectively.
  • Mechanism: On stable-wrong instances, rewards nearly coincide, with ˆAi ≈0, so group-relative RL provides no gradient toward missing computations.The measured mixture is ∼60% stable-correct, 34% swing, and 6% stable-wrong; pass@1 is 0.796 and pass@50 is 0.941.
  • Scope: The dissociation holds within one task family, GRPO, 8B+LoRA, ≤800 steps, and lr 10−6–10−5, without addressing prolonged-RL regimes.The transfer asymmetry remains open despite the clean score-failure characterization.

5 Interfaces Are Cheap, Knowledge Is Not

The model contains post-edit knowledge but cannot initially use the state+action interface. Oracle-scored action training wires the interface, whereas dense answers install additional knowledge; explicit reasoning provides no measurable benefit and may reduce knowledge.

  • Interface wiring: The dense-SFT model’s post-edit response is 0.667/0.680, but applying “add braces to story 3” drops performance to 0.513/0.533, a −0.15 interface gap on both tiers.The model is installed with knowledge but is not wired to the state+action interface.
  • Interface wiring: Action-format GRPO closes tier-2’s interface gap exactly: act 0.533 →0.687 versus desc 0.687, with discordant 15:15 and p=1.0.Tier-1 performance is 0.76 versus 0.73, and forward and post-intervention axes reach parity.
  • Interface wiring: SFT on 4k oracle-labeled action pairs closes tier-2’s gap from act 0.533 →0.760, p = 6 × 10−7, at zero knowledge tax: forward 0.740 versus 0.747.On tier-1, post-intervention improves by +0.140, with discordant 26:5 and p = 2 × 10−4, reaching 0.867.
  • Knowledge installation: Either oracle-based method wires the interface, but only dense answers install additional knowledge at every tested score budget, making scores sufficient for routing and nothing else measurable.Dense supervision also wires the interface wherever answers can be demonstrated, while pure SFT yields the best all-around model.
  • Reasoning channel: Three pre-specified tier-2 contrasts find no reasoning-channel benefit; repairing the channel costs knowledge by −0.073, p=0.013, while thinking-on buys nothing.Think-RL versus matched-compute nothink-RL ties, favoring not thinking.

6 Ceilings: Frontier, Specialist, and the Data-Precision Frontier

The trained 8B surpasses the frontier model with 32-shot prompting and matches a specialist, while six learners converge near 0.73–0.78 because data precision, not task difficulty, binds tier-2 performance. More boundary-focused data improves this frontier more efficiently than additional volume.

  • Frontier model: Claude Opus 4.8 reaches 0.633/0.500 (tier-1/2), with at most +0.10 and +0.073 over the first-order formula in significant cells.The gains are significant for tier-1 forward (p = 10−4) and tier-2 post-intervention (p = 0.013), but nonsignificant in the other cells.
  • Frontier model: With 32 oracle-labeled worked examples, Opus declines from 0.500 to 0.433 on tier-2, while the trained 8B exceeds the frontier by +0.300.Opus’s change has p = 0.087, whereas the trained 8B’s advantage has p < 10−4.
  • Specialist baseline: The specialist GNN matches the final text-reading model on tier-2 at 0.733/0.733 forward and 0.680/0.680 post-intervention.On fresh streams, the pooled difference is bounded within ±0.06 (n=1200, n.s.), with the point estimate favoring the LLM.
  • Data-precision frontier: Six learners converge at 0.733–0.780 on tier-2, with no significant pairwise difference, defining a data-precision frontier rather than a task ceiling.The task is noiseless, while 34% of tier-2 instances have top-2 drift margins under 10%, making per-story error limits explain the observed band.
  • Data-precision frontier: Tier-1 saturates by n = 1k, whereas tier-2 gains only +0.02 at the last doubling; a difficulty-matched second round adds +0.09.The second round is worth three doublings, indicating that information near the decision boundary is scarcer than additional volume.

7 Related Work

Prior work probes or manipulates language-model world models, supervision, and verifier-scored reinforcement learning, whereas OraclePhys intervenes on answer form under fixed data and controlled oracle grading. In structural engineering, it extends translator-plus-solver and scalar-beam studies to controlled fine-tuning on spatial structural response.

  • World models in language models: World-model probing finds structured internal state, while audits show good prediction can coexist with a poor model of the generating process.OraclePhys contrasts with this literature by intervening rather than merely observing.
  • World models in language models: OraclePhys intervenes on the objective’s answer form while freezing byte-identical data, unlike prior data-property or supervision-content manipulations.Earlier work manipulated data properties, fixed-answer traces, or logic-tuning formats without byte-identical inputs or counterfactual grading.
  • Limits of verifier-scored RL: OraclePhys measures the contested capability question for verifier-scored RL with one verifier and one start, finding that scores route while dense answers dominate.This tests the premise that RL needs the right SFT seed and provides a positive characterization of score-based routing versus answer-based capability.
  • LLMs for structural engineering: Prior structural-engineering systems translate problems for hardcoded solvers, while Hage and Buehler (2026) RLtune a compact LLM on scalar beam statics.OraclePhys is presented as the first LLM trained on spatial, per-story structural response and the first controlled study of one.

8 Conclusion

OraclePhys makes what fine-tuning teaches a controlled variable using structural mechanics as a model organism. It finds that answer form determines learning: written or score-filtered answers teach, whereas advantage-weighted scores only route.

  • OraclePhys turns what training teaches into a controlled variable, using structural mechanics as a model organism.
  • The label’s answer form determines what fine-tuning teaches.
  • Written or score-filtered answers teach, while advantage-weighted scores route.This conclusion holds within the recipes tested.

Limitations

The study’s conclusions are limited to 8B + LoRA recipes, synthetic solvable-physics domains, behavioral evidence, ordinal evaluation axes, and narrow evaluation and statistical scopes. Full-parameter training, prolonged reinforcement learning, mechanistic probing, broader domains, tool use, and larger-seed studies remain open.

  • Experimental scope: All results use 8B + LoRA, synthetic solvable-physics testbeds, and GRPO at ≤800 steps with group sizes up to 50.Full-parameter training and prolonged-RL regimes remain open; the replication supports recipe robustness, not domain generality.
  • Interpretation scope: “Installs a forward model” is an operational behavioral claim, not a mechanistic conclusion supported by linear probes.The probes decode equally from every arm, including behaviorally inert O, at the surface baseline; probing during generation remains open.
  • Evaluation scope: Post-edit structures belong to a family seen in training, and all evaluation axes are ordinal rather than value-prediction axes.Magnitude-trained arms may install value models these axes do not credit; “installs less” therefore applies only on these axes.
  • Evaluation scope: Descriptions use one fixed template and one in-context control, while the frontier comparison covers one model, one k, and text-only evaluation.A paraphrased evaluation surface survives, but many-shot evaluation and tool use are untested.
  • Statistical scope: Gate arms carry three seeds, whereas control-suite arms K/V/G/S are single runs; both decisive GRPO cells carry three seeds.“Statistically indistinguishable” denotes failure to reject except where a TOST-style bound is provided.

Ethics Statement · A Reproducibility · B Testbed details

OraclePhys uses a fully synthetic, deterministic testbed with exact reproducibility infrastructure and explicit limits against real structural-design use. Its benchmark spans procedurally specified structural problems, hardened distributions, a heat-conduction analogue, and released oracle-scored data artifacts.

  • Ethics Statement: OraclePhys is fully synthetic, uses no human subjects, annotators, or personal data, and its simplified 2-D linear FE model must not guide real design or code compliance.The released testbed and checkpoints are research artifacts for studying training objectives, not engineering-practice tools.
  • A Reproducibility: Every instance stream is deterministic, results files store per-instance correctness arrays, and one script regenerates every paper number from archived results.Greedy-decoding benchmark reruns reproduce aggregates exactly; comparisons use exact McNemar tests with paired bootstrap confidence intervals.
  • B Testbed details: Training problems contain 3–4 stories, while the OOD pool contains 5–6 stories under held-out geometry, with designs varying sections, connections, bracing, and damage events.Problems specify geometry, constraints, and lateral loads with vertical profiles Fx(s) ∝sα.
  • B Testbed details: Tier-1 sampling includes exploitable design regularities, while Tier-2 hardening randomizes column profiles, weakening events, and load shapes.The governing-story classifier scores 0.63 from bag-of-words features, 0.39 from geometry-only features, against 0.18 chance.
  • B Testbed details: The oracle ranks stories using load divided by a stiffness proxy, while a heat-conduction isomorph extends the testbed beyond structural mechanics.The proxy aggregates column sections, connection type, brace areas, and damage multipliers; direct design-object features score 0.53/0.47 on tier 1 and 0.45/0.47 on tier 2 for localization/post-intervention.
  • B Testbed details: Released JSON exports provide self-contained descriptions, exact prompts, oracle truths, calibration prompts, fingerprints, and a pure-numpy parser for external scoring.Each axis includes 150 instances, with 80 for extrapolation, and the benchmark includes governing-story and full-ranking truths where defined.
  • B Testbed details: Each released file includes integrity anchors, including tier-1 reference fingerprints of 0.5333/0.4733 and a SHA-256 hash over all truths.Paired or counterfactual streams regenerate deterministically from released code, and per-instance correctness arrays enable external model comparisons.

C Training details · D Alternative-explanation sweep for §4 · E Thinking-channel details (§5)

The study fixes answer-form supervision and training recipes while testing whether optimization, exploration, capacity, or thinking-channel effects explain the observed physics transfer. Across these controls, written labels install transferable capability, whereas scalar-reward GRPO improves in-pool reward without detectable out-of-distribution gains, and thinking adds a measurable bridge cost without reliable benefit.

  • C Training details: The three gate arms share byte-identical structure descriptions and differ only in their question–answer suffix, using controlled LoRA fine-tuning settings.The setup uses Qwen3-8B, with Llama-3.1-8B-Instruct for family replication, and audits token-level leakage.
  • C Training details: OraclePhys-30K contains seven answer-form arms with 3,760 byte-identical rows each, plus approximately 4k action-interface pairs and a hardened second-round curriculum.Every label except the S placebo is generated programmatically from the finite-element oracle’s solution.
  • D Alternative-explanation sweep for §4: A graded reward combines 60% full-profile rank correlation with 40% localization credit, while a parse floor preserves gradient toward well-formed answers.The reward already includes both profile-shape and direct-localization signals, so scalar densification does not provide label content.
  • D Alternative-explanation sweep for §4: Reward rises 0.80 →0.93 and in-pool localization 0.74 →0.88, but OOD localization moves only +0.01, showing optimization can satisfy the objective without transferable physics.A headroom-matched relaunch reproduces the pattern: reward 0.62 →0.69, in-pool localization 0.41 →0.55, and OOD +0.013.
  • D Alternative-explanation sweep for §4: The round-2 unsolved-prompt recipe changes OOD localization from 0.667 →0.687 (p=0.15), while oracle-written answers gain +0.09 to +0.14 on every axis.These results argue against exploration failure or parametric capacity as sufficient explanations for the transfer gap.
  • E Thinking-channel details (§5): Direct-answer SFT causes 100% of thinking-mode generations to hit the token limit, whereas a 4k rationale bridge restores think-mode termination with zero clipping.The bridge uses automatically generated “evidence lines →verdict →answer” chains grounded in structure features and finite-element ground truth.
  • E Thinking-channel details (§5): Thinking-on performance changes are unreliable: forward improves +0.033 (p=0.33), while post-intervention changes −0.013 (p=0.81).The contrasts use n=150 instances and exact McNemar tests.
  • E Thinking-channel details (§5): Think-RL scores 0.707 versus 0.760 for nothink-RL, with 95% CI [−0.133, +0.027] and p=0.27, favoring not thinking point-estimatively.The comparison uses the same SFT-hard ancestor; the gap is commensurate with the bridge tax.

F Full results tables

Tables 2–9 provide the paper’s core per-arm results, tiered evaluations, cross-domain tests, paired comparisons, powered endpoint contrasts, and paraphrased-template evaluations. All reported values regenerate from archived result files, with exact paired testing and confidence intervals specified where applicable.

  • Core evaluations: Tables 2–4 report results for tier-1 and hardened tier-2 frames, plus a heat-conduction isomorph, across localization and extrapolation axes.Table 2 uses n=150 per axis and 80 for extrapolation; Table 3 uses the same sample sizes, while Table 4 is single-seed.
  • Core evaluations: Table 2 separates raw outputs from a shared 2-shot format-calibration block to distinguish missing knowledge from output-format failure.Cross-arm judgments use the calibrated fmt2 rows reported in Table 5.
  • Statistical comparisons: Table 5 gives exact McNemar tests and 95% paired-bootstrap confidence intervals for key tier-2 comparisons, with Benjamini–Hochberg FDR across 32 tests.Rows prefixed “fmt2:” refer to the tier-1 stream under the 2-shot format calibration.
  • Statistical comparisons: Table 8 reports powered direct endpoint contrasts on fresh n=600 tier-2 streams pooled across three disjoint OOD axes, including TOST-style equivalence within ±0.05.The comparisons use canonical seed-42 checkpoints, exact McNemar tests, and 95% paired-bootstrap confidence intervals.
  • Robustness evaluations: Table 9 evaluates frozen checkpoints on a paraphrased linguistic surface while preserving the same held-out instances, information, questions, and parser.The re-rendering uses running prose, reordered fields, and reworded connection, bracing, and load phrasings without retraining.
Loading 2608.17162v1…