Source-linked AI summary

Two locked tests of phase-structure features for transition prediction

Abraham Chachamovits

arXiv:2609.00335v1cs.CL

TL;DR

The paper tests whether phase-derived features improve endpoint ranking beyond a baseline that excludes them. Across two pre-specified studies, neither locked advancement rule was passed, so the extra empirical ranking lift was not found; the theoretical paper was not retracted.

  • Problem

    The paper addresses whether phase-derived components carry extra predictive information about commitment or contradiction endpoints beyond a baseline without those features.

  • Method

    It runs two pre-specified tests: a frozen, sealed PC-2-versus-baseline AUROC comparison and a fifteen-treatment development study using a conjunctive five-block selection rule.

  • Results

    Both locked tests were null: Study 1’s paired AUROC difference was +0.00087, and zero of fifteen Study 2 treatments advanced.

  • Takeaways & Limitations

    The extra ranking lift was not found under the rules written in advance, while the theoretical paper’s mathematical claims were not under test and were not retracted.

  • Takeaways & Limitations

    Study 2 produced a development-gate null rather than a final out-of-sample claim about held blocks b5 or b6.

Abstract

from arXiv · show

A published theoretical account of phase structure in rotary attention was subjected to two pre-specified empirical tests of whether phase-derived features improve prediction of a commitment or contradiction endpoint over a baseline that does not receive those features. Study 1 froze a contradiction-category pipeline and scored a sealed primary comparison of PC-2 against baseline. On 1,136 eligible cases the paired AUROC difference was +0.00087. The 99% interval included zero, and the difference did not reach the pre-specified threshold of +0.05. Advancement was not passed. Study 2 developed fifteen layer treatments on open blocks b0-b4 only (1,415 transitions, 20x5 grouped folds). A locked conjunctive rule required a positive PC-2 mean-repeat increment, a positive increment on at least four of five seed blocks, and a positive mean of those five differences. No treatment advanced. The official selection is null. The theoretical paper is not withdrawn. The extra ranking lift was not found under the rules locked in advance.

1. Introduction

The paper tests whether phase-derived features improve endpoint prediction, while distinguishing representational coherence from institutional admissibility. Two pre-specified studies implement this test, and both fail to advance.

  • Rotary attention is described as position rotation in paired planes, yielding magnitude-weighted cosine score terms and phase-sensitive trajectory measures.
  • Predictive value was tested by comparing phase-derived components with a baseline denied those components under rules written before decisive scoring.
  • The two studies used a frozen contradiction pipeline and a development-only pack with a conjunctive leave-one-seed-block-out rule.
  • The paper reports protocols, official numbers, and frozen record hashes without reopening held blocks or treating the layer-25 remainder as confirmation.

2. Study 1 methods

Study 1 froze fitting and scoring procedures before sealed evaluation. Its primary comparison tested PC-2 against baseline using a paired AUROC estimand and a two-part advancement rule.

  • Study 1 used a signed first-study protocol comparing PC-2 with a baseline-only readout on an authorized contradiction-category partition.
  • The primary estimand was the paired difference in AUROC, with advancement requiring a 99% interval excluding zero and a difference of +0.05 AUROC.
  • Ten thousand bootstrap resamples used seed 260725507, while PC-1 was descriptive and replication was scored only for direction.
  • Fitting closed before sealed scoring, with primary and replication roles serialized, hashed, reload-verified, and fit only on authorized open partitions.

3. Study 1 results

Study 1 did not pass its advancement rule: PC-2 added only a negligible, statistically compatible-with-zero AUROC increment over baseline, and replication did not rescue the result.

  • +0.00087 was the paired AUROC difference between PC-2 and baseline across 1,136 eligible primary cases.The primary interval was [−0.00464, +0.00692], and the +0.05 threshold was not met.
  • The 99% interval included zero, so the primary interval conjunct was not satisfied.
  • The replication difference was +0.00019 with the same sign, but it remained far below threshold and did not rescue the failed primary conjuncts.

4. Study 2 methods

Study 2 evaluated phase components on open development blocks using fold-internal features, grouped repeated cross-validation, and a locked conjunctive selection rule.

  • Study 2 compared PC-1 and PC-2 with a baseline stack for transition and commitment-endpoint prediction using open blocks b0–b4 only.
  • The development set contained 1,415 eligible transitions from 480 groups under a 20-repeat by 5-fold grouped schedule.
  • Features were estimated fold-internally, and block models plus nested stackers produced out-of-fold predictions without retaining raw layer arrays after checkpointing.
  • A treatment advanced only when PC-2 exceeded baseline in mean-repeat AUROC, at least four of five seed blocks, and the positive mean of those block-wise differences.

5. Study 2 results

Study 2 evaluated fifteen phase-feature treatments under a locked conjunctive advancement rule, and none advanced. Layer 25 was the only treatment with a positive mean-repeat PC-2 increment, but it was favorable on only three of five seed blocks.

  • Zero of fifteen treatments advanced; the official selection was null.
  • +0.000272 was the only mean-repeat PC-2 increment above baseline, produced by layer 25.
  • Layer 25 had a favorable sign on three of five seed blocks, below the required four-of-five condition.
  • Mean-repeat Δ was defined as PC-2 minus baseline AUROC in the locked selection table.

6. Frozen records

The paper records SHA-256 fingerprints for the cited pipeline, result, selection, and completion artifacts. These fingerprints identify the frozen records used in the study documentation.

  • SHA-256 fingerprints identify the records cited in the paper.
  • The frozen record set includes the Study 1 pipeline freeze and sealed result.
  • The frozen record set also includes Artifact 4 result and completion fingerprints.
  • Artifact 5 selection and completion are separately fingerprinted in the record.

7. Discussion

Both locked tests were null, while the theoretical paper’s mathematical claims were not under test. The discussion limits Study 2 to a development-gate verdict and prohibits reopening held blocks or relaxing the locked rule.

  • Both locked tests were null; Study 1’s baseline ranked well and Study 2 showed no consistent block-wise phase direction.
  • The tests evaluated phase-feature improvement for a ranking endpoint, not the theoretical paper’s phase decomposition or local presoftmax stability bound.
  • The results are consistent with separating internal coherence from institutional admissibility.
  • Study 2’s verdict is a development-gate null, not a final out-of-sample claim about held blocks b5 or b6.
  • Artifact 5 used the stored candidate table because the labeled transition table was not exported as a standalone file, and that dependence is declared.
  • Blocks b5 and b6 must not be opened, and Layer 25 must not be promoted on the basis of +0.000272.

8. Conclusion

The paper concludes that two pre-specified phase-feature tests are complete without advancement or selection. Its contribution is the pair of locked records showing that extra empirical lift was not found under the advance-written rules.

  • Study 1 did not advance and Study 2 selected nothing in the two completed pre-specified tests.
  • The contribution is the pair of locked records and the demonstration that extra empirical lift was not found under rules written in advance.

Author note

The work was developed independently at ENTRUST AI, with no competing interests declared, and reports the first completed empirical tests of one operational reading of the cited theoretical program.

  • The work was developed independently at ENTRUST AI, and no competing interests are declared.
  • The experiments are described as the first completed empirical tests of one operational reading of Chachamovits (2026), Sections 11–13.

Appendix A. Locked protocol for a later study

Appendix A specifies a locked protocol for a later study of external governance, distinct from the already tested AUROC comparison and not reopening Study 2. The study would use pre-registered sessions, fixed contract predicates, multiple locked endpoints, and a conjunctive advancement rule requiring fewer inadmissible terminals without excessive loss of trajectory diversity.

  • The appendix is a protocol, not a result, and may run only after the manuscript is dated and hashed; it does not reopen Study 2.
  • The later study asks whether an external governance predicate can reduce inadmissible terminal transitions without requiring uniformly higher internal coherence or collapsing trajectory diversity.
  • The protocol excludes PC-2 versus baseline AUROC, Layer 25, Study 2 blocks b5 and b6, and any relaxation of the Artifact 5 rule.
  • New governed and ungoverned sessions are evaluated under a pre-registered contract whose predicates are fixed before scoring.
  • Primary endpoints are inadmissible-terminal rate, contract-based false-accept and false-reject rates, trajectory diversity, and coherence used only as a covariate.
  • Governance advances only if inadmissible terminals fall by a pre-declared margin while trajectory diversity stays within a pre-declared loss bound; coherence change is neither necessary nor sufficient.
  • If the pre-declared margin is not met, the study is reported as null, and the protocol cannot be edited after the first scored session.
Loading 2609.00335v1…