Source-linked AI summary
Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers
Louis Yiven Zhu
TL;DR
The paper asks whether readers can determine what changed in frontier AI safety frameworks from developers’ own revision accounts. It builds a versioned corpus and coding procedure to measure silent change, finding that most material changes are undisclosed and that visibility varies with account form and change direction. It concludes that existing revision duties require justification but not the commitment-level enumeration needed for auditability.
Problem
Existing accountability regimes require framework updates or justifications, but do not require accounts to make material revisions legible to readers.
Method
The study assembles public framework versions and revision accounts, traces commitments across version pairs, and codes materiality and announcement status with a frozen codebook.
Results
67% of material changes are silent under the strict standard and 53% under the lenient standard; silence is established as directionally asymmetric but regime differences are suggestive.
Takeaways & Limitations
Revision duties should require enumerations of what changed, because justifications explain why a framework changed without making commitment-level revisions auditable.
Takeaways & Limitations
Independent inter-coder agreement was not reported, and the account typology may be affected by provider-selected regimes and selection among only eight analysed pairs.
Abstract
from arXiv · showhide
Frontier AI developers publish safety frameworks that commit them to evidencing whether their models are dangerous. The European Union and California now treat these documents as instruments of accountability, and both already impose duties on their revision. Neither requires the revision to be legible, in the sense that a reader could learn from the developer's own account what changed. We introduce the silent revision rate, the share of material changes to a framework's commitments that the developer's published account does not identify, and we release the versioned, hash-pinned corpus needed to compute it. The corpus contains every public version of the safety frameworks of the twelve developers that have published one, together with each provider's changelog, redline or announcement. We trace 710 commitment instances across twelve consecutive version pairs, code them against a frozen codebook, and adjudicate 244 individually. Three findings follow. First, 67% of material changes (95% CI 62 to 72) are silent under a strict standard and 53% under a lenient one, falling to 49% at section granularity. Second, silence appears to track the form of the account, since narrative announcements run at 74% against 63% for itemised changelogs, whereas account length in words barely matters; on the test that respects nesting the difference is suggestive. Third, 77% of traced changes weaken or remove a commitment, and in seven of eight pairs weakenings are more often silent than strengthenings. The statutory remedy therefore exists and specifies the wrong artefact. A justification explains why a framework changed, an enumeration states what changed, and only the latter makes revision auditable. We argue that publication duties should carry an enumeration duty, which one provider already meets, voluntarily and incompletely.
1 Introduction
The paper asks whether readers can identify material commitment changes from developers’ own revision accounts, and introduces a measure and corpus to answer that question. It argues that current revision legibility is low and uneven, despite frameworks’ growing accountability role.
- Frontier AI safety frameworks state how developers will decide whether models are too dangerous to train or release.
- The study asks how much change readers can identify, whether account form corresponds to silence, and whether visibility depends on change direction.
- The silent revision rate measures the share of material commitment changes that a developer’s account does not identify.The authors pair this measure with a corpus of public framework versions and revision accounts.
- The corpus contains 43 labelled framework versions, 9 companion documents, and revision accounts across twelve developers.The contribution includes confidence-interval estimates across twelve version pairs and a granularity check on the denominator.
2 Background
Prior work studies safety-framework content, disclosure quality, auditing, measurement, and organisational secrecy, but does not measure whether later document revisions are themselves disclosed. This paper positions silent revision as an outsider-oversight measure of the standards that audits rely on.
- Existing framework assessments generally score documents at a single point in time rather than tracking revisions longitudinally.
- Disclosure research measures whether information categories are reported, but not whether later document revisions are disclosed.The paper identifies revision disclosure as a second-order property of transparency.
- Auditing research supplies the institutional frame by linking public disclosure to credible third-party oversight.
- The closest methodological precedent tracks privacy-policy change, while AI Safety Watchtower documents unannounced edits without defining a tested measure.This paper adds providers’ own accounts as the comparator needed to measure revision legibility.
3 Corpus
The corpus collects every publicly released version of twelve developers’ catastrophic-risk frameworks after the Seoul Summit, together with each provider’s revision account and provenance records. It distinguishes four account forms and includes companion documents to track relocated commitments.
- The corpus includes framework versions and provider accounts such as changelogs, version histories, redlines, and announcement posts.System and model cards are excluded because they report point-in-time implementation.
- Each file was established from primary sources, stored with its extraction, SHA-256 hash, and provenance, and checked against cited passages.
- The manifest contains 52 rows: 43 framework rows and 9 companion-document rows.Seven providers have at least two labelled versions and enter the drift analysis.
- Companion documents help distinguish commitment removal from relocation, but withdrawn earlier texts can prevent readers from verifying reported changes.
- Revision accounts comprise redlines, itemised changelogs, narrative announcements, or no account.Across nineteen consecutive labelled pairs, these forms occur in five, six, four, and four pairs respectively; redlines show textual changes but not necessarily materiality or direction.
4 Method
The method traces commitments across consecutive framework versions, codes material changes and account visibility using a frozen codebook, and estimates silence under strict and lenient standards. It also tests account form and direction while acknowledging incomplete independent reliability coding.
- Unitising and tracing: Commitments are statements binding providers to practices, coded into nine categories and traced across consecutive versions.
- Materiality: A change is material when it alters scope, threshold or trigger, actor, obligation strength, disclosure scope, or consequence.
- Visibility coding: Changes are classified as announced, partially announced, or silent according to whether the provider’s account identifies the commitment and direction.
- Measurement: The silent revision rate divides silent changes by material changes, with strict and lenient readings differing by treatment of partial announcements.Rates are also recomputed at section granularity and reported with Wilson score intervals.
- Statistical tests: The analysis uses an exact permutation test for account form and rank correlations for account length and item count.The design avoids treating clustered changes within pairs as independent observations.
- Adjudication: 244 of 710 rows were individually adjudicated, changing 4 outcome codes and 54 announcement codes.The remaining 466 rows received a spot-check that changed one row.
- Reliability: Independent second-coder agreement was not reported at submission, leaving reliability as the method’s principal incomplete element.The authors identify this as limiting the precision of interpretive codes.
5 Results
Across revisions with an account, most material commitment changes are not identified by providers, with silence varying by account form and change direction.
- RQ1: Magnitude: 67% of 383 material changes were silent under the strict reading, compared with 53% under the lenient reading.At section level, the strict silent rate was 49%.
- RQ2: Form: 74% of changes were silent for narrative accounts versus 63% for itemised accounts, although the nesting-respecting test made this difference only suggestive.Account length did not track silence; the lenient difference disappeared.
- RQ3: Direction: 77% of 299 traced changes weakened, removed, or relocated commitments.Weakening was more often silent than strengthening in seven of eight pairs.
- Regulatory timing: After the revision-duty date, strict silence was 0.71 versus 0.61 before it, while the weakening share remained similar.The comparison carries no causal claim because only three pairs preceded and five followed the date.
6 Discussion
The findings show that revision legibility is measurable but low and directionally patterned. Because justifications explain why a framework changed rather than what changed, revision duties require explicit enumeration to make commitments auditable.
- Findings: Two-thirds of material change was silent under the strict reading and one-half under the lenient reading.Revision legibility was computable for every pair with an account.
- Interpretation: The direction result was established within seven of eight pairs, while the account-regime pattern was only suggestive under the nesting-respecting test.The regime pattern fit structural secrecy, but the direction pattern did not.
- Interpretation: Weakening-related silence fit decoupling better than structural secrecy, although the authors did not infer intent from changelog omissions.The account described additions more readily than changes that loosened commitments.
- Regulatory implication: Both regulatory regimes require revision-related publication but specify justifications rather than commitment-by-commitment enumerations.The paper argues that enumeration is the artefact needed for auditability; Anthropic provides a voluntary, incomplete example.
7 Limitations
The paper identifies unresolved reliability and inference limits: coding agreement and first-pass sampling parameters remain unreported, provider-selected account regimes may confound comparisons, and silence cannot establish motive or benchmark against other standards.
- Reliability: Inter-coder agreement is not yet reported, so interpretive codes rely on one adjudicated coding and precision is bounded.The authors state that this limitation does not affect the corpus or the direction of the findings.
- Reliability: The first-pass coder’s sampling parameters were not fixed, leaving seed re-runs outstanding.
- Account regimes: Provider-selected account regimes and only eight pairs leave selection effects unresolved, especially because Anthropic’s sole non-redlined revision is its largest.Redlines score zero by construction, while the provider chooses the account regime.
- Inference: The corpus lacks a comparison class, so the paper cannot determine whether two-thirds silent change is high relative to other regulated standards.The authors also state that silence measures external visibility, not motive.
8 Conclusion
Frontier safety frameworks now carry legal accountability across two regimes, but their revisions remain difficult to audit because providers need not enumerate what changed. The paper measures that revision legibility and finds substantial, directionally asymmetric silence.
- Conclusion: Two legal regimes make frontier safety frameworks load-bearing, motivating a corpus that reads their revisions across time.
- Conclusion: Readers cannot identify two-thirds of material change from providers’ own accounts, and invisible changes skew toward loosening.
- Conclusion: Revision visibility depends on whether an account enumerates changes, although neither legal regime requires enumeration.
Appendices
The appendices define the coding framework, document the commitment-matching and change-classification rules, reproduce calibration materials and cases, and provide supporting documentation for reproducibility.
- Reference materials: Appendix A defines terms, abbreviations and acronyms, while Appendices B and C reproduce the frozen codebook and calibration examples.
- Coding rules: A commitment requires the provider as grammatical subject and a commitment verb, modal, or descriptive present-tense practice; conditional commitments count.
- Coding rules: Commitments are matched across versions by category and object despite renumbering, relocation or rewording; unmatched later commitments are added.
- Coding rules: Changes are classified as retained, strengthened, weakened, removed, relocated or added, with mixed-direction changes marked W and MIX.
- Materiality: Materiality covers scope, triggers, actors, obligation strength, disclosure scope and consequences, excluding non-substantive editorial changes.
- Operationalisation: The coding rejects textual-edit counts and external quality rubrics because they respectively inflate editorial changes or measure framework level rather than commitment change.
- Announcement coding: Announcement status distinguishes identified directional changes, partial mentions without direction, and silent changes when the account does not identify the commitment change.
- Calibration cases: Appendix F verifies discussed cases against primary documents, including an announced Anthropic weakening and a silent OpenAI weakening.
D Statistical derivations
The statistical appendices specify interval estimation, regime comparisons, reliability measurement and category-level reporting. They document high silence, suggest a form-related difference under the primary nested test, and report that most traced changes weaken or remove commitments.
- Interval estimation: Wilson score intervals estimate silence rates from k silent changes among n material changes using z = 1.96.The interval is preferred because it remains within [0, 1] and performs better for small n than the Wald interval.
- Regime comparison: The primary exact permutation test assigns narrative labels across 56 nested subsets; the observed pooled strict difference is 0.106 with share 0.071 at or above it.
- Direction: 77% of 299 traced material changes weaken, remove or relocate commitments, with Wilson interval (0.72, 0.81).
- Inference limits: The appendix reports nine category rates with intervals but performs no tests or inference on differences between categories.
- Reliability: Reliability is planned through Krippendorff’s α, computed separately for materiality, outcome and announcement status with bootstrap 95% intervals.
E Robustness analyses
Robustness analyses show that silence remains substantial across granularity, account form, timing, and provider-level direction measures. Itemised accounts are associated with less silence than narrative accounts, while weakening changes remain common.
- Granularity: 49% of material changes are silent at section granularity under the strict rate, compared with 34% under the lenient rate.A majority rule instead yields a 66% strict section-level rate.
- Account length and enumeration: The strict silence rate correlates −0.17 with account length in words and −0.58 with the number of discrete items.The word-count relationship is weak, whereas the item-count relationship is stronger and negative.
- Timing: Strict silence is 0.61 before 1 January 2026 and 0.71 afterward, while the weakening share is 0.78 before and 0.76 afterward.The appendix reports these splits around the later version’s date.
- Direction by pair and provider: Nine of twelve version pairs show a weakening majority, and the leave-one-provider-out weakening share ranges from 0.70 to 0.79.The pooled direction result remains within this range when each provider is excluded in turn.
F Verbatim text for every change discussed in the main text
The traced examples show commitments being weakened, removed, relocated, or substantively altered, with provider accounts sometimes naming only a broader category or omitting the change. These examples illustrate why revision accounts must enumerate the affected commitment rather than merely describe a general policy shift.
- Framework restructuring: Anthropic’s monitoring-pretraining commitment changed alongside a broader restructuring of the RSP toward more realistic unilateral commitments.The post announced the class-level restructuring rather than enumerating this specific commitment.
- Additional traced changes: Other examples include a dropped open-release intention, a relocated staff issue-raising pathway, and softened or removed commitments involving pauses, weight deletion, and risk thresholds.These cases span silent, partial, and non-classified account treatments.
- Assessment cadence: A fixed six-month assessment minimum became event-triggered, while Microsoft’s changelog named the cadence change without identifying its direction.The adjudication classified this as a partial announcement.
- Thresholds and governance: Google DeepMind removed a quantitative R&D acceleration anchor without an account, and broadened a deployment gate while replacing a named governance body with an unnamed function.The first change was partially announced; the second was classified as a weakening and silent.
- Domain restructuring: Instrumental Reasoning Level 2 was removed and described only as incorporated into another domain, producing a partial announcement rather than an itemised identification.The provider account was absent in the traced record.
- Criterion replacement: Naver’s account identified replacement of one performance-based criterion with separate context, use-case, and impact criteria.The adjudication classified this change as announced.
M Reproducibility statement
The release provides a versioned corpus, manifest, revision accounts, codebook, coding sheet, scripts, and hashes for recomputation and provenance. Reproducibility is strongest for the corpus and manifest, while coding reliability remains incomplete because second-coder agreement was not reported.
- Release contents: The repository includes retrieved framework and companion documents, plain-text extractions, hashes, a manifest, revision accounts, a frozen codebook, and a 710-row coding sheet.The manifest records provider, version, date, source, retrieval provenance, and hash.
- Reproducibility tiers: Every reported statistic and appendix table is recomputed from the released coding sheet by the released scripts.The paper states that running the scripts reproduces the reported numbers exactly.
- Reproducibility tiers: The corpus and manifest are reproducible in the strongest sense because files are hash-pinned and retrieval sources are recorded.Provider documents are retained as published, including replaced documents as new rows.
- Reproducibility tiers: The first-pass coding used a language-model system, and adjudication was performed by the first author, so the coding sheet is reproducible only by rerunning the described procedure.A rerun may differ until planned seed reruns and second coding measure agreement.