Source-linked AI summary
A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
Liudmila Rozanova, Alexander Temerev
TL;DR
The paper tests whether Voynichese glyphs, blank-delimited strings, and separators can be treated as letters, words, and uniform word spaces. Using transliteration analyses, matched controls, and boundary validation, it finds structure concentrated within forms and at graded boundaries rather than in prose-like token succession.
Problem
The manuscript’s recurring marks, gap-separated strings, and transcription conventions are often assigned the linguistic roles of letters, words, and word spaces without direct evidence.
Method
The study analyzes the Zandbergen–Landini transliteration with matched prose, cipher, and pseudo-text controls, quire-level resampling, sequence learning, and image-coordinate boundary validation.
Results
Voynichese structure concentrates within forms and at graded boundaries rather than in prose-like token succession, while models reproduce several but not all profile features.
Takeaways & Limitations
Glyphs, tokens, and separators should remain descriptive categories until models demonstrate that they function as letters, words, and word spaces.
Takeaways & Limitations
The analyses do not establish that preserved or learned boundaries are lexical word boundaries, nor that the resulting learned-unit inventory is final.
Abstract
from arXiv · showhide
The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves. Glyph regularity is too strong for one-to-one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire-stable scale of recurrent multi-symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2-10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control. Blanks fall into two regimes: the separators transcribers marked uncertain behave like word-internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich-imitating cipher and a self-citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge-glyph coupling or the open, hapax-rich vocabulary (70% singleton types against 41% and 59-60%). Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.
1 Introduction
The paper argues that Voynichese’s recurring marks, blank-delimited strings, and separators should not be treated as letters, words, or uniform word spaces without testing. It introduces separate, testable analyses of glyph structure, token function, and separator behavior to determine what the data warrant.
- Motivation: Voynichese provides recurring marks, gap-separated strings, and transcription conventions, but transcription alone does not establish linguistic units.The observations become countable through transcription without thereby becoming letters, words, or word spaces.
- Unit hypotheses: The paper tests three conventional identifications: EVA glyphs as plaintext letters, separator-delimited strings as words, and all separators as identical word boundaries.These operational choices are convenient but embed linguistic interpretations in the input.
- Tests: Glyph, token, and separator hypotheses are evaluated separately using entropy and learned-unit analyses, lexical and sequential tests, and transcription, image-coordinate, and space-erasure tests.The methods distinguish glyph substitution, token vocabulary and order, and whether separators reflect physically or structurally different junctures.
- Related work: The study builds on prior reports of low character entropy, cross-token glyph dependencies, weak word-to-word predictability, and other non-uniform sequence and word-structure effects.The paper presents much of this raw material as previously documented while testing the unit interpretations systematically.
2 Data and methods
The study fixes explicit transcriptional units and tests them across multiple representations, matched controls, and quire-aware resampling. It reserves interpretive terms such as letters, words, and word boundaries until supported by evidence.
- Corpus and sampling: 3,880 lines from 206 pages, 50 bifolios, and 16 quires supply the boundary corpus, with 29,482 intra-line separators and 3,634 continuation breaks.Two heavily represented quires contain 47% of qualifying lines, making 16 top-level clusters an upper bound on effective independent units.
- Representations and terminology: The analysis distinguishes symbol, token, and separator from the interpretive terms letter, word, and word boundary, which require additional evidence.Frequent EVA composites are collapsed for the primary representation, while entropy, boundary, and BPE analyses use stated cleaning rules and corresponding symbol counts.
- Token-order analysis: 32,747 tokens from 3,950 paragraph lines support adjacent-token information tests under observed and uncertainty-merged tokenisations.The null permutes token order within lines 100 times while preserving frequencies, line membership, and line lengths; controls match the token count and line-length sequence.
- Boundary analysis: Boundary association scores compare glyph pairs across separator classes against within-token and independent-pair anchors, with edge-composition nulls testing marginal glyph preferences.The normalised index sets I = 1 to the within-token anchor and I = 0 to the independent-pair anchor, and is explicitly not a lexical-word probability.
- Entropy analysis: 34,175 clean tokens enter entropy analysis of composite-collapsed and decomposed streams after removing spaces and line breaks as symbols.The composite-collapsed stream is primary because common multi-character EVA conventions become single glyph-like units; decomposed EVA tests representation sensitivity.
3 Results
Voynichese has a plausible, open vocabulary but exceptionally weak token-order predictability, with sequential structure concentrated at token edges rather than token identities. Uncertain separators also behave as word-internal junctures, both statistically and physically, across the manuscript.
- Token sequence: 69.7% of Voynich types are singletons, while 7,022 distinct forms and 33.2% coverage by the 50 most frequent forms remain within control ranges.The result is not a lack of reusable forms, but weak prediction of neighboring token identities.
- Token sequence: 0.79% of token entropy is contributed by observed token order, below botanical Latin’s 2.02% and other continuous-text controls’ 3.59–9.72%.Merging all 2,352 uncertain separators reduces the order share further, from 0.79% to 0.54%.
- Token edges: Token-edge glyphs carry the robust sequential signal: adjacent tokens are 12–16% more often within one edit than shuffled, while the token-order deficit concerns where order resides.The edge-composition pattern is concentrated at boundary glyph pairs rather than whole-token identities.
- Separator classes: 0.494 is the internality score for uncertain separators, versus 0.029 for certain separators, a 0.465-unit difference that persists across all 16 quires.The uncertain-minus-certain contrast remains between 0.411 and 0.479 when each quire is omitted in turn.
- Physical separation: 0.1285 is the mean normalized box gap at certain separators, versus 0.0043 at uncertain separators; gap alone recovers labels with AUC 0.9053.The contrast persists after controlling for flanking glyph pairs and generalizes in leave-one-folio-out validation.
4 Robustness and limitations
Robustness checks support an early, replicated multi-symbol scale and graded separator boundaries, but they do not establish lexical status, an exact merge count, or universal language-level claims. The evidence is limited by local token-order measures, few clusters and controls, transcription-dependent coordinates, and estimator assumptions.
- Scope and limitations: Token-order evidence is deliberately local, and eight continuous-text controls plus 16 quires cannot exhaust long-range dependencies, linguistic typology, or language-level variation.The study therefore emphasizes effect sizes, cluster resampling, leave-one-cluster stability, matched-length controls, and the exact scope of substitution invariance.
- Separator evidence: Coordinate-based recoverability validates a transcription distinction, not lexical status; direct-pixel discrimination is weaker, and the coordinates lack a published construction protocol.Coordinates are quantised to a 3-pixel grid, and only junctures split by both tokenisation systems enter the analysis.
- Separator evidence: The association index partly embeds the within-token pairing it measures, so the result depends on convergence with box gaps, ink gaps, and space-erased unit learning.The edge-composition null removes marginal edge preferences but not the pairing bias; the v101 check supports reproducible weak-space distinctions without lexical analysis.
- Segmentation robustness: The held-out-quire minimum at 32 supports an early multi-symbol scale, but the exact merge count and 88-unit inventory remain analysis-dependent.BPE units may be ligatures, motor chunks, affixes, cipher groups, or recurrent template fragments rather than linguistic units.
- Segmentation robustness: 1.311 bits at k = 0, 1.049 at k = 32, 1.046 at k = 64, and 1.163 at k = 128 preserve a broad 32–64-merge trough.The trough remains after collapsing nine standard composites and after erasing separators; certain spaces are strong distributional boundaries, while others are graded.
5 Discussion
The discussion rejects treating glyphs, tokens, and separators as letters, words, and word spaces, replacing those assumptions with recurrent multi-symbol units, weak token ordering, and graded structural boundaries. Controls show that these diagnostics reject calibrated concatenative substitution but remain compatible with verbose homophonic respacing and content-free generation.
- Interpretive implications: The results reject direct identification of observed glyphs, tokens, and separators with letters, words, and word boundaries.These linguistic categories are hypotheses about observable manuscript structures, not their established identities.
- Interpretive implications: A transcription symbol has conditional entropy too low for fixed one-to-one substitution, while recurrent multi-symbol units persist across robustness checks.The learned scale survives Currier splits, quire resamples, composite collapse, space erasure, and out-of-quire applications.
- Interpretive implications: A token has plausible vocabulary structure, but preceding-token identity predicts it less than in every matched control, despite internal regularity and edge constraints.This distinguishes static vocabulary resemblance from word-like sequencing.
- Controls and scope: The joint rejection excludes fixed one-to-one letter and whole-word substitutions, while calibrated concatenative ciphers retain positive language differentials that Voynichese lacks.Concatenative homophonic ciphers show +0.13 to +0.52 against surrogates and +0.05 to +0.13 against the wrong language; Voynichese averages −0.01 and never exceeds +0.08.
- Controls and scope: The Naibbe control reproduces glyph entropy, learned scale, boundary permeability, and token order within a factor of 1.5, showing the attack does not reject verbose homophonic respacing.Its Latin is invisible to the attack, which therefore targets the calibrated concatenative family specifically.
- Controls and scope: Weak token order does not exclude meaning, while the self-citation generator reproduces several diagnostics, making them consistent with content-free generation.Lists, labels, paradigms, litanies, tables, and formulaic notations can carry content without continuous-prose sequencing.
6 Conclusion
The conclusion rejects treating Voynichese glyphs, tokens, and blanks as letters, words, and uniform word spaces without demonstrated boundary functions. It instead identifies recurrent multi-symbol units, weak token order, edge-glyph coupling, and separator regimes as the profile models must explain.
- Conclusion: Voynichese glyphs, tokens, and blanks should remain glyphs, tokens or surface forms, and separators until models demonstrate letter, word, and word-space functions.Treating the observed categories as letters, words, and uniform word spaces builds an untested decipherment into the data.
- Conclusion: 32 merges mark the out-of-quire minimum and 64 the full-corpus minimum for a quire-stable scale of recurrent multi-symbol units.Fine-grained symbol regularity resolves onto an early scale of recurrent multi-symbol units rather than one-to-one letter-like symbols.
- Conclusion: Blank-delimited forms carry less adjacent identity information than every continuous-text control, while their edge glyphs carry more than every prose control.This contrast holds at any vocabulary cap and distinguishes token interiors from token edges.
- Conclusion: A Voynich-imitating verbose cipher and a self-citation generator reproduce the 2.7-bit conditional entropy, learned-unit scale, and below-1% token-order signal.The paper evaluates how much of its measured profile two published accounts already earn, rather than treating any single entropy or frequency statistic as decisive.
- Conclusion: Neither published account establishes letters, words, or word spaces, even when reproducing entropy, units, and weak order sufficient to explain resistance to standard attacks.Models matching word lengths, Zipf-like frequencies, or visual appearance alone explain familiarity, not the manuscript’s full structure.
7 Data and code availability
The ZL transliteration, controls, coordinate files, analysis archive, data, code, and fixed-seed outputs are publicly available in a documented repository. The package accompanies the arXiv version, while Beinecke page images used for the direct-pixel audit are not redistributed.
- Repository and provenance: The public repository provides the ZL transliteration, control corpora, per-token coordinate files, and analysed snapshot at commit 66f8adaadc120f93e8a4c906685ba66a9d3ab847.Source provenance is documented in the repository and cited records.
- Reproducibility package: The publication package contains one driver per analysis, including scripts for boundary, entropy, coordinate, token-order, edge-order, calibrated-attack, control, and syllabic tests.It also includes the corresponding data, control corpora, and archived fixed-seed outputs.
- Reproducibility package: Every driver fixes its seeds and accepts input roots explicitly, while the Beinecke page images used for the direct-pixel audit are not redistributed.The package includes transcribed Naibbe tables, Greshko’s sample ciphertext, and generated self-citation streams under data/controls.
A Estimator and reproducibility details · A.1 Boundary-index specification
The boundary-index specification defines five reproducible boundary pools and a collapsed 34-symbol vocabulary, with smoothing, marginals, and bootstrap recomputation fixed in advance. Boundary labels, separator positions, and continuation line breaks determine pool membership and table inclusion.
- A.1 Boundary-index specification: The collapsed boundary-corpus vocabulary contains 34 observed symbols.This vocabulary underlies the within-token bigram table.
- A.1 Boundary-index specification: Add-0.5 smoothing is applied to every cell of the 34 × 34 within-token bigram table before normalisation.The smoothing precedes normalization for all table cells.
- A.1 Boundary-index specification: The same position-independent unigram vector supplies both PMI marginals and the independent-product expectation.This keeps the marginal and expectation calculations tied to one unigram specification.
- A.1 Boundary-index specification: Bootstrap duplicates of a quire duplicate its lines and counts before all tables and anchors are recomputed.Resampling therefore operates at the quire level before downstream quantities are regenerated.
- A.1 Boundary-index specification: The five reported boundary pools use ZL comma/dot labels, separator position within the line, or continuation line breaks.Uncertain and certain use the labels; first and mid use position; continuation joins consecutive nonparagraph-first lines on the same page.
- A.1 Boundary-index specification: Last-of-line intra-line separators are included in certain and uncertain totals but omitted as a separate main-table row.They contribute to the totals without receiving their own displayed row.
A.2 Coordinate fixed effect
The analysis estimates the pair-fixed effect after removing within-group means and uses folio-level bootstrap resampling with an iterative two-way sensitivity check.
- Estimation: The pair-fixed-effect coefficient is estimated after demeaning box gap and uncertain-label indicators within collapsed final-to-initial glyph-pair groups.Both variables are demeaned within the same glyph-pair groups before estimation.
- Resampling: Bootstrap resampling draws folios rather than individual boundaries and recomputes pair means when a folio is sampled repeatedly.Repeatedly sampled folios have all observations duplicated in the bootstrap sample.
- Sensitivity: The two-way sensitivity estimate alternately demeans gap and label within glyph-pair and folio groups until convergence.This sensitivity procedure combines glyph-pair and folio fixed-effect demeaning.
A.3 Blind direct-pixel audit
A blind direct-pixel audit of six folios found that uncertain boundaries have smaller ink-to-ink gaps than certain boundaries. The effect appears consistently across folios but is substantially weaker than the box-based discrimination.
- Blind direct-pixel audit: The audit measured ink-to-ink gaps directly on 700-pixel-wide Beinecke IIIF derivatives while remaining blind to the certain/uncertain labels.The frozen candidate pool contained 328 eligible intra-line boundaries, including 305 certain and 23 uncertain boundaries.
- Blind direct-pixel audit: 5/5 folios retaining both label types showed the predicted sign, with a one-sided sign-test p = 0.031.An equal-folio permutation test gave p = 0.021 one-sided and p = 0.066 two-sided.
- Blind direct-pixel audit: 14 boundaries were excluded during blind QC, including 12/277 certain and 2/23 uncertain cases, with Fisher p = 0.29.The two excluded uncertain cases were algorithmic no-ink failures.
- Blind direct-pixel audit: 4.98 pixels for certain gaps versus 3.14 pixels for uncertain gaps, a difference of +1.84 pixels.The distributions overlap heavily, making this ink-based contrast weaker than the box AUC of 0.905.
A.4 Additional robustness detail
Robustness checks preserve the token-order separation across form caps and quire refits, while boundary-regime differences survive strict coordinate alignment and line-length controls. The inference is manuscript-wide rather than universal, with quire-level replication but limited independence.
- Token order: 0.48%–0.83% is the observed-space estimate range across 16 leave-one-quire-out refits, versus 0.23%–0.55% for merged space.The high-cap Voynich estimate reaches the bias floor while the separation remains.
- Uncertainty model: 9.4 times larger is the bootstrap standard error under quire resampling than under line resampling, corresponding to about 89 times the variance.The result means line-level independence is not an adequate uncertainty model for corpus-level effects; leave-one-quire-out checks show replication within every observed top-level unit.
- Line-length control: +0.031 [−0.027, +0.082] is the internality estimate after restricting to the middle 50% of line lengths, while token length is −0.003 and entropy is −0.01.The restriction covers 33–54 glyphs across 156 pages and 1,768 lines and removes the apparent within-page gradients.
A.5 Calibration of the attack differential and the two further controls
The calibrated attack differential was small and similar for right- and wrong-language models in the Naibbe control, while the self-citation generator matched vocabulary sparsity only imperfectly and varied widely across seeds.
- Differential calibration: Five-seed calibration rebuilt each synthetic cipher and applied identical BPE, Brown-merge, and eight-restart mapping-search pipelines to real and order-3 surrogate ciphertexts.The differential was evaluated against right-language Latin and wrong-language German models.
- Naibbe cipher: +0.002 attack differential against held-out Latin and +0.013 against German were obtained for the Naibbe cipher.Per-seed Latin differentials were +0.006, −0.001, −0.041, +0.021, and +0.038; German included −0.008 and −0.010 for Pliny.
- Naibbe cipher: +0.01 to +0.06 differential at k = 128 occurred equally for the wrong language, while erased-space units crossed 2.2% of hidden token boundaries and 7.1% of hidden Latin word boundaries.The corresponding parenthetical rate for hidden token boundaries was 1.9%.
- Self-citation generator: 57–60% singletons was the self-citation mechanism’s saturation range across settings, below the 70% singleton rate reported for Voynichese.The port produced 2,210 types and 62% singletons, versus the published sample’s 2,228 types and 54%; global statistics varied widely between seeds.
A.6 Learned units against published inventories … A.10 Reproduction
Across learned-unit inventories, token edges, syllabic controls, and reproducibility procedures, the analyses reject simple equivalences between Voynichese units and letters, words, or syllables. The results instead identify cross-boundary structure, weak token-order dependence, and reproducible scale-sensitive diagnostics.
- A.6 Learned units against published inventories: 100% of compound units learned by 64 merges are slot-legal, while one third of Stolfi-model occurrences are fragments crossing layer boundaries.The crossing fragments chiefly involve mantle-plus-crust suffixes and crust-prefix-plus-core groups.
- A.7 Currier strata and line-start detail: 0.528, 0.179, and 0.203 are the first, other, and pooled line-start JSD values, versus relabelling nulls of 0.006, 0.002, and 0.001.The quire-bootstrap 95% interval for the first-minus-other difference is [+0.31, +0.40].
- A.8 Edge-glyph and token-identity order: full table: 0.218 bits is the observed-space Voynich edge information before correction, with y→q contributing 0.097 bits.Restricting to interior token pairs leaves the information at 0.226 bits.
- A.9 Syllabic-transcription hypothesis: four matched tests: 0.105 and 0.126 are the pinyin language-match indices for Currier A and B, against surrogate values of 0.138 and 0.117.The differential is approximately zero, compared with +0.39 for a genuine cipher over its own surrogate.
- A.10 Reproduction: 1,500 bootstrap draws, 300 matched-length windows, 100 shuffles, 100 quire partitions, 50 randomisations, and five cipher seeds are the stated reproduction defaults.Drivers take explicit input roots, fix random seeds, accept --json-output, and are documented with exact commands and smoke-test flags.