Source-linked AI summary
The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models
Elzbieta Dawidek
TL;DR
The article examines whether frequency-based BPE tokenization preserves linguistically relevant structure in inflectional Polish. It analyzes tokenization patterns and grammatical form anchoring, finding that BPE stabilizes frequent written fragments rather than complete grammatical categories or forms, while proposing layered representations for more stable Polish modeling.
Problem
The paper asks whether BPE token boundaries preserve the grammatical structure and sentence relations encoded inside Polish word forms.
Method
The article analyzes Polish BPE segmentation across diagnostic words, texts, word families, inflectional forms, and subject-dialogical examples, using grammatical form anchoring and valency-oriented analysis.
Results
BPE stabilizes frequent graphemic fragments and statistical traces of grammar, but does not guarantee preservation of complete inflectional forms or their grammatical functions.
Takeaways & Limitations
More adequate Polish modeling requires sublexical stabilization alongside morpho-inflectional anchoring, sentence-valency representation, and subject-dialogical stability.
Takeaways & Limitations
The study does not directly measure how particular tokenization patterns affect generation quality and treats the grammatical “I” as an interpretive extension requiring separate dialogical study.
Abstract
from arXiv · showhide
This article analyzes BPE tokenization in Polish as a test case for the limits of statistical segmentation in an inflectional language. It asks whether frequency-based tokenization preserves linguistically relevant units, including orthographic form, phonemic and syllabic segmentation, derivational structure, inflectional endings, grammatical form, and the speaking subject. The material includes diagnostic words, a children's text, selected forms from the Preamble to the Constitution of the Republic of Poland, word-family tests, and examples with Polish diacritics and nasal vowels. BPE tokenizers may produce segments that coincide with syllabic or morphologically interpretable divisions, but remain dependent on the frequency of written forms. They do not systematically map orthographic representation onto phonemic structure or context-dependent phonetic realization. The results show that BPE stabilizes frequent surface fragments of grammatical exponents rather than grammatical categories themselves. A form such as ustanawiamy is not merely a sequence ending in -y, but a verbal form anchored in conjugation, person, number, tense, mood, and aspect. The article develops the concept of grammatical form anchoring. In Polish, forms such as poszlam, zrobilam, or bylam can establish the position of the speaking subject without an explicit pronoun. In interaction with AI, this exposes a further problem: a language model does not possess a stable grammatical "I", but reconstructs it contextually and may mirror the user's forms or shift grammatical gender. Roclawski's segmentation-flexional forms are proposed as a diagnostic framework for evaluating tokenization boundaries. More stable modeling of Polish may require sublexical stabilization, anchoring grammatical form in the inflectional system, representing sentence patterns and verbal valency, and maintaining the grammatical "I" in dialogue.
Introduction
The article examines whether frequency-based BPE tokenization preserves linguistically relevant structure in Polish, an inflectional language where words encode extensive grammatical information. It argues that BPE stabilizes written fragments rather than grammatical forms, phonological relations, sentence structure, or the speaking subject.
- Introduction: Polish tokenization must be evaluated by whether boundaries preserve grammatical structure, not merely by how many tokens represent each word.Polish word forms can simultaneously encode person, number, gender, case, tense, mood, aspect, and syntactic relation.
- Introduction: Grammatical form anchoring extends to the speaking subject because Polish first-person forms can express an implicit “I” without the pronoun ja.The article treats this subject-dialogical dimension as an interpretive extension rather than a separate empirical tokenization group.
- Introduction: The article proposes multilayered Polish modeling that combines sublexical stabilization, inflectional anchoring, sentence patterns, valency, and continuity of the subject in dialogue.Token-count efficiency is computationally important but does not by itself establish linguistic adequacy.
- Introduction: BPE operates at a graphemic-frequency level, stabilizing frequent written sequences rather than phonemic, phonetic, morphological, or grammatical units.Stable segments may locally coincide with linguistic units without preserving their function in the Polish language system.
- Introduction: Frequent endings such as -y or -ich may represent only fragments of grammatical exponents, while forms such as ustanawiamy and ogólnoludzkich require full inflectional interpretation.The relevant analyses include conjugation or declension and categories such as person, number, tense, mood, aspect, case, and gender.
- Introduction: Polish spelling and pronunciation diverge in cases such as morze/może, lód, nóż, and nasal-vowel forms, so tokenization cannot be assumed to encode phonemic or context-dependent phonetic structure.BPE may distinguish forms with equivalent segmentation-phonemic structure because their written sequences differ, while orthographic graphemes such as ą and ę have context-dependent realizations.
2. Methodology
The study uses a qualitative, exploratory analysis of words and short texts across multiple linguistic reference levels and three descriptive tokenizer environments. Its diagnostic material tests written, phonemic, phonetic, syllabic, morphological, syntactic, and subject-related relations without constituting a full Polish tokenizer benchmark.
- 2. Methodology: The study qualitatively compares BPE boundaries with graphemic, orthographic, segmentation-phonemic, phonetic-realizational, syllabic, logotomic, morpho-inflectional, and grammatical-sentential levels.The basic units are individual words or short test texts, with grammatical form anchoring analyzed as a separate interpretive category.
- 2. Methodology: The study is exploratory and diagnostic rather than a full benchmark, and reproducibility requires exact tool names, library versions, testing dates, token lists, identifiers, and unchanged inputs.Tokenization results may vary across tokenizer families, model versions, and testing environments.
- 2. Methodology: Three observed environments are labeled Bielik/APT4 Tokenizer, OpenAI-current Tokenizer, and OpenAI-legacy Tokenizer.The labels are descriptive because public documentation does not transparently specify the exact tokenizer versions used.
- 2. Methodology: The analysis records visible segmentation, token identifiers, token counts, character counts for short texts, and technical separator information.Newline-only or technical-separator tokens are omitted when they are not part of the examined word, and ambiguous visual boundaries are marked for verification.
- 2. Methodology: Five diagnostic groups include a high-frequency children’s text, complex official forms, a kazać word family, phonologically diagnostic words, and forms with nasal vowels and diacritics.The materials include examples from the Polish constitutional Preamble and tests of recurring surface fragments across related words.
- 2. Methodology: The reference levels are kept distinct because BPE operates on written form and frequency, whereas Polish also has segmentation-phonemic, phonetic-realizational, morpho-inflectional, and syntactic-valency structure.Orthographic analysis covers letters, digraphs, diacritics, and conventional written sequences; segmentation-phonemic analysis is inspired by Rocławski’s diagnostic framework.
1. Graphemic-frequency segmentation
Graphemic-frequency segmentation occurs when token boundaries are driven primarily by frequent written sequences. Such segments may be multi-character fragments but are not thereby phonological, phonetic, or morphological units.
- 2. Graphemic-frequency segmentation: Graphemic-frequency segmentation stabilizes recurring written fragments such as zię, ich, ani, aza, uj, or emy rather than linguistic units as such.In wdzięczni, segments such as zię, cz, or ni may reflect graphemic sequence frequency rather than recognized linguistic structure.
2. Syllabically convergent segmentation
Syllabically convergent segmentation occurs when token boundaries partially or fully coincide with syllable boundaries. This convergence does not establish phonemic, phonetic, or morphological alignment.
- 2. Syllabically convergent segmentation: The boundary in ław|ka coincides with the syllabic division ław-ka but not with the phonetic-realizational sequence ł-a-f-k-a.The mismatch reflects devoicing of w to f before k.
3. Segmentation convergent with logotomes
BPE segments may locally coincide with functional word components in Rocławski’s sense, but this convergence does not mean that BPE implements logotomes.
- BPE boundaries can partially coincide with functional word components used to describe word structure.The convergence concerns statistical segmentation of written form, not implementation of Rocławski’s logotomes.
- Such local correspondence may result from the frequency of the written form rather than representation of segmentation structure.
- Coincidence with a linguistically interpretable unit therefore requires cautious interpretation.
4. Morphologically interpretable segmentation
Some token boundaries coincide with prefixes, bases, suffixes, endings, or other interpretable fragments, but this does not establish preservation of the full grammatical form.
- Token boundaries may partially coincide with a prefix, base, suffix, ending, or other systemically interpretable fragment.Examples include final -ich in ogólnoludzkich and -ć in przekazać.
- A token can stabilize part of an inflectional exponent without anchoring the form in its inflectional paradigm.
- The presence of a morphologically suggestive segment is therefore not sufficient evidence that the full grammatical form has been preserved.
5. Mixed segmentation
Mixed segmentation combines several relations between tokens and linguistic structure, but the resulting pattern is not automatically better or worse.
- Mixed segmentation can combine syllabic-logotomic, written-unit, and other structural relations within one word.In pro|ś|ba, pro may be syllabic-logotomic while ś locally corresponds to a written unit associated with a sound or phoneme.
- The segmentation pro|ś|ba does not reflect phonetic realization because ś is voiced to ź before the voiced consonant b.
- Mixed segmentation should be interpreted as evidence of combined local relations, not automatically as superior or inferior tokenization.
6. Contextually non-phonetic segmentation
BPE can stabilize orthographic or inflectional fragments while failing to represent context-dependent pronunciation, complete grammatical forms, or grammatical categories themselves. The section also distinguishes interpretive categories for these partial correspondences and for model-generated shifts in grammatical form.
- 6. Contextually non-phonetic segmentation: BPE stabilizes orthographic representation in forms affected by devoicing, voicing assimilation, or written–pronunciation differences, rather than context-dependent phonetic realization.Examples include chleb, lód, nóż, prośba, and ławka.
- 6. Contextually non-phonetic segmentation: BPE may preserve graphemic sequences in forms such as morze/może, lód/nóż, and words with ą or ę without representing their phonological or realizational relations.
- 6. Contextually non-phonetic segmentation: In ustanawiamy, final -y may be stabilized while the full personal ending -my is split across fragments.The ending has not disappeared; only some fragments may be statistically stable.
- 6. Contextually non-phonetic segmentation: A single-token word is computationally efficient but does not by itself provide access to segmental, morphological, or inflectional structure.
- 6. Contextually non-phonetic segmentation: A statistical shadow of grammar requires a segment to satisfy at least two criteria, including recurrence, a potentially relevant structural position, or stable appearance across tokenizers.
- 6. Contextually non-phonetic segmentation: Isolated non-recurrent segments without a linguistic function are treated as graphemic-frequency fragmentation, while gendered form shifts belong to utterance generation rather than single-word tokenization.
2. graphemic-frequency fragmentation — when segmentation results from the frequency of written form but has no clear linguistic function;
The analysis treats Polish BPE tokenization as graphemic-frequency segmentation whose local linguistic correspondences do not establish grammatical form anchoring. It therefore evaluates token boundaries through their relations to linguistic description, sentence structure, and the speaking subject rather than token count alone.
- 2. graphemic-frequency fragmentation: The analysis distinguishes local token correspondences from grammatical anchoring, examining the relation between tokens and Polish linguistic-description levels.A token may converge with a linguistic fragment without constituting evidence of grammatical form anchoring.
- 2. graphemic-frequency fragmentation: The study is qualitative, exploratory, and purposive, using diagnostic examples to reveal segmentation mechanisms rather than estimate their frequency across Polish.Its material is selected to make mechanisms especially visible, not to support population-wide statistical estimates.
- 2. graphemic-frequency fragmentation: The study’s tokenizer comparisons are version- and environment-dependent, requiring model, library, date, input, token-list, and identifier details for full replication.The environment labels distinguish observed settings but do not substitute for official architectural documentation.
- 2. graphemic-frequency fragmentation: Interpretive classifications depend on linguistic expertise, while tokenization evidence cannot directly establish hidden model representations or full linguistic competence.The analysis separates graphemic, phonemic-segmentation, phonetic, and model-representation claims, and treats multilayer architectural proposals as a research framework rather than an implementation.
- 2. graphemic-frequency fragmentation: The analysis asks whether BPE remains graphemic-frequency based or locally converges with syllabic, logotomic, morphological, or phonemic segmentation.It also tests whether such convergence is systematic and whether BPE stabilizes linguistic units or merely recurring written fragments.
- 2. graphemic-frequency fragmentation: A central question is whether tokenization preserves grammatical form anchoring and morphological-family relations rather than only surface fragments.The analysis extends this issue to the grammatical “I” in human–language model interaction.
- 2. graphemic-frequency fragmentation: The sixth question is interpretive, marking a transition from tokenization data to grammatical-form stability in dialogue.The first five questions concern tokenization data and empirical Groups A–E.
- 2. graphemic-frequency fragmentation: The intended conclusion moves beyond token count toward preserving relations among written form, segmental structure, grammatical form, sentence position, and speaker position.The results are not claims about whether BPE recognizes language in a psycholinguistic sense.
3. Results
BPE tokenization in Polish stabilizes frequent written-form fragments rather than linguistic units or grammatical categories themselves. Its boundaries may locally resemble syllabic or morphological divisions, but they do not systematically preserve phonemic structure, grammatical form, or derivational relations.
- BPE remains fundamentally graphemic and frequency-based, stabilizing written sequences rather than phonological, phonetic, morphological, or grammatical units as such.
- Tokenization is not simply letter-based: multi-character fragments may stabilize, while boundaries can locally coincide with syllabic, logotomic, or morphological divisions without recognizing those structures.
- In diagnostic words, tokenizer differences reflect the depth of written-form segmentation rather than recognition of phonetic structure.
- Children’s text remained subword-segmented: the current tokenizer used 26 tokens for 83 characters versus 30 for the legacy tokenizer, without producing whole-word representation.
- Official and morphologically complex forms show that stable fragments such as -y, -ich, -ć, and -ni preserve exponent traces, not full grammatical forms or inflectional categories.
- Word-family tests similarly show recurring fragments such as aza, az, zak, pok, ć, any, uj, and emy without consistent preservation of derivational or inflectional structure.
4. Discussion
The discussion shows that BPE stabilizes frequent written fragments, but local agreement with syllabic, phonemic, or morphological units does not ensure grammatical anchoring. Polish tokenization must therefore be assessed across interacting linguistic levels rather than by token count alone.
- Polish tokenization must preserve relations among orthography, phonemic segmentation, phonetic realization, inflection, sentence structure, and the speaking subject.
- BPE operates on graphemic frequency, so local token stability does not establish equivalence with linguistic units.Tokens may resemble syllables, logotomes, morphemes, or endings while remaining products of frequency-based compression.
- Examples such as lód, nóż, morze, and może show that BPE stabilizes spellings without systematically representing phonemic or phonetic relations.The analysis distinguishes written sequences from /u/, context-dependent realization, and phonemic equivalence.
- Apparent segmentation correctness can hold at one level while failing at another, as with syllabic ław|ka and pro|ś|ba amid voicing alternations.
- Frequent fragments such as -y and -ich form a statistical shadow of grammar rather than preserving full grammatical categories.In ustanawiamy and ogólnoludzkich, interpretation requires conjugation or declension features beyond the visible ending.
- Rocławski’s framework is useful diagnostically because it identifies divergences among written form, segmentation, syllables, logotomes, and pronunciation without describing BPE’s operation.
5. Limitations of the Study
The study is qualitative, exploratory, and purposive, with limitations in material scope, reproducibility, classification, linguistic-level separation, and architectural interpretation. It identifies diagnostic mechanisms but does not provide a comprehensive benchmark or direct evidence of downstream generation effects.
- The purposive diagnostic material is insufficient for a full Polish tokenization benchmark across genres, styles, and frequencies.A larger and more diversified corpus would be needed for benchmark construction.
- Tokenization results depend on tokenizer versions, model families, libraries, interfaces, and testing dates, requiring complete technical reporting for replication.
- The expert classification of segmentation types would benefit from a second coder, explicit decisions, and procedures for resolving ambiguity.
- The analysis concerns input segmentation and potential consequences, not hidden representations or the model’s full linguistic competence.
- Phonetic realization is treated as a reference level distinct from both BPE and Rocławski’s segmentation-phonemic theory.
- The grammatical “I” and the proposed multilayer architecture remain interpretive extensions requiring separate dialogue studies and technical implementation research.The article does not yet determine how the proposed layers should be implemented.
- The article does not directly measure whether tokenization causes specific inflectional, syntactic, or dialogical generation errors.That question requires a separate experiment combining tokenization analysis with response evaluation.
6. Conclusions
The conclusions argue that Polish BPE should be evaluated by the linguistic relations it preserves, not merely by token economy. More adequate modeling requires coordinated sublexical, inflectional, sentence-valency, and subject-dialogical layers.
- Polish tokenization quality depends on preserved relations among written form, segmental structure, grammatical form, sentence structure, and the speaking subject.
- BPE stabilizes frequent orthographic sequences, while Rocławski’s theory serves as a diagnostic framework for divergences among writing, segmentation, and pronunciation.
- Stable token fragments in forms such as ustanawiamy and ogólnoludzkich do not equal anchored grammatical forms requiring inflectional features.
- Stable Polish modeling must connect forms with inflectional paradigms, verbs, and sentence argument structures.
- Because Polish can encode the speaking subject in verbal morphology, dialogue modeling must address contextual reconstruction and possible gender shifting.
- A more adequate approach combines sublexical stabilization, morpho-inflectional anchoring, sentence-valency representation, and subject-dialogical stability.The proposal is a multilayered research direction rather than a completed implementation.
Word Reference levels Main diagnostic function Interpretive comment
The diagnostic register separates reference levels for interpreting tokenization. Its examples show that frequent written fragments and diacritics may be stable without preserving morphological families or context-dependent phonetic realization.
- The appendix distinguishes orthographic, segmentation-phonemic, phonetic-realizational, syllabic, logotomic, and morpho-inflectional reference levels.
- Fragments such as -ani, -y, -ich, -ć, and -ni should be treated as graphemic-frequency effects unless linked to grammatical function.
- Word-family fragments such as aza, az, zak, pok, ć, any, uj, and emy do not demonstrate preservation of derivational and inflectional relations.
- Diacritics and written sequences in ręka, kąt, kąpiel, and wąski require separating orthographic representation from context-dependent phonetic realization.
Appendix A5. Interpretive Categories Used in the Register
Appendix B records the observed token segmentations and metadata used in the analysis, while distinguishing qualitative interpretive use from the requirements of full technical replication. Its examples show that BPE segmentations depend on written-form frequency and stabilize orthographic sequences without constituting phonetic or fully grammatical representations.
- Appendix B. Register of Observed Token Segmentations: Appendix B provides a minimal register of observed segmentations, token counts where available, visual data, and basic test metadata.The register does not include full token identifiers for all examples.
- Replication requirements: The appendix has an auxiliary control function, but full replication additionally requires token lists, identifiers, tokenizer and library versions, tool specifications, and test dates.The stated counting method should also distinguish technical tokens, separators, and newline characters.
- Observed examples: Observed examples show that segmentation depends on the frequency of written forms rather than consistently encoding linguistic structure.The register includes children’s text, inflectional forms, morphologically complex forms, and the kazać / pokazać / zakazać word family.
- Nasal vowels and diacritics: BPE stabilizes Polish diacritics and frequent written sequences as orthographic units, so forms such as rę|ka and ką|piel should not be read as phonetic representations.The appendix explicitly characterizes these segmentations as graphemic-frequency segmentations that do not reach the phonetic-realizational level.
- Replication scope: Missing token identifiers limit full technical replication, although the qualitative analysis remains unaffected because it concerns observed segmentation in relation to Polish linguistic-description levels.The outputs remain verifiable because the tokenizers and testing tools are publicly available or reusable in documented environments.