This is what I have to say about the claims the paper makes:
Summary of the claim
The paper computes nine character-level metrics over the manuscript in four transcriber layers, four natural-language texts, one Naibbe ciphertext, and two shuffle controls, and concludes that the manuscript's statistical profile is characteristic of verbose, positionally constrained encoding and is not reproduced by natural language, so research attention should move from linguistic typology to cryptographic mechanism.
The implementation is careful. Exclusion lists, splat-marker handling, OOV bucketing, numeric folio sorting and a public repository are all handled properly, and testing across transcriber layers is the right instinct. The objections below concern what the measurements can support.
1. The positive result is circular
The Naibbe tables were derived with the manuscript's word shapes as the target, and Greshko already reports that the output reproduces flattened glyph frequencies, low entropy, rigid word structure and positional patterning. The statement that no parameters were tuned describes this study's procedure, not the provenance of the parameters, which were fitted upstream. A mechanism reverse-engineered from a target cannot serve as a test of that target; it can demonstrate sufficiency, which the source paper already did.
Calling the cipher a control inverts the term. A control is a case whose value on the outcome is not fixed in advance by its construction. The paper compounds this by declining to advance the cipher as a historical proposal while retaining the inference that the manuscript was likely produced this way, which places the historical plausibility of six tables and per-letter randomization over 38,000 tokens beyond scrutiny while keeping its evidential benefit.
The corresponding omission is decisive. A copying-and-modification process predicts low conditional entropy, small contextual gain and high edit-distance overlap without being fitted to anything, and is the one candidate in the field that could have failed this test. Its absence from the control set, and from the bibliography, leaves the conclusion resting on a single fitted artifact. Shuffle controls are not a substitute: no one proposes them.
2. The paper names the wrong effect
Computing medial minus initial positional entropy per corpus:
Quote:Condition Initial Medial Med-Init
------------------------------------------------------------
VMS recipe, recto->verso 3.1665 3.3003 +0.134
VMS recipe, verso->recto 3.1657 3.2967 +0.131
VMS herbal (f1-f66) 3.0711 3.3294 +0.258
VMS intermediate (f67-f102) 3.0649 3.3620 +0.297
VMS all folios 3.1514 3.3556 +0.204
VMS, Friedman layer (F) 3.1719 3.3366 +0.165
VMS, Currier layer ( C) 3.0892 3.3581 +0.269
VMS, Stolfi layer (U) 3.0526 3.3886 +0.336
Naibbe ciphertext 3.1748 3.3772 +0.202
------------------------------------------------------------
Esperanto 4.0379 3.9938 -0.044
Latin 4.0717 3.8966 -0.175
Finnish 3.8337 3.7874 -0.046
English 4.0870 3.9636 -0.123
------------------------------------------------------------
Within-token shuffle 6.0139 4.7258 -1.288
Global shuffle 6.1495 4.7464 -1.403
------------------------------------------------------------
Positional entropy, bits per event. Values from Kinnison (2026);
the third column is my subtraction. Every Voynich condition and
the cipher have medial above initial; every natural language and
both shuffles have medial below.
Every Voynich condition and the cipher have medial above initial; every language and both shuffles have medial below. Relative to its own baseline, what is depressed in Voynichese is the initial position. Initial position also separates more strongly in absolute terms: a gap of roughly 0.8 bits against the language controls, against 0.5 bits at medial.
The paper designates medial position as its interpretive focus on the stated ground that it separates best. That criterion is post hoc, no structural reason for expecting the medial position to be informative is offered, and the selection is incorrect by its own standard. The sign of the initial-medial difference is the one clean discriminator in the results and goes unremarked.
3. Four results contradict the paper's own conclusions
- Final-position entropy. Claimed to be constrained relative to natural-language controls. Finnish is 2.3588, below every Voynich condition but one, and well below the herbal section (2.6116) and the Stolfi layer (2.5911).
- Levenshtein ≤1, token-weighted. English 97.0764; VMS all folios 97.0719; recipe (96.20) and herbal (95.85) below English. This column separates nothing. Only the type-weighted column does, and there the manuscript (83–91%) lies between the languages (57–75%) and the cipher (99.84%), not with the cipher.
- Bigram legality rejection. Naibbe 0.0063 < Finnish 0.037 ≈ Latin 0.039 ≈ English 0.052 < VMS Friedman 0.089 < VMS all folios 0.201 < VMS Currier 0.341 < Esperanto 0.360 < VMS recipe 0.433 < VMS herbal 0.565 < VMS V→R 0.844 < VMS Stolfi 1.024. The manuscript occupies the top of the range and the cipher the bottom, by two orders of magnitude.
- ΔBPC. Naibbe (0.2876) falls outside the Voynich range (0.2165–0.2774) on the paper's second headline metric.
The pattern is systematic: the cipher matches on every measure dominated by glyph inventory and word shape, which is what its tables were fitted to, and diverges on both measures of lexical productivity, which they were not. That is the signature of a fitted artifact and is evidence against the identification.
4. Metric definitions
Final-position entropy is a smoothing artifact. The event has a single observable outcome, so unsmoothed it contributes zero bits. Under Laplace the reported value is (c+1)/(c+V) transformed, a monotone function of the training count of each token-final character. The metric ranks final-character inventory diversity, not conditional uncertainty. The subsection defending the non-zero values presents the artifact as the consequence of a design choice rather than as the thing requiring justification. It should be dropped, or redefined so both outcomes are observable: given a character anywhere in a token, the probability that the token ends after it.
The positional and n-gram measures are mutually inconsistent. With mean token length near 5, the positional values imply roughly 3.1 bits per event against BPC2 of 2.2763 for the same text and model order. The methods note that smoothing vocabularies differ by event type and never reconcile the two. As reported, positional bits are not comparable to BPC, across event types, or across corpora with different inventories.
Levenshtein ≤1 conflates recurrence with proximity. Distance 0 qualifies, so a text repeating five words scores 100%. English reaches 97% because roughly half its tokens are high-frequency function words. Reporting distance 0 and distance exactly 1 separately would make the column informative; as defined it measures token coverage, which for the language controls tracks morphological richness.
The nine metrics are not nine diagnostics. ΔBPC is the difference of BPC2 and BPC3; the positional measures decompose the same bigram model by event type; the two Levenshtein columns are one computation under two weightings. The table represents roughly three underlying quantities. Presenting them as a battery of independent confirmations overstates the evidential weight, and the figure plotting medial entropy against ΔBPC plots a quantity against a function of itself, so the clustering is largely built in.
5. Uncontrolled parameters
Training-set size. A Laplace-smoothed trigram model on a partition of a few thousand tokens is dominated by smoothing mass, inflating BPC3 and compressing ΔBPC. The manuscript's partitions are small; the language controls are full books. No token counts appear anywhere in the paper. The remedy is to subsample the controls to matched token counts and plot both measures against training size.
Split asymmetry. The manuscript receives an interleaved recto/verso split, so every test page has a training page from the same leaf. The controls receive a sequential 50/50 split, separating chapters, characters and topical vocabulary. Interleaved splits are systematically easier. This biases every held-out estimate in the direction of the conclusion, and the claim of identical processing across corpora is not met.
Alphabet inventory. Cross-entropy scales with vocabulary size through the smoothing denominator. The transcription table demonstrates the effect: Stolfi's larger inventory raises BPC2 by roughly 0.15 bits and legality rejection elevenfold on the same text. That within-manuscript spread on the legality metric exceeds the manuscript-versus-English gap, so the profile is not invariant across transcriptions in the sense claimed.
OOV incidence. No counts are reported for any condition. The smoothing vocabulary is whatever the training half happened to contain, and therefore varies between corpora, between split directions and between sections, while results are compared to four decimal places. Rare glyphs entering the bucket is also a candidate explanation for the Stolfi layer's elevated rejection rate.
6. The partition averages over the variable of interest
The sectional analysis uses folios 1–66, 67–102, 103–116 and all folios, described as conventional divisions. This is not the conventional illustration-based partition, nor a quire partition, nor a Currier partition. Folios 1–66 contain Herbal A and Herbal B; 67–102 merges astronomical, biological, cosmological and pharmaceutical material and both Currier languages; only 103–116, a single quire, is homogeneous, and it required no decision.
Both mixed blocks average over the same range of variation, so their pooled statistics will resemble each other whether or not the text varies. The reported cross-sectional stability is a property of the aggregation. Adding an all-folios condition, and describing it as showing that the features persist when section boundaries are removed entirely, pools the pooled.
The initial and final positional metrics are computed on exactly the glyph positions Currier's tests are built from — qo- initially, -dy and word-final glyph distribution. So the columns most likely to reveal sectional variation are the ones averaged across the boundary.
The conclusion nonetheless asserts that the A/B differences are surface variation within a single evolving system, that mixed pages contribute modestly to edit-distance overlap, and that they do not affect the entropy values, ΔBPC or the medial result. None of this was measured, and the last two are quantitative claims about an experiment that was not run. Bowern and Lindemann, cited elsewhere in the paper, report h2 of 2.17 for A against 2.01 for B, attribute the difference to qo- and -dy frequency, and find convergence only after deleting those affixes.
Several of the objections above are one design decision surfacing in different places: the block partition, the recto/verso split, the aggregation justified by statistical power, the unmotivated positional decomposition and the untested A/B assertion all follow from averaging over the manuscript's principal axis of variation.
7. Inference and presentation
No confidence intervals, variance estimates or significance tests appear anywhere. Values are reported to four decimal places; for the percentages a single token is roughly 0.01% of the test partition, so the third decimal is below the resolution of one observation. The claim that the cipher and the manuscript are statistically indistinguishable under the diagnostic measures is made without any inferential test and is contradicted by the legality and type-overlap columns. Appealing to statistical power as the justification for coarse aggregation does not cohere in a paper that performs no test. A folio-level bootstrap would settle most of this.
The paper's interpretive claims are placed in table captions rather than in the running text, so the chain from measurement to conclusion is never set out in a form that can be followed in sequence. The claim of a pronounced medial collapse first appears as a caption assertion, with no baseline named, in a table whose own columns show medial exceeding initial.
8. Documentation
The control corpora appear only as bare identifiers in table headers, in two formats, with no titles, authors, dates, genres or lengths, and are not among the files listed as deposited, so the natural-language table cannot be reproduced by executing the released suite. The transcription is Takahashi's, carried as layer H in the Landini–Stolfi Interlinear file; Takahashi is not cited, and the file is cited without version, URL or access date although its contents determine which layers the robustness check can use. The identity of the Latin text matters twice over, since it is also the Naibbe plaintext, and it is never given.
9. What would make the study work
Five experiments, all within reach of the existing code and data.
- Add a copying-process generator to the control set and run the identical pipeline. This is the comparison that could discriminate between mechanisms.
- Partition on Currier's assignments and rerun. Report the initial and final columns separately. Also report metrics quire by quire, letting the manuscript supply its own ordering.
- Subsample the language corpora to matched token counts and plot BPC and ΔBPC against training size. Give the controls an interleaved split at comparable granularity.
- Measure edit-distance overlap as a function of separation between source and target material: adjacent tokens, same line, same page, same quire, distant quires. A stationary encoding predicts a flat curve; a copying process predicts decay. Report distance 0 and distance 1 separately.
- Bootstrap over folios and report intervals; drop the final-position metric or redefine it.
Assessment
The measurement that the paper's abstract foregrounds — that Voynichese word shapes are unusually predictable from adjacent glyphs — has been established since the 1970s and is not in dispute between any competing hypothesis. Self-citation predicts it, a verbose cipher predicts it, a slot grammar predicts it. A property every candidate explanation predicts cannot discriminate among them.
Two results in the paper do discriminate: the sign of the initial-medial difference, and the manuscript's position at the extreme of the legality rejection range alongside its incomplete type overlap. Both indicate that the constrained position is the word boundary and that the manuscript's vocabulary keeps producing forms its own earlier text does not contain. The first is misdescribed and gives the paper its title; the second is explained away with four untested auxiliary hypotheses. As it stands the conclusion is not supported by the design, and the recommendation that the field redirect its attention rests on one comparison with a single artifact reverse-engineered from the object under study.