The Voynich Ninja

Full Version: The Voynich Manuscript’s positional entropy collapse
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3
Happy Thursday, everyone.  Let's try to avoid the thread needing to be locked.  A general plea first to all who are reading - please avoid using LLMs to compose your posts (as opposed to producing a strict translation of something you've already written).  We can often spot this way easier than you think even without the notorious em dashes thanks to the unnatural cadence/rhetorical devices they employ.  It just isn't great for back-and-forth discussion.  It can make people reluctant to engage with you out of worry you're a mere postbox between them and an LLM, and it may also increase the risk of slop.  

Two specific points here:

1. OurYou are not allowed to view links. Register or Login to view. are clear.  Ad hominem comments and flaming are prohibited.  Dunsel, your response to Torsten You are not allowed to view links. Register or Login to view. looks like flaming to me.  There was the option to thank him for spotting any typos and make the point that a private email would have been kinder and more appropriate. You've chosen to escalate this issue in an angry way, and there was no need at all to start swearing.  Nor was there any need to attack the forum.  I do not believe that "all too often" we criticize references rather than papers.  

2. I think we can all agree that typos and similar minor errors are not worthy of discussion here and would be better dealt with via private correspondence.  But at least one of the references that Torsten disputes does not seem to be a typo, and yet it's been defended as a typo.  I find that concerning.  There is a big difference between: 
  • Benett, 1971, "Scientific and engineering problems of the Voynich Manuscript, Proceedings of the American Philosophical Society 115(5):395–412
  • Benett, W. R., Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall 1976.

If Torsten is right that the first is entirely incorrect because there is no mention of the Voynich manuscript in that journal; and the pages cited do not correspond to any article in that journal; and that the correct citation is a section in a book and not a journal...that is going to provoke questions because LLMs are known beyond the Voynich for providing references that are partially or wholly fictitious. 

Like you, Dunsel, I lack Torsten's enviable access to certain journals, so as far as I'm concerned, there could well be an alternative explanation.  e.g. perhaps you made a typo for the volume number and Benett actually published a journal article in pages 395-412 of volume 115(6) as well as the 1976 book.  But I did just spend 30s checking the bibliography in Claire and Luke's free access entropy article, and their citation of Benett is for the book:  
  • Bennett, W. R. (1976). Scientific and engineering problem-solving with the computer, chapter 4. Language, pages 103–198. Prentice Hall series in automatic computation. Prentice Hall, Engle wood Cliffs, N.J.

So I have to ask:  did you use an LLM to provide the bibliography without checking up on it?  If not, can you please quickly explain why the reference appears so wrong, and then we can all move on from this and start discussing your paper's conclusions?
(09-09-2026, 10:47 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.5. You are not allowed to view links. Register or Login to view.. Cited for the claim that "character distributions at word boundaries differ markedly from those in medial positions, with medial transitions exhibiting especially high predictability."

Indeed this article does not support this claim or anything similar, I just read it.
This is what I have to say about the claims the paper makes:

Summary of the claim

The paper computes nine character-level metrics over the manuscript in four transcriber layers, four natural-language texts, one Naibbe ciphertext, and two shuffle controls, and concludes that the manuscript's statistical profile is characteristic of verbose, positionally constrained encoding and is not reproduced by natural language, so research attention should move from linguistic typology to cryptographic mechanism.

The implementation is careful. Exclusion lists, splat-marker handling, OOV bucketing, numeric folio sorting and a public repository are all handled properly, and testing across transcriber layers is the right instinct. The objections below concern what the measurements can support.

1. The positive result is circular

The Naibbe tables were derived with the manuscript's word shapes as the target, and Greshko already reports that the output reproduces flattened glyph frequencies, low entropy, rigid word structure and positional patterning. The statement that no parameters were tuned describes this study's procedure, not the provenance of the parameters, which were fitted upstream. A mechanism reverse-engineered from a target cannot serve as a test of that target; it can demonstrate sufficiency, which the source paper already did.
Calling the cipher a control inverts the term. A control is a case whose value on the outcome is not fixed in advance by its construction. The paper compounds this by declining to advance the cipher as a historical proposal while retaining the inference that the manuscript was likely produced this way, which places the historical plausibility of six tables and per-letter randomization over 38,000 tokens beyond scrutiny while keeping its evidential benefit.
The corresponding omission is decisive. A copying-and-modification process predicts low conditional entropy, small contextual gain and high edit-distance overlap without being fitted to anything, and is the one candidate in the field that could have failed this test. Its absence from the control set, and from the bibliography, leaves the conclusion resting on a single fitted artifact. Shuffle controls are not a substitute: no one proposes them.

2. The paper names the wrong effect

Computing medial minus initial positional entropy per corpus:

Quote:Condition                         Initial  Medial    Med-Init
------------------------------------------------------------
VMS recipe, recto->verso          3.1665   3.3003    +0.134
VMS recipe, verso->recto          3.1657   3.2967    +0.131
VMS herbal (f1-f66)               3.0711   3.3294    +0.258
VMS intermediate (f67-f102)       3.0649   3.3620    +0.297
VMS all folios                    3.1514   3.3556    +0.204
VMS, Friedman layer (F)           3.1719   3.3366    +0.165
VMS, Currier layer ( C)           3.0892   3.3581    +0.269
VMS, Stolfi layer (U)             3.0526   3.3886    +0.336
Naibbe ciphertext                 3.1748   3.3772    +0.202
------------------------------------------------------------
Esperanto                         4.0379   3.9938    -0.044
Latin                             4.0717   3.8966    -0.175
Finnish                           3.8337   3.7874    -0.046
English                           4.0870   3.9636    -0.123
------------------------------------------------------------
Within-token shuffle              6.0139   4.7258    -1.288
Global shuffle                    6.1495   4.7464    -1.403
------------------------------------------------------------

Positional entropy, bits per event. Values from Kinnison (2026);
the third column is my subtraction. Every Voynich condition and
the cipher have medial above initial; every natural language and
both shuffles have medial below.

Every Voynich condition and the cipher have medial above initial; every language and both shuffles have medial below. Relative to its own baseline, what is depressed in Voynichese is the initial position. Initial position also separates more strongly in absolute terms: a gap of roughly 0.8 bits against the language controls, against 0.5 bits at medial.
The paper designates medial position as its interpretive focus on the stated ground that it separates best. That criterion is post hoc, no structural reason for expecting the medial position to be informative is offered, and the selection is incorrect by its own standard. The sign of the initial-medial difference is the one clean discriminator in the results and goes unremarked.

3. Four results contradict the paper's own conclusions
  • Final-position entropy. Claimed to be constrained relative to natural-language controls. Finnish is 2.3588, below every Voynich condition but one, and well below the herbal section (2.6116) and the Stolfi layer (2.5911).
  • Levenshtein ≤1, token-weighted. English 97.0764; VMS all folios 97.0719; recipe (96.20) and herbal (95.85) below English. This column separates nothing. Only the type-weighted column does, and there the manuscript (83–91%) lies between the languages (57–75%) and the cipher (99.84%), not with the cipher.
  • Bigram legality rejection. Naibbe 0.0063 < Finnish 0.037 ≈ Latin 0.039 ≈ English 0.052 < VMS Friedman 0.089 < VMS all folios 0.201 < VMS Currier 0.341 < Esperanto 0.360 < VMS recipe 0.433 < VMS herbal 0.565 < VMS V→R 0.844 < VMS Stolfi 1.024. The manuscript occupies the top of the range and the cipher the bottom, by two orders of magnitude.
  • ΔBPC. Naibbe (0.2876) falls outside the Voynich range (0.2165–0.2774) on the paper's second headline metric.

The pattern is systematic: the cipher matches on every measure dominated by glyph inventory and word shape, which is what its tables were fitted to, and diverges on both measures of lexical productivity, which they were not. That is the signature of a fitted artifact and is evidence against the identification.

4. Metric definitions

Final-position entropy is a smoothing artifact. The event has a single observable outcome, so unsmoothed it contributes zero bits. Under Laplace the reported value is (c+1)/(c+V) transformed, a monotone function of the training count of each token-final character. The metric ranks final-character inventory diversity, not conditional uncertainty. The subsection defending the non-zero values presents the artifact as the consequence of a design choice rather than as the thing requiring justification. It should be dropped, or redefined so both outcomes are observable: given a character anywhere in a token, the probability that the token ends after it.
The positional and n-gram measures are mutually inconsistent. With mean token length near 5, the positional values imply roughly 3.1 bits per event against BPC2 of 2.2763 for the same text and model order. The methods note that smoothing vocabularies differ by event type and never reconcile the two. As reported, positional bits are not comparable to BPC, across event types, or across corpora with different inventories.
Levenshtein ≤1 conflates recurrence with proximity. Distance 0 qualifies, so a text repeating five words scores 100%. English reaches 97% because roughly half its tokens are high-frequency function words. Reporting distance 0 and distance exactly 1 separately would make the column informative; as defined it measures token coverage, which for the language controls tracks morphological richness.
The nine metrics are not nine diagnostics. ΔBPC is the difference of BPC2 and BPC3; the positional measures decompose the same bigram model by event type; the two Levenshtein columns are one computation under two weightings. The table represents roughly three underlying quantities. Presenting them as a battery of independent confirmations overstates the evidential weight, and the figure plotting medial entropy against ΔBPC plots a quantity against a function of itself, so the clustering is largely built in.

5. Uncontrolled parameters

Training-set size. A Laplace-smoothed trigram model on a partition of a few thousand tokens is dominated by smoothing mass, inflating BPC3 and compressing ΔBPC. The manuscript's partitions are small; the language controls are full books. No token counts appear anywhere in the paper. The remedy is to subsample the controls to matched token counts and plot both measures against training size.
Split asymmetry. The manuscript receives an interleaved recto/verso split, so every test page has a training page from the same leaf. The controls receive a sequential 50/50 split, separating chapters, characters and topical vocabulary. Interleaved splits are systematically easier. This biases every held-out estimate in the direction of the conclusion, and the claim of identical processing across corpora is not met.
Alphabet inventory. Cross-entropy scales with vocabulary size through the smoothing denominator. The transcription table demonstrates the effect: Stolfi's larger inventory raises BPC2 by roughly 0.15 bits and legality rejection elevenfold on the same text. That within-manuscript spread on the legality metric exceeds the manuscript-versus-English gap, so the profile is not invariant across transcriptions in the sense claimed.
OOV incidence. No counts are reported for any condition. The smoothing vocabulary is whatever the training half happened to contain, and therefore varies between corpora, between split directions and between sections, while results are compared to four decimal places. Rare glyphs entering the bucket is also a candidate explanation for the Stolfi layer's elevated rejection rate.

6. The partition averages over the variable of interest

The sectional analysis uses folios 1–66, 67–102, 103–116 and all folios, described as conventional divisions. This is not the conventional illustration-based partition, nor a quire partition, nor a Currier partition. Folios 1–66 contain Herbal A and Herbal B; 67–102 merges astronomical, biological, cosmological and pharmaceutical material and both Currier languages; only 103–116, a single quire, is homogeneous, and it required no decision.
Both mixed blocks average over the same range of variation, so their pooled statistics will resemble each other whether or not the text varies. The reported cross-sectional stability is a property of the aggregation. Adding an all-folios condition, and describing it as showing that the features persist when section boundaries are removed entirely, pools the pooled.
The initial and final positional metrics are computed on exactly the glyph positions Currier's tests are built from — qo- initially, -dy and word-final glyph distribution. So the columns most likely to reveal sectional variation are the ones averaged across the boundary.
The conclusion nonetheless asserts that the A/B differences are surface variation within a single evolving system, that mixed pages contribute modestly to edit-distance overlap, and that they do not affect the entropy values, ΔBPC or the medial result. None of this was measured, and the last two are quantitative claims about an experiment that was not run. Bowern and Lindemann, cited elsewhere in the paper, report h2 of 2.17 for A against 2.01 for B, attribute the difference to qo- and -dy frequency, and find convergence only after deleting those affixes.
Several of the objections above are one design decision surfacing in different places: the block partition, the recto/verso split, the aggregation justified by statistical power, the unmotivated positional decomposition and the untested A/B assertion all follow from averaging over the manuscript's principal axis of variation.

7. Inference and presentation

No confidence intervals, variance estimates or significance tests appear anywhere. Values are reported to four decimal places; for the percentages a single token is roughly 0.01% of the test partition, so the third decimal is below the resolution of one observation. The claim that the cipher and the manuscript are statistically indistinguishable under the diagnostic measures is made without any inferential test and is contradicted by the legality and type-overlap columns. Appealing to statistical power as the justification for coarse aggregation does not cohere in a paper that performs no test. A folio-level bootstrap would settle most of this.
The paper's interpretive claims are placed in table captions rather than in the running text, so the chain from measurement to conclusion is never set out in a form that can be followed in sequence. The claim of a pronounced medial collapse first appears as a caption assertion, with no baseline named, in a table whose own columns show medial exceeding initial.

8. Documentation

The control corpora appear only as bare identifiers in table headers, in two formats, with no titles, authors, dates, genres or lengths, and are not among the files listed as deposited, so the natural-language table cannot be reproduced by executing the released suite. The transcription is Takahashi's, carried as layer H in the Landini–Stolfi Interlinear file; Takahashi is not cited, and the file is cited without version, URL or access date although its contents determine which layers the robustness check can use. The identity of the Latin text matters twice over, since it is also the Naibbe plaintext, and it is never given.

9. What would make the study work

Five experiments, all within reach of the existing code and data.
  1. Add a copying-process generator to the control set and run the identical pipeline. This is the comparison that could discriminate between mechanisms.
  2. Partition on Currier's assignments and rerun. Report the initial and final columns separately. Also report metrics quire by quire, letting the manuscript supply its own ordering.
  3. Subsample the language corpora to matched token counts and plot BPC and ΔBPC against training size. Give the controls an interleaved split at comparable granularity.
  4. Measure edit-distance overlap as a function of separation between source and target material: adjacent tokens, same line, same page, same quire, distant quires. A stationary encoding predicts a flat curve; a copying process predicts decay. Report distance 0 and distance 1 separately.
  5. Bootstrap over folios and report intervals; drop the final-position metric or redefine it.

Assessment

The measurement that the paper's abstract foregrounds — that Voynichese word shapes are unusually predictable from adjacent glyphs — has been established since the 1970s and is not in dispute between any competing hypothesis. Self-citation predicts it, a verbose cipher predicts it, a slot grammar predicts it. A property every candidate explanation predicts cannot discriminate among them.
Two results in the paper do discriminate: the sign of the initial-medial difference, and the manuscript's position at the extreme of the legality rejection range alongside its incomplete type overlap. Both indicate that the constrained position is the word boundary and that the manuscript's vocabulary keeps producing forms its own earlier text does not contain. The first is misdescribed and gives the paper its title; the second is explained away with four untested auxiliary hypotheses. As it stands the conclusion is not supported by the design, and the recommendation that the field redirect its attention rests on one comparison with a single artifact reverse-engineered from the object under study.
I was busy with other obligations and made the mistake of not following this thread. 

So if I understand it correctly:

1) Torsten has a direct style that many find unpleasant. His posts also read like they are AI-assisted, which, true or not, makes them difficult to read. 
2) The paper under discussion has big issues with citations, similar to what happens when AI is misused for research.
3) The people that approved of the publication saw nothing wrong with this.
4) Torsten did the work they should have done, resulting in unpleasant ad hominem comments from the author.

While I would, on a personal level, find Torsten's posts easier to read if he used a different style, the more pressing issue seems to be that a paper with potential fake citations passed peer review. Tavi and Nablator seem to confirm this?
(10-09-2026, 02:40 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.This is what I have to say about the claims the paper makes

I'm sorry, but if this post was your first here, imo it would be instantly flagged as an AI generated/assisted post. I completely understand if you are using these tools for translation, or phrasing in some cases, but this goes so much further beyond that. The entire structure of the text reads exactly like an LLM response, even down to which words are put in bold.
(10-09-2026, 02:40 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view....

Serious question, because I'm curious about your posts and articles and not implying any wrongdoing: how do you manage to get such a high AI score (90-100%) with You are not allowed to view links. Register or Login to view. or any other reliable AI detector such as GPTZero? Are they entirely generated or rewritten by AI or just translated? I never managed to get anything detected remotely close to 100% AI (it's usually 0% or close to 0%) when it's just translated by AI.

This is since your article "The challenge of Analyzing a Dynamic Text", not in your earlier posts and articles.
(09-09-2026, 08:46 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Unfortunately, with my first attempt at peer review, I didn't preprint this paper.  However, it does a comparison to English, Finnish, Esperanto, Latin, Naibbe cypher and shuffled controls.  It basically uses a train/test method where it will train on, for example, odd pages and then test on even pages or train one Voynich section (i.e. recipes) and test on another (i.e. herbal).

Considering all that has been determined about the structure of the Voynichese "words", it should be undisputed now that, IF Voynichese is conjectured to be some roughly phonetic encoding of a natural language -- like an original orthography, or a simple substitution cipher of the language's standard orthography -- then the language must be a monosyllabic one.  

Then why do people still compare it only against polysyllabic languages?  Angry

And, for the millionth time: character-based statistics -- such as digraph frequencies and correlations, and per-character entropy -- are properties of the spelling and encryption system, NOT of the underlying language.  Any claim that "character statistic X is too low|high for a natural language" is nonsense.

All the best, --stolfi
Quote:the more pressing issue seems to be that a paper with potential fake citations passed peer review

I think it's not our problem, it's Cryptologia problem. They actually belong to some mega badass publishing corporation with a lot of resources. They should care of it.

I also agree that Torsten commentary feels like generated by AI. I think it is somehow awkward to nitpick at someone's mistakes probably coming from AI usage and then create your acussation using AI as well which results in some heavy text hard to swallow.

I really would like if we discussed the actual content and not citations. And we did it in our simple, human words.
By the way the article is behind paywall so 80% of people (like me) are here discussing stuff that they haven't seen.
Thread locked. Dunsel will be contacted privately.
Torsten, thank you for bringing potential AI abuse to general attention. That said, it would be nice if you could be a bit clearer about your own (alleged) AI usage as well.
I have received a lot of correspondence around this thread, and will add two more things:

* Torsten did not necessarily want to point out problematic AI usage - this inference was entirely my own.
* Lisa adds that "citation accuracy is the author's job entirely and exclusively. Journals and editors don't check them...no one has that kind of time." I called this the reviewers' job earlier, but I agree that it is the author who should take this responsibility upon themself. Clearly this is where it went wrong.
Pages: 1 2 3