The Voynich Ninja

Full Version: The Challenge of Analyzing a Dynamic Text
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
The Challenge of Analyzing a Dynamic Text: Why the Voynich Manuscript Resists Conventional Interpretation

You are not allowed to view links. Register or Login to view. has a new paper by Torsten Timm

Here is the abstract:
Quote:The Voynich Manuscript (MS 408, Beinecke Library, Yale University) has resisted all attempts of conventional interpretation for over a century. This paper argues that the persistent failure of cryptographic, linguistic, and statistical approaches stems from a shared foundational assumption: that the manuscript’s text was produced by a static system—whether a fixed cipher, a natural language grammar, or a stable encoding scheme. Analysis of the manuscript’s word network and vocabulary evolution reveals that this assumption is untenable. 
The text exhibits continuous development throughout the manuscript, with vocabulary, character usage, and word frequency distributions shifting gradually from beginning to end. This dynamic character is evidenced by a single, highly connected network of similar words spanning the entire manuscript, a strong correlation between word frequency and number of similar word forms and asymmetric vocabulary distribution indicating directional evolution. These properties are naturally explained by the self-citation hypothesis, which proposes that the text was generated through iterative copying and modification of previously written words—a process confirmed as the intuitive human strategy for producing meaningless text beyond approximately 100 words (Bowern & Lindemann, 2021; Gaskell & Bowern, 2022). The dynamic perspective reconciles apparently contradictory observations and provides a framework for understanding why the manuscript has resisted every analytical approach premised on static rule systems.
Hey, this is a very interesting read. It seems to be able to explain a lot. I wonder what more testable implications this "self-citation hypothesis" gives rise to.
In many places you refer to [13] by Timm and Schinner. Is that paper publicly available somewhere?
I think the Timm & schinner paper is paywalled but the github is not.

Paper: A possible generating algorithm of the Voynich manuscript Torsten Timm & Andreas Schinner Pages 1-19 | Published online: 25 May 2019
You are not allowed to view links. Register or Login to view.

Ninja thread with link to paper and github /additional materials
The Voynich Ninja › Voynich Research › News > [Article] A possible generating algorithm of the Voynich manuscript
You are not allowed to view links. Register or Login to view.

Discussion thread about the paper
The Voynich Ninja › Voynich Research › Analysis of the text > Discussion of "A possible generating algorithm of the Voynich manuscript"
You are not allowed to view links. Register or Login to view.
My paper "The Challenge of Analyzing a Dynamic Text: Why the Voynich Manuscript Resists Conventional Interpretation" was published at Cryptologia:

You are not allowed to view links. Register or Login to view.
You’re assuming a static system.

Four points on this:

1. What if the cipher evolved as it was being written? According to my model, that’s entirely possible, because during the encryption process, the cipher might have revealed that certain sections didn’t fit together. This was gradually “perfected.” Actually, that's exactly what you'd expect from a 200-page book by Chiffre...

2. If we assume that it’s a homophone cipher with free choice, we know that people develop habits simply based on what they prefer to write, and so a homophone cipher can shift those habits.

Regarding “m”—example: Imagine someone has been writing letters for months. At the end of a line, according to the rules, they may either write out “und” or use the abbreviation like “&”—both are allowed and have equal value. At first, they almost always write out “und,” but over some month, he increasingly use “&.”

3. The theory of humors was the organizing principle of 15th-century medicine. Herbal texts of the time classified plants according to warm/cold/moist/dry). An author writing such a work works his way through an ordered scale. And an ordered scale in the content produces, on the surface, exactly what your cosine-similarity curve shows: not a thematic block with a hard boundary, but smooth transitions. The “continuous gradient without a discontinuity” would then not be an evolving system, but simply a table of contents arranged from cold to warm or something similar.

4. Spelling Variations in Authentic Manuscripts:

Spelling in medieval texts is not uniform (I call this ‘spelling anarchy’). The same word can appear in several variants within a single manuscript (the best-known example being the German ‘und’: “und, un, unde, unnd,” etc.), and the preferred form often changes over the course of longer works. I have examined this using the “Breslauer Arzneibuch” and the “Admont Bartholomew Codex”: In the Admont Codex, for example, the “si/sy” pair gradually shifts from about 60% “si” at the beginning to about 10% at the end. Other pairs change abruptly, with several pairs changing at the same points in the text—which suggests a change in scribe or source text rather than a matter of habit; and that, too, could apply to the VMS.

However, I cannot completely rule out transcription effects in the editions I have used. Just a note.

But the fundamental phenomenon is beyond doubt, and I have observed it myself in shorter manuscripts: The surface spelling shifts, gradually or abruptly, while the underlying language remains perfectly stable German.

So if a gradual superficial change were evidence against a stable system, this criterion would classify authentic 15th-century German manuscripts as lacking a system. (Just kidding Wink )


If you take all four of these factors together, the observed evolution could just as easily arise naturally.
(10-09-2026, 06:00 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view....

You describe four systems that vary, but each has a stable core: a cipher that exists before being refined, a homophone table with shifting preferences, a fixed humoral ordering, a stable German with spelling variation. Remove the variation and the system remains.

The Voynich text is different. The change is not variation around a stable system — it is a directional shift. Word families like ⟨chol⟩/⟨chor⟩ dominate the early sections but decline in frequency as new families like ⟨chedy⟩/⟨shedy⟩ rise, with documented intermediate forms along the way. New word families like ⟨qokedy⟩/⟨qokeedy⟩ emerge that have no counterpart earlier in the manuscript. Even new glyphs appear: ⟨m⟩ is absent from early herbal folios and becomes frequent later. That is not substitution preferences shifting over a stable vocabulary — it is the vocabulary and the writing system themselves expanding in one direction.

This predicts a testable difference. In each of your four models, normalizing the variation should recover the stable core — the cipher table, the plaintext language, the humoral categories. For the Voynich, nobody has succeeded in doing this, despite decades of attempts. The self-citation model explains why: there is nothing stable to recover.
You chose "m" as your example of the writing system itself expanding. That is exactly the case I have examined in detail.

"m" does the same thing as a scribe who may end a line either with "und" written out in full or with the abbreviation "&" — same value, free choice. If you count what stays constant, you get: the number of lines ending on that word does not change at all. Only the preference between the two spellings shifts. Right?

And that is exactly how "m" behaves. You cite it as a new glyph extending the system. But on folios without "m", nothing is missing. Two measurements, one pipeline: First, the broad class of line closings {n, r, l, m} stays at ~50% whether or not a folio uses "m" at all (permutation p=0.75) — the slot itself is conserved. 

Second, where am-family closings (am, dam, okam) are absent, aiin-family closings are correspondingly more frequent — within Currier A alone p=0.006, within Currier B alone p=0.0002, so this is not just an A-versus-B difference. The gap does not spread across the rest of the alphabet, as copying from local templates would predict — if I understand it correctly. But no: it shows up in the exchange partner. And the swap is no one-letter edit — though "am" does look a little more like "aiin" than "&" looks like "und". Still: two differently shaped spellings serving the same closing function — a fixed slot, a wandering preference, measured on your own example. 

"m" may simply be an abbreviation sign, like the closing flourishes of medieval scribes — Stolfi suggested exactly this relation: m as an abbreviation for iin-groups, particularly at line end, with dam/am as the counterparts of daiin/aiin.

(Along your chedy gradient: class share flat, ρ=−0.06; m-share within the class rising, ρ=+0.27, p=0.0005. Line-internal token endings show no such difference (control p=0.38) — the effect sits exactly at line-final position. Folios ≥8 lines; the picture is robust at page level ≥5 lines as well.)

On directional change: a cipher table that can be extended during use can produce exactly this signature. New entries can only be used after they have been added — can't they? So new families logically appear only later, seemingly "out of nowhere". And if old entries are kept — so as not to endanger decipherability — early families persist. That is your asymmetry, word for word: your own temporal-arrow argument ("variants can only appear after their source exists") applies to a growing table just as much as to copying. Direction alone cannot distinguish between these two cases! Sorry.

And part of the direction need not be evolution at all, but content. Any manuscript whose subject matter is arranged systematically produces a vocabulary gradient by itself: terminology tied to one part of the subject declines while terminology tied to the next rises, with a mixed transition zone in between. A table of contents arranged along a scale looks, from the outside, exactly like "evolution" — except that it isn't.
(What distinguishes the two models in principle: under a growing table, the structural rules must hold for old and new families alike. And they do: adjacent gallows remain practically nonexistent in all four quartiles of your gradient, and "qo" stays ~99% token-initial throughout. Self-citation does not by itself explain why the strongest combinatorial constraints remain invariant while the vocabulary distribution drifts.)

On your test: normalizing the Breslau spellings works for exactly one reason — we know German (the two of us especially Wink). We recognize "vnder", "under", "unnd", "un" as the same word because we know the word. For an unsolved cipher, precisely this recognition is missing. The failure to recover plaintext is therefore predicted either way — whether the text is self-citation or an unsolved homophonic cipher. A test that comes out the same under both hypotheses decides nothing.

What is possible without the key is distributional normalization — finding glyphs that behave as one class. The m/aiin relation above is exactly that: a recovered distributional equivalence class within the stable core, with p-values. Your model has no reason to expect such conservation classes. Mine predicts that more of them should exist. The m/aiin relationship is one such case.