The Voynich Ninja

Full Version: Why and how the text could be Bavarian
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36
Why this finding matters for the VMS and the VBM

The following comparison is not just one more statistical similarity between the Voynich Manuscript and a historical German text. It shows something much more revealing: a double structure.

The inventory of Voynich cores behaves like the consonant inventory of a natural language. And, interestingly enough, medieval Bavarian gives the best fit here.
The frequent part of the distribution, however, is compressed exactly the way one would expect from a cipher system built around the VMS families. Once those families are structurally unpacked, the distribution once again lands astonishingly close to: Bavarian.

So the VMS cores show both at the same time: a language-like long tail and a cipher-like compressed head. That much is not exactly news.

The interesting part is that this apparently contradictory combination can be explained rather elegantly by the Vowel Bridge Model (VBM), and it only takes a very small number of additional assumptions.

I chose the Stars section as the comparison base because it contains running text and is structurally a better match for my comparison corpus. It yields 7,081 cores, which I compared with equally sized samples from a corpus of 14 Bavarian texts from the 15th century. The Bavarian corpus comes from the Reference Corpus of Early New High German and basically preserves the historical dialect spelling.

You are not allowed to view links. Register or Login to view.
The first comparison gives:

Code:
Measure                          VMS        Bavarian
Distinct types                  1,095      1,123  (+/- 22 )
Hapax types                      709        667   (+/- 23 )
Share of most frequent type     11.65%     5.84   (+/- 0.27%)
Types for 50% coverage          15         33.6  (+/- 1.2)
Types for 80% coverage          135         190   (+/- 5.7)
Types for 90% coverage          387         436   (+/- 13.4)
± denotes one sample standard deviation across the 1,000 random samples.

The first two figures show the language-like side of the result. With the same amount of material, the VMS has essentially the same number of distinct core forms as the Bavarian corpus. The number of forms that occur only once - hapaxes - is also in the same range.

Type count and hapax count are not independent. Both react to rare forms. But taken together they show that the VMS does not have a tiny core inventory endlessly recycling the same handful of pieces. Behind the tightly regulated surface sits an extraordinarily broad reservoir of rare structures.

And that gives us the first important result.

Put differently:
The VMS has enough distinct and rare cores to represent a natural, historically variable consonant stream. That flatly kills again the claim that the VMS cannot contain language!

I compared these consonant clusters with Latin as well, using the same kind of structured pipeline (functional-word filtering, ending strip, geminate collapse) rather than raw text. Even with this fair treatment, Latin falls far short. Even without any deeper analysis, it became obvious that the Romance languages are distributed much more densely, simply because they contain more vowels and fewer consonant clusters. Which is actually pretty logical.

I compared it with Czech as well, a consonant-rich language, and that does not fit either. There is one caveat: I had no reliable way to determine what the above-mentioned families would correspond to in Czech. So that test has to be treated with caution!

I also compared it with Alemannic. That fits much better again, although not quite as well as the Bavarian text. Swabian also fits well, but worse than Alemannic. This suggests that a Germanic core structure works better overall.

The apparent contradiction - and a key to the cipher

The head of the distribution, however, does not fit at all at first. This is actually a well-known Voynich problem.

In Bavarian, the most frequent cluster accounts for about 5.84% of all occurrences. In the VMS, the most frequent reduced core accounts for 11.65% - almost twice as much.

To cover half of all Bavarian material, about 34 different types are needed. In the VMS, only 15 are needed at first.

So we get a strange picture:

The rare structures look like a large natural-language inventory.
The frequent part looks as if many different forms had been collapsed into a small number of strings.

This contradiction is exactly what makes the VBM interesting. The model does NOT claim that the visible Voynich cores are untouched plaintext clusters. It claims that they are organised, and partly compressed, by vowel bridges, boundary glyphs and families.

Independent checks on consonant clusters at word endings and word beginnings in the plaintext corpus also show that the core positions are strongly conditioned by classes and families. They do not combine like free letter strings.

Where the oversized head comes from

The three most frequent reduced cores are:

k: 825 occurrences
d: 780 occurrences
t: 490 occurrences

Together they make up roughly 30% of all cores.

What matters here is that these forms almost never occur in isolation. They are nearly always(!) - tied to a right-hand family state:

Code:
Core    Bound to a family or to y
k              98.4%
d              97.9%
t              98.4%

In the original analysis, the right-hand families were removed. As a result, forms such as k+aiin, k+ain, k+ar, k+al, k+y were all counted as plain k! That does not mean that all of these forms are identical in the cipher, or that they necessarily belong together. It only means that our first counting method threw them into the same drawer.

So I stopped doing that. Instead of counting every one of those cases as the same naked k, d or t, I counted the core together with the family context in which it actually occurs: k|aiin, k|ain, k|ar, k|y, d|aiin, d|ar, d|y, and so on.

That raised the obvious question: what if these glyphs are not supposed to stand apart from the families at all? 
What if k+aiin is a different "k-state" from k+ain?

Given the extremely strong binding, and given what a cipher of this kind would look like, that is not a wild assumption. It is the obvious thing to test.

Important: nothing is altered here. Nothing is added. This is information that is actually present in the manuscript, and the strength of the binding confirms that it matters. The information was simply lost when the families were stripped and forgotten.

Once different k, d and t "states" are distinguished by their family context, the distribution suddenly changes:

Code:
Measure      VMS before    After separating k/d/t states
Top-1          11.65%             5.82%
Cov50              15                29
Cov80            135                164
Cov90            387                419

The most frequent VMS type now stands at 5.82%, almost exactly on the Bavarian value of 5.84%.

That is the first real "aha" moment:

The Voynich peak, which was almost twice as high, nearly disappears once the already visible family states are no longer artificially collapsed into single categories called k, d and t.

So the sharp head was not necessarily a property of the underlying language. It was, at least to a large extent, the result of family compression - or, more precisely, of our first core count being too crude.

The E-family explains the broad shoulder of the distribution

After separating k, d and t, the extreme peak is already explained. But between roughly rank 10 and rank 70, the Voynich distribution is still more concentrated than the Bavarian one.

And that is exactly where the E-family sits.

In the strict Stars definition - one anchor atom followed by one of ten fixed kernels (e, ee, eee, eo, eeo, ed, eed, eod, eeod, eeed) - it contains 72 distinct types and 1,599 occurrences. These include forms such as k.ee, k.eed, k.ed, k.e, t.ee, t.ed, ch.e, ch.ed.

These 72 types are not evenly distributed. Just 19 E-types among the first 70 ranks contain 89.4% of the entire E-family mass.

So the E-family has two zones: a small, very frequent inner group, and a large, thinly populated fringe of rare combinations.

That explains its peculiar effect on the overall distribution. The frequent inner group forms a broad shoulder in the head. The rare fringe contributes to the long hapax tail at the same time.

The E-forms are also tied to their right-hand context, though not uniformly: forms ending in -ed and -eed are followed by y in 81.2% and 88.9% of their occurrences respectively, while -eod is followed by y in only 70.0% of cases - predominantly, but not "almost always" across the board.

So the E-family is not decorative free variation. It is a regulated matrix of connection states.

I then applied the same principle to the other side. For the E-family, I distinguished only the two directly observable states "before y" and "not before y".

The result is:

Code:
Measure      VMS with family states      Bavarian
Types                1,157            1,123    (+/- 22
Hapaxes                721              667    (+/- 23
Top-1                   5.82%            5.84  (+/- 0.27%
Cov50                   34               33.6   (+/- 1.2
Cov80                  182              190    (+/- 5.7
Cov90                  449              436    (+/- 13.4
± denotes one sample standard deviation across the 1,000 random samples.

Now it is no longer just one or two figures lining up:

The most frequent type has almost exactly the same share.
Almost exactly the same number of types is needed to cover 50% of all occurrences.
Practically the same number of types is needed to cover 90%.
The total type count remains in the same range.
The long hapax tail survives.

Across the first 70 ranks, the mean deviation of the cumulative distribution drops from 12.54 to 1.83 percentage points. That is a reduction of roughly 85%.

What is actually unusual here

The strong result is not simply: "Bavarian and Voynich have roughly the same number of types."

The stronger result is this: the VBM could explain at the same time why one part of the distribution already fits Bavarian from the start, and why another part does not fit at first.

According to the model, the rare tail comes from the variety of the underlying consonant clusters. The compressed head is created because frequent cluster states are organised into families and then collapsed by an overly crude analysis.

So the model does not merely explain a match. It explains the exact shape of the mismatch.

Summary

Methodologically, this matters much more than a lucky curve match. A similarity found after the fact can always be dismissed as coincidence. But here the error sits exactly where the independently observed Voynich families dominate:

k, d and t create the extreme peak; the E-family shapes the range up to roughly rank 70; the remaining rare combinations preserve the long tail.

You could put it like this: the apparent failure of the model turns out to be the fingerprint of the family mechanism - and therefore of the cipher.

Or even more simply: under the visible Voynich distribution lies a language-like distribution. The families squeeze the frequent part without destroying the rare tail.

What the test tells us about the cipher

This result also changes how the families themselves have to be interpreted.

They are apparently not just disposable endings that can be cut away and forgotten. They probably carry information about which exact cluster state is meant.

The relevant unit would therefore not simply be k, but something more like: k|aiin, k|ar, k|y k.ee|y, k.ee|NON-y.

That means a later decipherment cannot look only at naked cores. Core, family and right-hand connection have to be treated as one combined cipher unit, potentially carrying different values.

So the statistical comparison does not only provide an argument for natural language. It also gives a concrete clue as to how the cipher organises its most frequent consonant structures.

A cautious conclusion

This result proves neither a specific plaintext reading nor Bavarian beyond doubt. Alemannic texts also came reasonably close in the previous controls, so the more cautious linguistic label is "historical Upper German".

Romance and Latin comparison texts have much shorter cluster inventories; Old Czech lies between the Romance and Upper German values.

The family test was also developed on the same Stars section in which the deviation was observed. So this is not yet an independent blind test.

Of course splitting frequent types makes any distribution flatter. That is not the point.

The point is that the particular split demanded by the visible family structure makes the head, the tail, the type count and the hapax count line up with Bavarian at the same time.

That is the weird part.

Despite all the caveats, what I have shown here is substantial:

At core level, the VMS has the diversity of a natural Upper German consonant stream. Its strange surface distribution can largely be explained by an independently visible family system which, quite typically for a cipher, reveals a kind of compression mechanism.
Are benched gallows cth ckh cph cfh and oddballs like cthh other etc… all rolled up to t and k?

In my mind there’s always been the possibility of a column row based table or set of tables hiding away either syllables or I suppose consonant clusters of 1-3 characters with a separate o table particularly designed for labels or numbers.

All that to say I like the idea of requiring two components working together to decipher a plaintext.
(11-07-2026, 03:12 PM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.At core level, the VMS has the diversity of a natural Upper German consonant stream. Its strange surface distribution can largely be explained by an independently visible family system which, quite typically for a cipher, reveals a kind of compression mechanism.

Please don't take this as an attempt to deride your approach, I admire the effort and the persistence. However, I have found an interesting parallel between how I see your theory and Stolfi's Chinese theory. The more effort you put into defending your theories the less convincing they look to me. I've tried to understand why, and now I realize it's quite simple. It's absolutely clear that both you and Prof. Stolfi are very convinced in the merits of your respective theories and both of you are highly motivated to look for evidence supporting the respective theories. So with every new post that doesn't actually contain any good direct evidence, it seems more and more likely to me that no definitive evidence can be found at all. It's as if you were looking for gold in a particular place and meticulously kept excavating and panning for gold grains and with each new shiny speckle that you find and that turns out not to be gold the observers will inevitably become less and less credulous of the whole mother lode claim.
(11-07-2026, 04:00 PM)Grove Wrote: You are not allowed to view links. Register or Login to view.Are benched gallows cth ckh cph cfh and oddballs like cthh other etc… all rolled up to t and k?

In my mind there’s always been the possibility of a column row based table or set of tables hiding away either syllables or I suppose consonant clusters of 1-3 characters with a separate o table particularly designed for labels or numbers.

All that to say I like the idea of requiring two components working together to decipher a plaintext.

No. They are not collapsed into k or t.

In this analysis, the following are treated as separate atomic / one glyphs: cth, ckh, cph, cfh, as well as ch, sh, and
qo. 

And yes, that's basically how the cipher must work—the research I've done here makes that clear. However, I don't see any evidence of numbers, at least not yet Wink
Hey Jojo, does your theory sitll work when the Bavarians turn into something close to the Stuafer?

Hey Jojo, does your theory sitll work when the Bavarians turn into something close to the Staufer?
(11-07-2026, 05:32 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.
(11-07-2026, 03:12 PM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.At core level, the VMS has the diversity of a natural Upper German consonant stream. Its strange surface distribution can largely be explained by an independently visible family system which, quite typically for a cipher, reveals a kind of compression mechanism.
Please don't take this as an attempt to deride your approach, I admire the effort and the persistence. However, I have found an interesting parallel between how I see your theory and Stolfi's Chinese theory. The more effort you put into defending your theories the less convincing they look to me. I've tried to understand why, and now I realize it's quite simple. It's absolutely clear that both you and Prof. Stolfi are very convinced in the merits of your respective theories and both of you are highly motivated to look for evidence supporting the respective theories. So with every new post that doesn't actually contain any good direct evidence, it seems more and more likely to me that no definitive evidence can be found at all. It's as if you were looking for gold in a particular place and meticulously kept excavating and panning for gold grains and with each new shiny speckle that you find and that turns out not to be gold the observers will inevitably become less and less credulous of the whole mother lode claim.

Yes, that assumption might arise, because Voynich theories often lose their persuasiveness the more flexible they become.

But there’s a significant difference! What I’m doing here is not adding arbitrary assumptions merely to rescue the model, because otherwise the basic system simply wouldn’t work!

The system already works at the most basic level, because it can explain a significant portion of the VMS structures with just two assumptions. According to Ockham’s Razor, that’s already a very good foundation.

But the system is clearly not just that simple—that’s only the first step. Quite obviously, the cores are encrypted and not in plaintext; otherwise, the cipher would have been deciphered 100 years ago.

What I’m doing now—and this is where you seem to lose interest—is approaching this cipher from various angles of statistical analysis. This is the classic approach in cryptology. There’s no other way. How else could one understand exactly how a cipher works?

And unlike most other theories, this approach works very well and simply with the VBM.

The last example I mentioned demonstrates this in a striking way—through this approach, the cipher has exposed itself. 

Not understanding that is a mystery to me... somehow - but okay...
(11-07-2026, 06:08 PM)sōlstəs Wrote: You are not allowed to view links. Register or Login to view.
Hey Jojo, does your theory sitll work when the Bavarians turn into something close to the Staufer?

If you mean the “Staufer,” I assume you're referring to the Swabian-Alemannic language area, because otherwise they wouldn't fit into that time period.

As mentioned, other Germanic languages are also possible; they fit too. But Bavarian seems to be the best fit, at least among the ones I've tested. Probably because it’s one of the most abbreviated Germanic dialects.
I have moved the You are not allowed to view links. Register or Login to view. to oshfdk's reference to the Chinese Theory to the Chinese Theory thread, since this attracted a further reply, and we'll get digressed from JoJo's Bavarian claims.
To counter the impression that I’m “watering down” my theory with all these posts, I’d like to briefly clarify the structure of my posts. 

At first, there was really only a single basic assumption or rule:

Rule 1: Vowels are represented as bigrams, separated in the middle by word boundaries.

Rule 2: The Consonant-nuclei are encrypted.

The second rule is actually just a conclusion: Since the remaining consonant clusters cannot be plain text or a simple substitution due to the family structures, they must be encrypted.

With this alone, I can explain the various oddities of the VMS:

1. Using the VBM, I can explain why frequent word repetitions are simply a linguistic consequence of the VBM.

e.g.:
<f75r.13,+P0> pchedy.keedy.qokedy.<->qokedy.qokedy.qokedy.qokain.olshedy

You are not allowed to view links. Register or Login to view.
 
<f8v.8,+P0> okchol k sh.<->chol.chol.chol.cthaiin.dain

You are not allowed to view links. Register or Login to view.

2. These rules of VBM immediately give rise to some of the structures of VMS

- The uniform, almost static word distribution
- The limited word lengths
- The almost repetitive occurrences of similar words
- Words that differ by only one letter
- Repetitive phrases
- A structured nature of such text
- And the many hapax, which contradict any normal linguistic structure

You are not allowed to view links. Register or Login to view.

3. The 7 rules for word boundaries become a principle of the VBM

This is because the VBM essentially reduces these rules to a simpler principle: word boundaries are separations of vowel bigrams, VL.VR, where the left side consists of only a small number of glyphs.
(This rule is later expanded; see below.)

4. The VBM produces significantly low entropy at the glyph level
If the extremely frequent vowels are encoded as recurring bigrams, lower glyph-level entropy is exactly what should be expected.

5. The VBM can very easily explain this strange long-range effect in the VMS.

Why does “shedy” follow “qo” 49% of the time and “tedy” only 25% of the time? How can a glyph in a normal language have a long-range effect extending beyond three glyphs and a space?

This then becomes a perfectly normal phenomenon in a natural language:

You are not allowed to view links. Register or Login to view.

6. The strukture of the chipher

So I added some more rules in You are not allowed to view links. Register or Login to view. , because, of course, it can’t be that simple. These rules concern the families and the cores.

Rule 3: s="und" (and)—this resulted from frequency analyses, but isn’t actually that important, it change the counts only a little.

Rule 4: The  aiiin, aiin, ain, aiir, air, am, ar, or, al, ol families were (on a trial basis) treated as suffixes. To better understand the roots, I omitted these families in the VMS and removed these suffixes from the Bairish text.

Result: The distribution of types and happaxes suddenly resembled that of the Bairish corpus. But the distribution did not.

This led to two additional rules:

Rule 5: The glyphs “k,” “t,” and “d” are markers defined by the families that follow them.
That is to say: in “kaiin,” “k” is something different than in “kain” or “keed.” So one must distinguish between the different “k”s. The same applies to “d” and “t.”

This is not arbitrary. These three glyphs are almost completely tied to families or y:
k: 98.4%; d: 97.9%; t: 98.4%

Once these states are separated, the most frequent VMS type drops from 11.65% to 5.82%, which matches the Bavarian value: 5.84%.

That is not a minor cosmetic improvement.

Rule 6 And then came Rule 6, which isn’t really a rule in its own right, but merely an extension of Rule 5: The glyphs that follow the E family—that is, y or non-Y—must also be separated accordingly. This is logical and had nothing to do with trial and error. The cipher encodes multiple parts as a single unit.

And lo and behold, suddenly everything fits perfectly:

Code:
Measure      VMS with family states   Bavarian
Types                1,157            1,123   
Hapaxes                721              667   
Top-1                  5.82%            5.84 % 
Cov50                  34              33.6 
Cov80                  182              190   
Cov90                  449              436  

The highly peculiar distribution of glyphs—which supposedly doesn’t fit any language—fits perfectly with a language solely on the basis of these few assumptions. And not through completely outlandish rule constructions, but through very simple assumptions that arise  from the logic of the VMS structure and were NOT simply made up!!!

(I made the post You are not allowed to view links. Register or Login to view.so detailed in order to accurately present the background.)

So we’re dealing here with a basic rule and very few assumptions that can suddenly explain the most absurd anomalies of the VMS in a relatively simple and comprehensible way.

Now you have to pull out your calculator and think about how likely it is that such a simple basic rule could explain so many peculiarities of the VMS—and, with just a few additional assumptions, this head-and-tail distribution. Not because it’s proof, but because a relatively compact underlying architecture forces more and more previously separate anomalies into a single common mechanism. Taken individually, each one could be a coincidence, but together...

So I really don’t understand why my posts—and especially #321—are supposed to lead to a “watering down” of the discussion; it’s the exact opposite. It simply explains the underlying cipher and demonstrates, using the Bavarian corpus, that it can indeed work this way (which does not definitively prove that it is Bavarian, but only that VMS under the VBM could be a language.)

Let’s ask it the other way around: What other theory can achieve all this with so few assumptions?
(12-07-2026, 06:40 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.Let’s ask it the other way around: What other theory can achieve all this with so few assumptions?
If you number the letters with Roman numerals and turn the words into anagrams in order from largest to smallest, the result was similar to the VMS text (there will be doublets, triplets, and "words" that are similar to each other). If you choose the alphabet correctly, I believe you can get as close to Voynichese as possible on your first attempt.
This doesn't apply to your theory, but in my opinion, it's easier to do...
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36