Why this finding matters for the VMS and the VBM
The following comparison is not just one more statistical similarity between the Voynich Manuscript and a historical German text. It shows something much more revealing: a double structure.
The inventory of Voynich cores behaves like the consonant inventory of a natural language. And, interestingly enough, medieval Bavarian gives the best fit here.
The frequent part of the distribution, however, is compressed exactly the way one would expect from a cipher system built around the VMS families. Once those families are structurally unpacked, the distribution once again lands astonishingly close to: Bavarian.
So the VMS cores show both at the same time: a language-like long tail and a cipher-like compressed head. That much is not exactly news.
The interesting part is that this apparently contradictory combination can be explained rather elegantly by the Vowel Bridge Model (VBM), and it only takes a very small number of additional assumptions.
I chose the Stars section as the comparison base because it contains running text and is structurally a better match for my comparison corpus. It yields 7,081 cores, which I compared with equally sized samples from a corpus of 14 Bavarian texts from the 15th century. The Bavarian corpus comes from the Reference Corpus of Early New High German and basically preserves the historical dialect spelling.
VMS: Stars section (folios f103-f116). This section contains 49,788 atoms in 1,085 lines and 10,836 tokens.
ch, sh, qo, ckh, cth, cph and cfh were each treated as one unit.
Bavarian: a corpus of 14 North, Central and South Bavarian texts from the 15th century, taken from the Reference Corpus of Early New High German, ReF v1.0.2. It is an abbreviation-expanded version in which the historical dialect spelling is basically preserved.
The VMS pipeline.
1. Tokens are reconstructed per line from the atom stream (without spaces).
2. Family suffix list (checked longest match first): aiiin, aiin, ain, aiir, air, am, ar, or, al, ol.
3. A token is dropped entirely if: it is the last token of its line; it consists of exactly one atom "s" (Assumption: “s” = “and” (Tironian 7); or stripping a family suffix from it leaves nothing (a bare family word).
4. For every remaining token, in order: (a) if it is not the first token of its line, the token-initial bridge glyph is removed, together with any e-run directly attached to it (e.g. "qo", "che", "chee"); (b) on what remains, a family suffix is stripped if one matches; otherwise, if the token is not line-final and the following token was not dropped, the single token-final bridge glyph is stripped instead.
5. What remains is the core. If nothing remains, the token contributes an empty core (1,760 cases) and is not counted as a type.
6. After this decomposition, any resulting core that consists of exactly the single atom "s" is treated as the functional counterpart of "und" and removed from the core count (119 cases). This step is applied only after family/bridge stripping, since the "s" in question is what is left once the actual family or bridge glyph has already been correctly identified and removed.
Result: 7,081 non-empty cores, 1,095 distinct types, 709 hapax types.
Bavarian pipeline.
1. Words matching und, vnd, unde, vnde, vnnd are removed (functional counterpart of VMS "s").
2. Each remaining word is normalized, in order: ss -> s; sz -> s; v before a vowel -> f, remaining v -> u; j -> i; tz/cz -> z; ck/kh -> k; c before a,o,u -> k, c before e,i -> z; b -> p; t -> d; consecutive repeated letters collapsed to one; sch and ch each treated as one indivisible unit.
3. Ending strip, applied per word before concatenation: if the word is at least 3 characters long and ends in a vowel followed by n or r, the last two characters are removed - to remove the standard German endings (functional counterpart of the VMS family endings).
4. All processed words are concatenated with no separator, then segmented on every maximal run of non-vowel characters; each such run is one comparison type.
Method: 1,000 random samples of exactly 7,081 segments were drawn from the full Bavarian segment pool, and Types / Hapax / Top-1% / Cov50 / Cov80 / Cov90 computed for each. The tables report mean and standard deviation across the 1,000 samples.
The first comparison gives:
Code:
Measure VMS Bavarian
Distinct types 1,095 1,123 (+/- 22 )
Hapax types 709 667 (+/- 23 )
Share of most frequent type 11.65% 5.84 (+/- 0.27%)
Types for 50% coverage 15 33.6 (+/- 1.2)
Types for 80% coverage 135 190 (+/- 5.7)
Types for 90% coverage 387 436 (+/- 13.4)
± denotes one sample standard deviation across the 1,000 random samples.
The first two figures show the language-like side of the result. With the same amount of material, the VMS has essentially the same number of distinct core forms as the Bavarian corpus. The number of forms that occur only once - hapaxes - is also in the same range.
Type count and hapax count are not independent. Both react to rare forms. But taken together they show that the VMS does not have a tiny core inventory endlessly recycling the same handful of pieces. Behind the tightly regulated surface sits an extraordinarily broad reservoir of rare structures.
And that gives us the first important result.
Put differently:
The VMS has enough distinct and rare cores to represent a natural, historically variable consonant stream. That flatly kills again the claim that the VMS cannot contain language!
I compared these consonant clusters with Latin as well, using the same kind of structured pipeline (functional-word filtering, ending strip, geminate collapse) rather than raw text. Even with this fair treatment, Latin falls far short. Even without any deeper analysis, it became obvious that the Romance languages are distributed much more densely, simply because they contain more vowels and fewer consonant clusters. Which is actually pretty logical.
I compared it with Czech as well, a consonant-rich language, and that does not fit either. There is one caveat: I had no reliable way to determine what the above-mentioned families would correspond to in Czech. So that test has to be treated with caution!
I also compared it with Alemannic. That fits much better again, although not quite as well as the Bavarian text. Swabian also fits well, but worse than Alemannic. This suggests that a Germanic core structure works better overall.
The apparent contradiction - and a key to the cipher
The head of the distribution, however, does not fit at all at first. This is actually a well-known Voynich problem.
In Bavarian, the most frequent cluster accounts for about 5.84% of all occurrences. In the VMS, the most frequent reduced core accounts for 11.65% - almost twice as much.
To cover half of all Bavarian material, about 34 different types are needed. In the VMS, only 15 are needed at first.
So we get a strange picture:
The rare structures look like a large natural-language inventory.
The frequent part looks as if many different forms had been collapsed into a small number of strings.
This contradiction is exactly what makes the VBM interesting. The model does NOT claim that the visible Voynich cores are untouched plaintext clusters. It claims that they are organised, and partly compressed, by vowel bridges, boundary glyphs and families.
Independent checks on consonant clusters at word endings and word beginnings in the plaintext corpus also show that the core positions are strongly conditioned by classes and families. They do not combine like free letter strings.
Where the oversized head comes from
The three most frequent reduced cores are:
k: 825 occurrences
d: 780 occurrences
t: 490 occurrences
Together they make up roughly 30% of all cores.
What matters here is that these forms almost never occur in isolation. They are nearly always(!) - tied to a right-hand family state:
Code:
Core Bound to a family or to y
k 98.4%
d 97.9%
t 98.4%
In the original analysis, the right-hand families were removed. As a result, forms such as k+aiin, k+ain, k+ar, k+al, k+y were all counted
as plain k! That does not mean that all of these forms are identical in the cipher, or that they necessarily belong together. It only means that our first counting method threw them into the same drawer.
So I stopped doing that. Instead of counting every one of those cases as the same naked k, d or t, I counted the core together with the family context in which it actually occurs: k|aiin, k|ain, k|ar, k|y, d|aiin, d|ar, d|y, and so on.
That raised the obvious question: what if these glyphs are not supposed to stand apart from the families at all?
What if k+aiin is a different "k-state" from k+ain?
Given the extremely strong binding, and given what a cipher of this kind would look like, that is not a wild assumption. It is the obvious thing to test.
Important: nothing is altered here. Nothing is added. This is information that is actually present in the manuscript, and the strength of the binding confirms that it matters. The information was simply lost when the families were stripped and forgotten.
Once different k, d and t "states" are distinguished by their family context, the distribution suddenly changes:
Code:
Measure VMS before After separating k/d/t states
Top-1 11.65% 5.82%
Cov50 15 29
Cov80 135 164
Cov90 387 419
The most frequent VMS type now stands at 5.82%, almost exactly on the Bavarian value of 5.84%.
That is the first real "aha" moment:
The Voynich peak, which was almost twice as high, nearly disappears once the already visible family states are no longer artificially collapsed into single categories called k, d and t.
So the sharp head was not necessarily a property of the underlying language. It was, at least to a large extent, the result of family compression - or, more precisely, of our first core count being too crude.
The E-family explains the broad shoulder of the distribution
After separating k, d and t, the extreme peak is already explained. But between roughly rank 10 and rank 70, the Voynich distribution is still more concentrated than the Bavarian one.
And that is exactly where the E-family sits.
In the strict Stars definition - one anchor atom followed by one of ten fixed kernels (e, ee, eee, eo, eeo, ed, eed, eod, eeod, eeed) - it contains 72 distinct types and 1,599 occurrences. These include forms such as k.ee, k.eed, k.ed, k.e, t.ee, t.ed, ch.e, ch.ed.
These 72 types are not evenly distributed. Just 19 E-types among the first 70 ranks contain 89.4% of the entire E-family mass.
So the E-family
has two zones: a small, very frequent inner group, and a large, thinly populated fringe of rare combinations.
That explains its peculiar effect on the overall distribution. The frequent inner group forms a broad shoulder in the head. The rare fringe contributes to the long hapax tail at the same time.
The E-forms are also tied to their right-hand context, though not uniformly: forms ending in -ed and -eed are followed by y in 81.2% and 88.9% of their occurrences respectively, while -eod is followed by y in only 70.0% of cases - predominantly, but not "almost always" across the board.
So the E-family is not decorative free variation. It is a regulated matrix of connection states.
I then applied the same principle to the other side. For the E-family, I distinguished only the two directly observable states
"before y" and
"not before y".
The result is:
Code:
Measure VMS with family states Bavarian
Types 1,157 1,123 (+/- 22
Hapaxes 721 667 (+/- 23
Top-1 5.82% 5.84 (+/- 0.27%
Cov50 34 33.6 (+/- 1.2
Cov80 182 190 (+/- 5.7
Cov90 449 436 (+/- 13.4
± denotes one sample standard deviation across the 1,000 random samples.
Now it is no longer just one or two figures lining up:
The most frequent type has almost exactly the same share.
Almost exactly the same number of types is needed to cover 50% of all occurrences.
Practically the same number of types is needed to cover 90%.
The total type count remains in the same range.
The long hapax tail survives.
Across the first 70 ranks, the mean deviation of the cumulative distribution drops from 12.54 to 1.83 percentage points. That is a reduction of roughly 85%.
What is actually unusual here
The strong result is not simply: "Bavarian and Voynich have roughly the same number of types."
The stronger result is this:
the VBM could explain at the same time why one part of the distribution already fits Bavarian from the start, and why another part does not fit at first.
According to the model, the rare tail comes from the variety of the underlying consonant clusters. The compressed head is created because frequent cluster states are organised into families and then collapsed by an overly crude analysis.
So the model does not merely explain a match. It explains the exact shape of the mismatch.
Summary
Methodologically, this matters much more than a lucky curve match. A similarity found after the fact can always be dismissed as coincidence. But here the error sits exactly where the independently observed Voynich families dominate:
k, d and t create the extreme peak; the E-family shapes the range up to roughly rank 70; the remaining rare combinations preserve the long tail.
You could put it like this: the apparent failure of the model turns out to be
the fingerprint of the family mechanism - and therefore of the cipher.
Or even more simply: under the visible Voynich distribution lies a language-like distribution. The families squeeze the frequent part without destroying the rare tail.
What the test tells us about the cipher
This result also changes how the families themselves have to be interpreted.
They are apparently not just disposable endings that can be cut away and forgotten. They probably carry information about which exact cluster state is meant.
The relevant unit would therefore not simply be k, but something more like: k|aiin, k|ar, k|y k.ee|y, k.ee|NON-y.
That means a later decipherment cannot look only at naked cores. Core, family and right-hand connection have to be treated as one combined cipher unit, potentially carrying different values.
So the statistical comparison does not only provide an argument for natural language. It also gives a concrete clue as to how the cipher organises its most frequent consonant structures.
A cautious conclusion
This result proves neither a specific plaintext reading nor Bavarian beyond doubt. Alemannic texts also came reasonably close in the previous controls, so the more cautious linguistic label is "historical Upper German".
Romance and Latin comparison texts have much shorter cluster inventories; Old Czech lies between the Romance and Upper German values.
The family test was also developed on the same Stars section in which the deviation was observed. So this is not yet an independent blind test.
Of course splitting frequent types makes any distribution flatter. That is not the point.
The point is that the particular split demanded by the visible family structure makes the head, the tail, the type count and the hapax count line up with Bavarian at the same time.
That is the weird part.
Despite all the caveats, what I have shown here is substantial:
At core level, the VMS has the diversity of a natural Upper German consonant stream. Its strange surface distribution can largely be explained by an independently visible family system which, quite typically for a cipher, reveals a kind of compression mechanism.