The Voynich Ninja

Full Version: Why and how the text could be Bavarian
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44
[attachment=17288]
Bavarian – the divine language?

If you speak the various German dialects yourself, and then hear them spoken, you can tell where they come from.

For those who don’t speak German, there’s no point in translating the text. If you don’t understand German, the dialects certainly won’t make things any easier.

But the fact remains: the German text in the VM is definitely Bavarian.

The written VM text certainly isn’t.

Translated with DeepL.com (free version)
Addendum:
Perhaps I should also mention that there are four German dialects in the caption.
Berlin, Cologne, Hamburg, and, well, Bavarian. (Perhaps Munich)
@ Age Thanks for the joke Wink

...but I’m just as convinced that it isn’t Latin as you are that it isn’t Bavarian. One of us – or both of us – will be proved right

...aber ich bin genauso überzeugt davon, dass es kein Latein ist, wie du, dass es kein bairisch ist. Einer von uns oder beide werden Recht behalten.
The Length-Dependent Part of the Cipher

As I’ve been writing for some time now: I realized that I was still missing an important part of the cipher—the missing 20 percent....

I’ve found it. And with that, everything is now much clearer—and, unfortunately, also a bit more complicated:

In addition to the vowel bridge, the VMS also uses a length-dependent ciphering of the cores.

I can now also see how the cipherer went about it. There was likely initially a length-dependent cipher for tokens up to 3/4 glyphs in length and for those with longer tokens. This had to be the case because vowels tend to be the first and last letters of the token. This means that this principle only works if words are longer than three letters. I had already intuitively realized this and treated the short tokens differently from the longer ones—without realizing the logic behind it. This also explains why the short tokens are so very different from the longer ones (something that has been known for a long time).

After all, he naturally needed separator characters as well—for example, when one word ended with a vowel and another began with one, and for other cases that arise automatically due to the VBM.

This gave him the idea to incorporate length dependence into additional tokens as well, and that’s how he came up with the counter letters: which, as in the Eva transcription, are read as “e” and “i.” But in principle, they’re just dashes.

At first, I did recognize the counting structure and ran through every conceivable variation, but I was always missing a part of the structure, and it just wasn’t consistent no matter how I approached this matrix.

Until it dawned on me that he had actually taken the length-dependent game even further: He had different lists for words of different lengths.

Example: So “ked” in a word that’s 5 glyphs long can mean something different than “ked” in a 6-glyphs word. He introduced the counter strings to create a second level, making the words “longer” and thereby placing them in a different length category—where they then have a different meaning.

VMS “ked” = “nd,” but now he needed a different consonant cluster, so he simply inserted a second “e,” resulting in “keed”—this could be (!) a completely different cluster that appears in a different list. This way, he expanded his selection while simultaneously making the cipher nearly unbreakable.

That explains why there are so many words that look the same but are missing a letter or have an extra one. Even though they look almost the same, they are completely different words.

And only in a final step did he think to himself: Oh, I can apply this to the vowels too. Because with them as well, you can see a length-dependent change in their context. That is, the vowel bridges behave differently depending on the length of the word.

And that’s exactly the aspect of the cipher that accounts for the missing 20 percent. Very exciting.

I know: I still have to prove it, but I'm working on it. I've found a few more clues, and I'm now putting it all together like a crossword puzzle. It's going to take some time.

Unfortunately, this theory (and that’s all it is so far) has completely thrown off the basic structure of my generators, so I’ll have to keep working on that as well. But since I’m on my own, that’s obviously a lot for one person to handle.


Conclusion:

But I think we should at least consider the idea that this cipher could habe a component that depends on word length.

It would explain why the cipher has driven so many cryptologists to the brink of madness—in addition to the VBM structure. The two together are nearly impossible to crack.

But it still roughly fits the ciphers of the time; only the complexity is a bit anachronistic.... I’ll admit that, but that’s what the numbers show...
I know it's a little hard to understand, given this length-dependence.

As an Example let’s take “aiin”—a word whose overrepresentation has already driven generations of VMS cryptologists to despair: 10 percent of tokens contain “aiin.” (which, in terms of frequency, matches the ending “en” in German perfectly...
In that regard, the “aiin” here is just an example meant to illustrate what this is all about.)

Now imagine this: “aiin” as a 4-glyph word (i.e., standing alone) means something different than “aiin” in a 5-glyph word, such as “daiin” or “kaiin.” And “daiin,” in turn, means something different in a 6-glyph word.

At that point, the frequency of “aiin” would break down into small clusters. Here’s a table illustrating this.

[attachment=17460]

From 0.12 to 3,42 percent—that’s much more manageable than 10 percent.

If you design this cipher this way, you can be sure that it will drive any poor cryptologist to the brink of madness!

And now you have to apply that to the other families...

Suddenly, the VMS is much less repetitive Wink
To illustrate what I mean:

Here is an overview of how this division of syllables can be recognised, starting with a rough outline: The ‘aiin’ family and the ‘e’ family only really begin with the 4-syllable word (which, in a way, is hardly avoidable in the case of the ‘aiin’ family, given that it also contains ‘ain’, ‘an’ and ‘n’); logically, the ‘e’ family can only really begin at 3). At the same time, the ‘r/l’ family drops off sharply at this 4-syllable threshold.

[attachment=17473]

We can therefore identify a clear length-dependent shift, a break that lies between 3 and 4.

The internal logic is also interesting: the ‘aiin’ family continues to rise the longer the word is, whilst the ‘e’ family forms a bell curve.

From this graph alone, despite all the normal length effects, it is already apparent that a certain dependence on length might be possible (the graph does not reveal anything more than that).
An example of the length dependence of vowels in the VBM:

A word ends in y. How often does the next word start with qo? Sorted by the length of that next word:

[attachment=17476]

it climb to 6, and collapse at 7

Why, and why so strong and suddenly?  Wink

The “qo” is replaced by other building blocks. Here’s the full list:

[attachment=17478]

A different bridge is used, one that relies more on y:o and y:ch.
y:che actually surges by almost three times, effectively offsetting the slump in qo
 Interesting, isn’t it?

At the very least, this shows that it exhibits a certain degree of length dependence.

The reason for this, of course, remains unclear. Whether it forms part of the cipher is, at the moment, merely my own interpretation. It could simply be a normal linguistic feature. But I have further evidence; more on that soon...
So I’ve calculated the E-family in terms of length. The table shows the percentage share.

The first row is: e alone (alone)

Anchor is the letter before the ‘e’. Run is from ‘e’ to ‘eee’.

[attachment=17484]

We can, of course, see a ‘step’ here, as the length of the ‘e’ runs is naturally part of the token’s length.

I then removed this by defining e = e1 / ee = e2 / eee / e3 as an atom (just as qo is an atom for me)

[attachment=17485]

then the staircase disappears and the subsequent pattern becomes homogeneous

That was, of course, to be expected. But it shows something:

1. The functional unit e is a kind of ‘composition symbol’. As soon as one counts runs as a single symbol, all minimum lengths collapse to two values and all building blocks lie within the same band L*3–6.
It almost looks as though the EVA transcription has transcribed the e-runs incorrectly, they are atoms. Unless – and this is the theory behind it – the cipher is length-dependent.

2 The token architecture thus appears considerably more structured than it would be without this ‘trick’.

Now the crucial question: why didn’t the writer simply choose three different glyphs: namely e1, e2 and e3? 
Because these would not have altered the length!

The same structure is found in the aiin Family... Wink

But it gets even better...
(30-08-2026, 07:55 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.But it gets even better...

I don't see anything strange in this. It seems to me logical that the writer just likes the easy-to-repeat strokes  i and  e . He is preferring to write in a style that is not too taxing. I mentioned something about this in a previous post

You are not allowed to view links. Register or Login to view.
(30-08-2026, 08:40 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.
(30-08-2026, 07:55 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.But it gets even better...

I don't see anything strange in this. It seems to me logical that the writer just likes the easy-to-repeat strokes  i and  e . He is preferring to write in a style that is not too taxing. I mentioned something about this in a previous post

You are not allowed to view links. Register or Login to view.

Yes, that’s certainly a plausible alternative theory. No doubt about it. But there are a few more points to consider regarding the length theory.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44