The Voynich Ninja

Full Version: Why and how the text could be Bavarian
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44
Things actually get a bit more complicated now, but I’ll keep it simple.

Theory 1: o / qo – the difference in length

o and qo share the same core vocabulary as prefixes.

93.1 per cent of qo-words use suffixes that also occur after the prefix o. Many pairs are almost balanced:


Suffix after o after qo
kaiin    211 265
kal      146 193
kar      135 157
keey    174 307
kain     138 279

‘qo’ is therefore not a distinct vocabulary; one might say it behaves like a brother to ‘o’

Hypothesis: ‘o’ and ‘qo’ differ in length, so the writer can influence the length by choosing ‘o’ or ‘qo’.

(If you look more closely, it becomes more complicated: in the short token range, ‘o’ still has the ‘r’ as a following letter, which ‘qo’ generally lacks, and of course ‘ol’, which is, however, a separate prefix of the same length as ‘qo’ but precedes other following letters (such as ‘ch’, for example). But this fits perfectly with the length theory.)

Theory 2: ch / sh – the difference in length

‘Ch’ and ‘sh’ are so similar that it has even been claimed they are one and the same character.

Strictly speaking, ‘sh’ is ‘ch’ with a macron in the form of a loop. This constitutes a kind of abbreviation and implies that one or more letters have been shortened here.

This means that ‘ch’ is 2 glyphs long and ‘sh’ is at least 3 glyphs long.

In a length-dependent cipher, ‘chedy’ would therefore be a 5-glyph cipher in one list, whilst ‘shedy’ would be a 6-glyph cipher in another list, and would thus mean something completely different.

Theorie 3 bank gallows - the reducer!

‘Gallows’ before ‘ch’ is free and occurs in abundance (kch, tch, pch — over 3,300 times)

‘ckh’ words have roughly the same overall width as ‘ch’ words when both are regarded as atoms, but they also contain a gallow.

The word absorbs the Gallows content at no extra cost. This makes the Gallows bank a good way to reduce length when it gets out of hand. And that is exactly what a length-dependent model would need. Especially if ‘ch’ is treated as a two-glyph word.

Summary

I am, of course, aware that this is merely a hypothesis at this stage. A hypothesis that undermines its own testability. This is because this length-dependency makes external statistical analysis very difficult, if not impossible, as one would actually need to know what the plaintext is. If ‘chedy 5l’ and ‘shedy 6l’ mean different things, or if ‘ked’ in a 5l token can mean something different from ‘ked’ in a 6l token, an external analysis is nearly impossible.

But interestingly, that would also be precisely the explanation for why no one has cracked this cipher so far. If, on top of that, the VBM (which fits this length-dependency quite well, as it removes prefixes such as ‘o’ / ‘qo’ etc. and suffixes like ‘y’ from the token) were also correct, the cipher would indeed be unbreakable using standard cryptographic methods without a good crib.

And back then there were homophone ciphers that played with length by using multiple dots or dashes or something similar – it’s not that far-fetched – though they probably didn’t exist at this level of complexity.
Your theories interested me because I arrived at roughly the same conclusions. In my mini‑cipher, o can turn into qo, and then I noticed that some people generally adhere to the idea of such a connection.
But I’ll allow myself to ask a question based on the fact that you’re gradually developing theories and new rules — isn’t your theory becoming more complex or overly precise?
(30-08-2026, 09:37 AM)ololololo Wrote: You are not allowed to view links. Register or Login to view.Your theories interested me because I arrived at roughly the same conclusions. In my mini‑cipher, o can turn into qo, and then I noticed that some people generally adhere to the idea of such a connection.
But I’ll allow myself to ask a question based on the fact that you’re gradually developing theories and new rules — isn’t your theory becoming more complex or overly precise?

Yes, and no.  Big Grin Wink

Well, I can see the problem, and I’m not sure whether it’s getting worse either. That would be a clear warning sign.

But in principle, I don’t work on the basis of: I’ve got an idea and now I’ll do everything I can to make it fit.

Actually, I try to base my work on the statistics / numbers I see. For example, the theory about vowel bridges and vowels arose from the figures relating to word transitions. That then fitted surprisingly well with the distribution of vowels in German texts.

And then it turned out that this theory fitted the VMS surprisingly well. For instance, using only the vowel rules, I was able to model all word transitions without any additional rules – even slightly better than these seven rules. That shows that something is working there.

However, I still hit a brick wall at some point. Time and time again, and I couldn’t understand why, because the rest fitted so well.

Then I realised that I was still missing a part of the cipher. Something I’d overlooked.

So I thought, I’ll do it the other way round – I’ll try to reconstruct the cipher. But the generators I designed also hit that wall. Statistics such as entropy, the Zipf curve and many others were good, mainly because I was using very few rules. But even here, those ‘20’ per cent were missing.

And then, looking at other statistics, it occurred to me that the cipher might be length-dependent – suddenly it all seemed to make sense. And I demonstrated this more and more clearly. By that point, I’d certainly already carried out around 100 statistical analyses...

Even so, I’m not ready to say: ‘This is definitely the solution!’ But it could be the missing piece of the jigsaw – perhaps not even for my VBM, but for someone else’s idea.

However, I can see that the VBM is a good way of tapping into many of the VMS’s special features. Does the length-dependent cipher make it more complicated? Yes. But it also makes it more coherent. I’m curious myself to see how things will develop from here.

So: yes and no Wink
(30-08-2026, 09:21 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.I am, of course, aware that this is merely a hypothesis

I still believe you are thinking too hard. This super-complex cypher theory would have been too awkward for the writer. The solution is probably more simple.

Complicated cyphers might have been okay for items of diplomatic correspondence that needed just to be made unreadable by any third party. They needed just to be decyphered once and read once. The VMS however claims to be some compendium of scientific knowledge, a reference book to be consulted often.

If the text is indeed meaningful then it would need to have to be easily readable. No hard algorithm to encode / decode. It would have been sufficient just to do enough to obfuscate the meaning.
(30-08-2026, 10:52 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.No hard algorithm to encode / decode.
Think of one-time pad. The essence of this cipher is as simple as possible — the key must match the length of the message and be used only once. But anyone who doesn’t have the key will never be able to decipher it.
The VMS cipher may well be difficult for us, but for the author it could have been easy, or at least worthwhile. He also had an advantage over us — he knew more about it than we did...
(30-08-2026, 10:52 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.This super-complex cypher theory would have been too awkward for the writer.
Here, I’ll say that I don’t know whether @JoJo_Jost is discussing the entire manuscript, Currier A, or Currier B. For Currier A, for example, such a cipher would be quite suitable, since the pages of Herbal A usually contain rather short text excerpts. The author didn’t bother with them much.

The argument about complexity only works when it makes the algorithm impossible or overly cumbersome. Don’t forget that you’re looking at a book that has defied decipherment for 600 years. Is it really worth relying on a substitution?
@ You are not allowed to view links. Register or Login to view. 

Deciphering it is very complex and complicated. But the cipher itself is relatively simple.

Here’s an example. This is an old, discarded form (note: the cipher doesn’t work like this (!))

That one wasn’t finished either - it was a test...

But overall, I’m trying to keep it as simple as possible; after all, it has to fit the context...


[attachment=17489]
I doubt anyone ever considered each e or I to be a single glyph. The e and I families were recognized from the start as just a means to transliterate the glyphs to something easier to work with.

There were those who believe that two e’s are a unit that can precede a single e, but I doubt that many others would agree.

There is also the @ weirdo which results from an a with the n fused to it. I’ve sometimes pondered whether the choice is 0-4 I strokes including the a. For the e family, the rare b can be preceded by eee for 4 ‘e’ strokes.

I like to think that they do indicate a count of places before or after the plaintext character(s) the following glyph would break to…. But hey, I’m just someone who expects there is a plaintext hidden by some kind of encryption system.
(30-08-2026, 12:23 PM)Grove Wrote: You are not allowed to view links. Register or Login to view.I doubt anyone ever considered each e or I to be a single glyph. The e and I families were recognized from the start as just a means to transliterate the glyphs to something easier to work with.

There were those who believe that two e’s are a unit that can precede a single e, but I doubt that many others would agree.

You can test this; the various ‘e’s are likely to contain meaningful information to a certain extent....
Another clear indication of length dependence.

[attachment=17497]
(It just occurred to me: “ain” refers to the entire “aiin” family (not just “ain”), and “edo” refers to the entire ‘e’ family—that is, “ed,” “eod,” etc. and calculated based on Eva ZB, excluding 57v and 116v, and each line containing more than 2 tokens)

Two route diagrams: one for words beginning with “o,” one for words beginning with “qo.”

First observation: After the gallows, the two routes follow almost the same number (yellow circles)
If you compare the values, you can see how close they are; the only outlier is the “t” to “e” family.

These continuation profiles are closer to each other than those of any other pair of units in the system. Whatever the choice between “o” and “qo” signifies, it reveals almost nothing about how the word continues—the word seems to virtually forget its bridge as soon as the gallows is written. (This supports what I wrote in the post above about “qo” and “o.”)

Second observation: “o” has the short endings o-> L = 19.8%, o -> r = 7.2% (ol, or, and related words)—where a word can simply end. This makes sense, of course; since “o” is only one glyph long, it’s needed for short words. “qo” has hardly any to practically no transitions in that direction: (l 4.7%, r 0.6%). Thus, nearly every “qo” word is forced to use the long mechanism.

In short, the choice of this prefix does say something about the length of the word. This fits perfectly into a length-dependent cipher.

(At the same time, this is further evidence that systems based entirely on direct language substitution cannot be correct. As far as I know, this does not exist in any language. It would be interesting to know whether this exists in Chinese, but I’m not familiar enough with the language to say.)

In summary: It is a choice that provides almost no information about what follows, but results in a clear difference in “initial permissions.” “o” is one glyph wide and serves as a bridge to complete short words. “qo” is two glyphs wide, excluded from the short forms, and therefore occurs in longer words: L4–L6 contain 84,8% of all tokens beginning with “qo”,while “o” is distributed more evenly

This is exactly the pattern one would expect if a word’s width matters: The same grammar underlies both bridges, but the one-glyph bridge serves the short words and the two-glyph bridge serves the longer ones. And in precisely such a system, it is possible to determine a word’s length class based on its first pair of glyphs. In all other systems, this is somewhat strange.
(31-08-2026, 09:24 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.Two route diagrams: one for words beginning with “o,” one for words beginning with “qo.”

First observation: After the gallows, the two routes follow almost the same number (yellow circles)
If you compare the values, you can see how close they are; the only outlier is the “t” to “e” family.

Interesting observations and a nice diagram.

Some time ago I did something similar: I calculated the geometric distances of Voynich glyphs according to the distribution of their previous and following characters (the 'routes' in your diagram) and I made some observations you might find interesting.

"k" and "t" are indeed very similar, both in regard to the preceding and the following characters. Their distances are distance_previous = 0.12, distance_following = 0.12 (the minimum possible distance value is zero, the maximum is sqrt(2) =~ 1.414). Only a few characters couplets score better on distance_following, but they are all trivial cases, for example the distance_following between "n" and "m" is 0.03. Only one non-trivial couplet scores better on distance_previous: "cth" "cph", at 0.10 (*) [the 'p' is not a typo].


Re-doing the same, but considering also "ot" and "ok" to be single characters we get:
[attachment=17508]
Here, 'k' and 't' mean 'k not preceded by o' and 't not preceded by o'.

Some couplets are even more similar on distance_following than 'k,t' in the original text, including 'ot, ok', slightly better at distance = 0.11 vs. 0.12. Others are more divergent, from just a little ('ot, t' and 'k, t') to rather higher ('ok, t' at 0.23).

The distances_previous are surprising because they are constantly much higher than the original 0.12 of 'k,t'. I can explain this to me (but I did not check!) only hypothesizing that 'k' and 't' are preceded by 'o' so frequently to dramatically bring down the distances, which could be taken to support the hypothesis that 'ok' and 'ot' are two 'atoms' of the script (which would mean 'qo' is not an 'atom', even if 'qok' and 'qot' could be).


Do whatever you want with the above!






(*) In the analysis I considered "cth", "ckh", "cph", "cfh", "ch" and "sh" to be single characters. Transcription Rf1a-n, full text, all uncertain spaces are spaces, all words including a '?' have been discarded.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44