The Voynich Ninja

Full Version: Voynichese is a numeric cipher?
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
Taking this opportunity, I would like to reiterate the main similarity between Roman numerals and Voynichese.
The Roman numeral system itself has strict positional rules, according to which IV = 4, VII = 7, but IIV = nonsense. 
A similar detail is present in the VMS text, and to explain it, I don’t need to resort to slot grammar. Simple examples: daiin = iinda, cthodol = dolctho, cthy = ycth. As you can see, in the second case, we end up with a non‑existent word. That is, this is the most basic rule of Voynichese — the letters must be arranged in a certain order to form a word.
However, the VMS alphabet is much larger than the basic set of Roman numerals. This means that some sequences are encrypted in a non‑standard way (as in the mini‑cipher: IX = or/ol, but IV = ok/ot), and some may change their spelling depending on their position in a string or word (for example, let’s assume that cthodol = XXXVIX. Four identical symbols are encrypted differently).
Of course, it is still possible that the Voynich writing system may have allowed a much wider range of glyph combinations and positional rules than those found in the manuscript. We can only infer rules from the sequences that happen to appear. Perhaps the source material, if there was any, simply had no need for the other possibilities.
(25-08-2026, 08:01 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Of course, it is still possible that the Voynich writing system may have allowed a much wider range of glyph combinations and positional rules than those found in the manuscript. We can only infer rules from the sequences that happen to appear. Perhaps the source material, if there was any, simply had no need for the other possibilities.
The situation, apparently, is exactly as it appears to be...
But it will still be useful to at least study this cipher. Again, I think it wasn’t invented by a brilliant cryptologist, which means it may be complex and redundant, but not that brilliant.
Encrypted abbreviations or parts of words 

@JoJo_Jost said that if the cipher alphabet includes abbreviations like -us, deciphering becomes difficult:
(12-07-2026, 09:34 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.7. two cipher tables—one with letters and the other with abbreviations. (which, however, makes decryption nearly impossible).
Let’s take the word from You are not allowed to view links. Register or Login to view. as an example: aralarar
[attachment=17384]
First, the letters are almost in order, except for the first ar. It comes first, even though it seems like it should be placed together with the other ar at the end of the word.repeated. 
Secondly. If we break the word down into bigrams, we’ll see that the word consists of 4 letters, 3 of which are repeated. It looks strange.
It can be assumed that the last arar is a separate number, and that is why it stands apart from ar. Then the word will consist of three letters. The word aer comes to mind — interestingly, the letters are arranged in ascending order in the alphabet. In the Voynichese word, they also seem to be arranged in ascending order, from the small ar to the large arar. But then it’s unclear why the letters suddenly went in ascending order if they should be in descending order. And this logic doesn’t change anything — if arar is greater than ar, ar should come behind arar, and in any case, it will be alararar... 
... But there is another explanation! Returning to my idea that the same symbols can encode different information, or may not encode it at all in some cases, we can assume that the first ar encodes something different from the final ar. This could be an example of an encrypted abbreviation, as in the word “continuus” = “9tinu9”, where the same “9” symbols represent different morphemes located in different places in the word. This explains why it stands alone — it may belong to a different alphabet, or the author simply wanted to make our work easier by allowing us to distinguish between the letters and abbreviations in the encrypted word  Big Grin

Of course, this word may have other interpretations, and all the ar's may be equal to each other, but the author simply wanted to put one ar at the beginning, as he didn’t want to see ararar; however, in this interpretation, it may tell us a little about the statistical counteraction by this cipher. 
@oshfdk expressed doubt that we can't distinguish nulls from real symbols*. Here is an excellent illustration of the fact that at present it is difficult to distinguish a null from a real symbol: ordinary frequency analysis will tell you nothing, because... you will simply count the number of all ar.

*Besides the fact that it sounds crazy, the very essence of a cipher with nulls is to blur the statistics of the symbols. He was talking about some kind of difference in the distribution between groups of symbols, but... what is this difference? For example, the statistics for the letter “e” in English differ from those for other letters. So, does that mean it doesn’t mean anything?

By the way, this example can also demonstrate that order is present both in words and in strings, as shown in the example in the following image, where before the symbol EVA s (the предполагаемое beginning of a new line) there is a beautiful series "lol kar shr r ol ol". In particular, we see the “reduction” of ar-r-r and lol-ol-ol.
Voynichese and Copiale and The main advantage of Voynichese

Up until this point, I have had some difficulty describing the hypothetical numerical cipher involved in the VMS. I doubted that such a cipher had any historical precedent or would even work.
But after studying the Copiale cipher a bit in my spare time, I thought it actually resembled my assumptions. Through comparison, I will try to describe the numerical cipher to you and let you see how certain parallels with Copiale might work:
1). ...each ciphertext character stands for a particular plaintext character, but several ciphertext character may encode the same plaintext character.
I assume something similar happens with numbers. Examples from my mini‑cipher: XV = o/qo and VI = an/dy. Again, as I was saying, the Voynichese script may allow for different ways of writing the same numbers.
2). In addition, some ciphertext characters stand for several characters or even a word.
I’m not sure about a few specific symbols, but if the numeric cipher is adapted for Latin, it could well encrypt some short prepositions or parts of words (for example, endings like -us, -rum).
Moreover, I assume that if a symbol encodes a contraction, it must encode something else as well.
3). Seven ciphertext characters encode the single letter e.
On the contrary, one character can encode several. I already learned about this, but I still put it in a separate point.
4). All the unaccented Roman characters encode a space.
No, there’s nothing like that in the numeric cipher. I suppose there isn’t such a wide range of symbols with a single meaning.

And now, what we have is a well‑developed theory of how the proposed numeric cipher works: Plaintext - Numeric Alphabet - Cipher alphabet
1). The cipher alphabet is not intended for encrypting letters, but for encrypting Roman numerals. It covers a basic set of symbols (I, V, X; larger numbers like C are also possible) and some of their combinations (similar to my mini‑cipher: VI = an/dy; V = a, I = n, but IV = ok/ot). Missing combinations are recorded by using the available symbols.
I think the principle behind forming special combinations is a bit arbitrary — for example, VI = dy, then logically IV = yd, but such a bigram apparently doesn’t exist in VMS. In my cipher, IV = ok/ot. Essentially, here we have the misleading letter o, which equals 15 individually, and the gallows. But we can do it more simply — write IV as yk/yt, simply replacing d with another symbol.
A very small set of symbols covers a wide range of meanings — letters, words, nulls, special characters, etc. Without any clues regarding the role of a specific symbol, a very complex cipher is obtained, but they can be easily introduced. For example, in my mini‑cipher, yk/yt is nonsense. You understand that y (equal to I individually) does not carry any meaning due to the dummy symbol k/t, which, in essence, is a service symbol in this case.
2). The numerical alphabet consists of a set of Roman numerals (which is very surprising and unexpected Exclamation ). Most likely, the author first selected the meanings for the letters and only then proceeded to expand the alphabet with new morphemes, words, etc.

This comparison shows that the hypothetical numerical cipher and Copiale are quite similar in principle. In turn, the Copiale cipher (in my opinion) is fundamentally similar to classical medieval homophonic ciphers. The principle of nulls, various signs and words is not new in itself:
[attachment=17506]
Thus, we can say that a numeric cipher would be quite modern for the 15th century, as it would be fundamentally similar to homophonic ciphers.

...Why didn’t we solve it then?
In my opinion, the reason for the resilience of VMS lies in its two fundamental features:

1). The cipher alphabet encrypts only numbersBecause of this, when discussing Voynichese, we are discussing numbers, but not letters. We don’t know what the numbers themselves encode or how they are distributed in the text (I mean some special rules that may not be visible to us). And we have no clear hint about the text. For example, could daiin be a letter, a preposition, an abbreviation, or a full word?
2). Lack of differentiation in symbols. This is a glaring feature. In the Copiale cipher, three categories of characters can be условно distinguished: Latin characters — nulls, Greek characters — vowels, abstract symbols — consonants (this is not entirely accurate, but it serves for clarity). And the turning point in deciphering was the removal of characters from one of these categories (Latin nulls) from the text. In the case of Voynichese, removing the symbols will not lead to anything, as we risk losing half of the text. If there are nulls, they may be symbols from the cipher alphabet that are not arranged according to the rules (such as ch in my examples of encrypted texts). But even this won’t free us from the burden of deciphering the meanings of the symbols…

What can statistics tell us?

Let’s try to compare the indicators of Zipf’s law and entropy for VMS and Copiale. 
The Copiale text, unprocessed for nulls, does not conform to Zipf’s law (this also happens due to homophones, but to a lesser extent). The stats just flatten out. But VMS conforms.
Copiale h1 = 5.5-6.0, Copiale h2 = 5.0. VMS h1 = 4.0, VMS h2 = 2.5 (or around two, I don't remember exactly).

This merely confirms the supposed features of the numeric cipher...
There’s clearly a catch in the signatures…
[attachment=17544]
I can’t decipher these words letter by letter, but I can identify some features when comparing the words and their hypothetic translations:

1). There is a connection between phlegm and yellow bile. The thing is that ph and f are sounds that are identical in terms of how they sound. If you replace ph with f, you get a perfect correspondence: Flegma — Flavus and otedyotol. It turns out that ot could theoretically be either F or Fl.
2). The two a’s in the word atra correspond to the two d’s in the word dchdy.
3). There are no correspondences with the word sanquis (olkchs). This may suggest that chs (variant of che?) is a single number.
magnesium Wrote:qokeedy qokeedy qokedy qokedy

Could encrypt something like the Roman numeral MMXX.
Actually, this is quite interesting. In this setup, the cipher takes on a different form — the letters represent digits, and combinations of similar symbols are numbers. Then repetition elements like ch (which I call nulls) can serve as an auxiliary sign indicating that a digit belongs to a specific number.

This is something to ponder. So, qokeedy qokeedy qokedy qokedy = MMXX. But what then is okedy? This could be the same X, but as a number, not just a single digit. For example, qokeedy qokeedy okedy = MX X.

Overall, we run into the same issue with the numerical cipher — a single number can be expressed using different methods. But then we lose all sorts of possible clues (in the form of i = one, for example), and the cipher becomes... virtually unbreakable. 

Essentially, this turns into a book cipher, or there’s some kind of algorithm that allows you to extract numbers from these words (maybe qokeedy is a set of numbers that add up to M).
An idea for a cipher with a kind of numeric component.
Taking the scheme of prefix - core - sufix.
The core represents a plain letter, and the set of sufix-prefix around the core moves the letter
certain number of positions, in the style of a caesar cipher.
This will reduce the field, 1 core per plain letter is enough, and the number of sets of prefix-sufix doesnt 
need to be high to have a high number of options to encode every letter.
123 of the 947 Voynich labels begin with EVA “ot,” which is about 13.0%. The count included labels with more than one character.

Does that mean 13% of the words are supposed to "F"?  Wink
(11-09-2026, 04:38 PM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.123 of the 947 Voynich labels begin with EVA “ot,” which is about 13.0%. The count included labels with more than one character.

Does that mean 13% of the words are supposed to "F"?  Wink
No one knows. Maybe it works for all words. Maybe only for Currier B. Maybe there are some hidden rules here. But there are still some coincidences.

(11-09-2026, 01:58 PM)Juan_Sali Wrote: You are not allowed to view links. Register or Login to view.An idea for a cipher with a kind of numeric component.
Taking the scheme of prefix - core - sufix.
The core represents a plain letter, and the set of sufix-prefix around the core moves the letter
certain number of positions, in the style of a caesar cipher.
This will reduce the field, 1 core per plain letter is enough, and the number of sets of prefix-sufix doesnt 
need to be high to have a high number of options to encode every letter.
Hmm, this is quite interesting. But can Core tokens really be encrypted letters? Or rather, can we try to come up with some kind of cipher alphabet?
Overall, you can imagine something like a Polybius square. Prefix shifts the letter horizontally, Suffix — vertically.
Pages: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21