30-08-2026, 09:21 AM
Things actually get a bit more complicated now, but I’ll keep it simple.
Theory 1: o / qo – the difference in length
o and qo share the same core vocabulary as prefixes.
93.1 per cent of qo-words use suffixes that also occur after the prefix o. Many pairs are almost balanced:
Suffix after o after qo
kaiin 211 265
kal 146 193
kar 135 157
keey 174 307
kain 138 279
‘qo’ is therefore not a distinct vocabulary; one might say it behaves like a brother to ‘o’
Hypothesis: ‘o’ and ‘qo’ differ in length, so the writer can influence the length by choosing ‘o’ or ‘qo’.
(If you look more closely, it becomes more complicated: in the short token range, ‘o’ still has the ‘r’ as a following letter, which ‘qo’ generally lacks, and of course ‘ol’, which is, however, a separate prefix of the same length as ‘qo’ but precedes other following letters (such as ‘ch’, for example). But this fits perfectly with the length theory.)
Theory 2: ch / sh – the difference in length
‘Ch’ and ‘sh’ are so similar that it has even been claimed they are one and the same character.
Strictly speaking, ‘sh’ is ‘ch’ with a macron in the form of a loop. This constitutes a kind of abbreviation and implies that one or more letters have been shortened here.
This means that ‘ch’ is 2 glyphs long and ‘sh’ is at least 3 glyphs long.
In a length-dependent cipher, ‘chedy’ would therefore be a 5-glyph cipher in one list, whilst ‘shedy’ would be a 6-glyph cipher in another list, and would thus mean something completely different.
Theorie 3 bank gallows - the reducer!
‘Gallows’ before ‘ch’ is free and occurs in abundance (kch, tch, pch — over 3,300 times)
‘ckh’ words have roughly the same overall width as ‘ch’ words when both are regarded as atoms, but they also contain a gallow.
The word absorbs the Gallows content at no extra cost. This makes the Gallows bank a good way to reduce length when it gets out of hand. And that is exactly what a length-dependent model would need. Especially if ‘ch’ is treated as a two-glyph word.
Summary
I am, of course, aware that this is merely a hypothesis at this stage. A hypothesis that undermines its own testability. This is because this length-dependency makes external statistical analysis very difficult, if not impossible, as one would actually need to know what the plaintext is. If ‘chedy 5l’ and ‘shedy 6l’ mean different things, or if ‘ked’ in a 5l token can mean something different from ‘ked’ in a 6l token, an external analysis is nearly impossible.
But interestingly, that would also be precisely the explanation for why no one has cracked this cipher so far. If, on top of that, the VBM (which fits this length-dependency quite well, as it removes prefixes such as ‘o’ / ‘qo’ etc. and suffixes like ‘y’ from the token) were also correct, the cipher would indeed be unbreakable using standard cryptographic methods without a good crib.
And back then there were homophone ciphers that played with length by using multiple dots or dashes or something similar – it’s not that far-fetched – though they probably didn’t exist at this level of complexity.
Theory 1: o / qo – the difference in length
o and qo share the same core vocabulary as prefixes.
93.1 per cent of qo-words use suffixes that also occur after the prefix o. Many pairs are almost balanced:
Suffix after o after qo
kaiin 211 265
kal 146 193
kar 135 157
keey 174 307
kain 138 279
‘qo’ is therefore not a distinct vocabulary; one might say it behaves like a brother to ‘o’
Hypothesis: ‘o’ and ‘qo’ differ in length, so the writer can influence the length by choosing ‘o’ or ‘qo’.
(If you look more closely, it becomes more complicated: in the short token range, ‘o’ still has the ‘r’ as a following letter, which ‘qo’ generally lacks, and of course ‘ol’, which is, however, a separate prefix of the same length as ‘qo’ but precedes other following letters (such as ‘ch’, for example). But this fits perfectly with the length theory.)
Theory 2: ch / sh – the difference in length
‘Ch’ and ‘sh’ are so similar that it has even been claimed they are one and the same character.
Strictly speaking, ‘sh’ is ‘ch’ with a macron in the form of a loop. This constitutes a kind of abbreviation and implies that one or more letters have been shortened here.
This means that ‘ch’ is 2 glyphs long and ‘sh’ is at least 3 glyphs long.
In a length-dependent cipher, ‘chedy’ would therefore be a 5-glyph cipher in one list, whilst ‘shedy’ would be a 6-glyph cipher in another list, and would thus mean something completely different.
Theorie 3 bank gallows - the reducer!
‘Gallows’ before ‘ch’ is free and occurs in abundance (kch, tch, pch — over 3,300 times)
‘ckh’ words have roughly the same overall width as ‘ch’ words when both are regarded as atoms, but they also contain a gallow.
The word absorbs the Gallows content at no extra cost. This makes the Gallows bank a good way to reduce length when it gets out of hand. And that is exactly what a length-dependent model would need. Especially if ‘ch’ is treated as a two-glyph word.
Summary
I am, of course, aware that this is merely a hypothesis at this stage. A hypothesis that undermines its own testability. This is because this length-dependency makes external statistical analysis very difficult, if not impossible, as one would actually need to know what the plaintext is. If ‘chedy 5l’ and ‘shedy 6l’ mean different things, or if ‘ked’ in a 5l token can mean something different from ‘ked’ in a 6l token, an external analysis is nearly impossible.
But interestingly, that would also be precisely the explanation for why no one has cracked this cipher so far. If, on top of that, the VBM (which fits this length-dependency quite well, as it removes prefixes such as ‘o’ / ‘qo’ etc. and suffixes like ‘y’ from the token) were also correct, the cipher would indeed be unbreakable using standard cryptographic methods without a good crib.
And back then there were homophone ciphers that played with length by using multiple dots or dashes or something similar – it’s not that far-fetched – though they probably didn’t exist at this level of complexity.
