The Voynich Ninja

Full Version: Interchangeable ch/sh and k/t characters
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6
(29-06-2026, 06:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Anyway, let me insist again that statistics (glyph and word frequencies, Zipfness, entropies, correlations etc) are not properties of languages, but of texts.  In any language one can have texts that have much lower or much higher word entropy than a typical novel or philosophical treatise.   

If a given text has a character entropy that is substantially lower or higher than the character entropy stat that is usually associated with texts in that language, then one should be able to explain how this has happened.  

The problem with Voynich solutions is that solvers* do not explain how their target language has such a low entropy when written in the Voynich script and would presumably in its usual script have a higher entropy.  In most cases, I would imagine they cannot explain, because their solution is actually not just a simple substitution system but a product of multiple, diverse, and fatally inconsistent decisions to add or remove letters in order to match Voynichese words with their target language, and none of this can be formalised and reproduced.  

That is why we bang on so much about character entropy here.   If the manuscript isn't meaningless, there must be an explanation, whether it's simple or complex.  But solvers don't bother to address it.

*One rare exception that comes to mind is You are not allowed to view links. Register or Login to view.with the fada in Irish but I don't think she mentioned ever measuring if that were enough to resolve the entropy problem, and I doubt it is.

[Edit:  Getting back on topic here, what I find interesting is how different /sh/ and /ch/ can behave while being so similar.  Despite making similar words like shedy, chedy, sheey, cheey, etc, they definitely have their own "personality" in the manuscript.  /sh/ in particular appears much more in the Top Row of paragraphs than you would expect from its performance in lower lines.   Meanwhile /ch/ tends to go from being a common initial and a rarer word-middle to being a rarer initial and a common word-middle, forming words such as opchey.]
(29-06-2026, 08:34 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.what you say is strictly true, but in practice the statistics of a text are (barring rare outliers, as the famous English novel written without a single 'e') mostly governed by the underlying language (which includes, of course, its specific orthography) and not by the text. Once a language has been chosen it's usually possible to say if a given text is written in that language or not just by checking the most basic statistics.

I insist that this claim holds only for texts of the same type, namely texts that consist mostly of discursive or narrative prose -- like novels, chronicles, philosophical treatises, newspaper and magazine articles, even textbooks.  Like all the books in your example.  

You cannot count on this claim if the text has a peculiar structure with little free prose, even if whatever is there is perfectly grammatical.  Like a ship's log book, a terse list of saints, recipes, or kings, a catalog of products, a gazetteer, a book of exercises of algebra...  

The most common words in such texts can be very different from those in discursive texts.  Otherwise common words like "the", "of", "and" may occur scarcely or never.  And any statistics of characters and bigrams are largely determined by their presence or absence in the few most common words in the text.

And, of course, character-based statistics are totally dependent on the orthography.  You may say that Italian with any orthography other than the standard one "is not Italian".  Even if that was a reasonable view in other contexts, it is not a useful position when trying to identify the language of a text with a completely original alphabet.  Even if the VMS is a transcription of I Promessi Sposi, it will not be in the standard orthography...

All the best, --stolfi
(29-06-2026, 10:55 PM)tavie Wrote: You are not allowed to view links. Register or Login to view.The problem with Voynich solutions is that solvers* do not explain how their target language has such a low entropy when written in the Voynich script and would presumably in its usual script have a higher entropy.

If a proposed solution is a simple one-to-one letter substitution cipher, the frequency distributions of glyphs (D1) and of digraphs (D2) of the encrypted text should be the the same as those of the plaintext, modulo the substitution.   (The zero-order entropy H0 is determined by D1 and the first-order entropy H1 is determined by D2, so looking at D1 and D2 should be more revealing than looking at H0 and H1 alone.)

But this is not the case if the claimed solution is anything more complicated than a one-to-one substitution cipher.   Then the D's (and H's) of the encrypted text may be quite uninformative.  One can easily construct a deterministic lossless encryption method that will take typical English prose text and produce an encrypted text with arbitrarily low H0 and H1 entropies. Or with H0 and H1 equal to the limit of log2(26) = 4.7 bits per letter.

Therefore, if one intends to use the distributions D1 and D2 (or the entropies H0 and H1) to test the validity of the solution, one should instead compute them on the decoded plaintext.  The result would be the same in the simplest case, but would be meaningful for any deterministic lossless cipher, even a codebook (nomenclator) one.

However, what should we conclude if the decoded plaintext is claimed to be (say) Latin, but its D's turn out to be very different from the typical Latin prose text?  If the decoded text turns out to be "phares genuit hesron hesron genuit ram ram genuit aminadab aminadab genuit nahasson nahasson genuit salmon salmon genuit booz booz genuit obed obed genuit isai isai genuit david" -- can we say that the solution is wrong because this text has weird D1 and D2?  Can we demand that the person proposing the solution provide an explanation for these anomalous statistics?

Well, no; we can only say that, statistically, the plaintext of the VMS happens to be a rathe atypical Latin prose text.   

In fact, if the proposed solution produces this text when fed lines f1r.1-5, it would be hard to argue that the solution is wrong.

But you say:

Quote:In most cases, I would imagine they cannot explain, because their solution is actually not just a simple substitution system but a product of multiple, diverse, and fatally inconsistent decisions to add or remove letters in order to match Voynichese words with their target language, and none of this can be formalised and reproduced.

So the issue is how to refute solutions that postulate a lossy encryption method -- whose decoding must involve guesses.  But looking at the D's (or H's) of the encrypted text will not help.  If the encryption is not a simple one-to-one substitution, these statistics of the encrypted text will be hard to relate to the statistics of the decoded plaintext; and anyway these would not tell us whether the solution is valid or not (see the example above).

I can't think of any mathemagical way to decide whether a non-deterministic decryption method is valid.  The question is basically whether the encryption method is "effectively" invertible: that is, if Alice encrypts a secret -- but meaningful and plausible -- plaintext T in the alleged language, and Bob decodes the encrypted text E, whether he will reliably obtain a text T' that is sufficiently close to T.  But Bob must allowed to chose the guesses any way he can, including You are not allowed to view links. Register or Login to view. or unlimited trial and error until he obtains a text that seems plausible to him.

So I think that the most (only?) reliable way to evaluate a proposed solution is to perform this challenge test.  Statistics seem pretty useless for that purpose.

All the best, --stolfi
tavie Wrote:[Edit:  Getting back on topic here, what I find interesting is how different /sh/ and /ch/ can behave while being so similar.  Despite making similar words like shedy, chedy, sheey, cheey, etc, they definitely have their own "personality" in the manuscript.  /sh/ in particular appears much more in the Top Row of paragraphs than you would expect from its performance in lower lines.  Meanwhile /ch/ tends to go from being a common initial and a rarer word-middle to being a rarer initial and a common word-middle, forming words such as opchey.]

Thank you, Tavie, that’s an excellent point!

In the You are not allowed to view links. Register or Login to view. I linked above, I compared word contacts for chedy/shedy and they appear to overlap so well that I would say there is no clear difference: they have the same adjacent words, with the possible exception of 'daiin'. For comparison, I included data for qokeedy, which is visibly different (e.g. ol.chedy ol.shedy are both frequent, while ol.qokeedy isn’t).
[attachment=16233]
Since we don’t have much Voynichese text, word stats are only possible for very frequent words, but from that experiment I get the impression that the system fails to clearly differentiate ch from sh.

A few years later, Patrick Feaster published results that describe the behavior you mentioned (You are not allowed to view links. Register or Login to view., 2022). He developed a simple and effective way to present glyph distribution in paragraphs, mapping all paragraphs to a square shape where the first line of paragraphs contributes to the top row of the square and the last line of paragraphs contributes to the bottom row (similarly, columns correspond to different left-right positions in lines, with the first and last column corresponding to first and last word in a line respectively).

Patrick included images for both sh-(blue) vs ch-(red), and k(blue) vs t(red). Stats for sh- and ch- are limited to word-start positions. 
This is part of his figure 3: “Rightward and downward distribution of words beginning [Sh] in blue and words beginning [ch] in red (left); words containing [k] in blue and words containing [t] in red” (right).
[attachment=16235]
One can easily see that, as you say, sh- dominates the first first line of paragraphs. The black column on the left is due to the well-known properties of the first word in lines, where peculiar word-types are produced by the apparent prefixing of characters like d- and gallows, so that line-initial benches are almost entirely absent; this equally affects ch- and sh-.
Pink - word-initial bench; green - bench elsewhere:
[attachment=16236]

For k/t, the divide has a diagonal shape: k prefers the bottom-left, t the top-right; in this case the first column is markedly different, with a purple color showing a similar frequency for both glyphs.

In principle, one could think that the vertical preferences of sh- and ch- could correspond to semantic features, like some sh- words being frequent at the start of paragraphs, but if the difference was semantic we would see marked differences in word contacts for sh- and ch-words, and we don’t. Also, sh- markedly prefers the first line only, there is not a gradual vertical decline in its frequency, as I would expect for something due to word meaning.
(29-06-2026, 11:29 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.I insist that this claim holds only for texts of the same type, namely texts that consist mostly of discursive or narrative prose -- like novels, chronicles, philosophical treatises, newspaper and magazine articles, even textbooks.  Like all the books in your example.  

I'm sorry Jorge, but I cannot agree with you, in fact my example uses a very diverse panel of texts.

Enzo Biagi, L'albero dai fiori bianchi, I guess a novel (did not read it), 1994
Italo Calvino, Il Barone rampante, 'fantasy' novel, 1957
Alessandro Baricco, Seta, novel, 1996
Edmondo de Amicis, Cuore, novelettes, ~1880
Ugo Foscolo, Notizie intorno a Didimo Chierico, I don't know, ~1810
Giacomo Leopardi, Canti, poetry, 1817-1836
Giordanio Bruno, La cena delle ceneri, philosophical dialogue, ~1583
Dante, Divina Commedia Inferno, poetry, 1304-1321
Carlo Goldoni, L'avaro, theatrical comedy, 1756
Eugenio Barbarich, La campagna del 1796 in Veneto, military history, 1910
Vittorio Alfieri, Della tirannide, political treatise, 1777
Appunti di anatomia, anonymous, notes from anatomy lessons, ?? 1970-2010 ??
Vittorio Alfieri, Bruto secondo, theatrical tragedy, 1789
Declration of human rights, Italian, ??
L'amore di Loredana, Luciano Zuccoli, novel, 1908

Notwithstanding the temporal span (almost 700 years!), the difference in genre (prose, poetry, theater), the difference in topics, the orthographical differences (I don't need the late medieval Dante here: for instance Italian used to have a diphtong 'uo' [uo] which has, since say ~1920, been replaced by 'o' [ɔ] in most words; or, the letter 'j' was used for the semi-vowel [j] while now the orthography uses 'i'), the lexical differences, the linguistical differences, in spite of all this just the bigrams statistic is enough to positively identify a text as Italian or not.

But yet again: if you feel my panel of texts is not diverse enough and have a different one to propose, in any language, link me the documents (.txt possibly) and I'll check.
(30-06-2026, 01:37 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.my example uses a very diverse panel of texts.

Those texts are "very diverse" only from the literary point of view; they are all very similar from the linguistic one.  They are all running narrative/discursive prose, with long complicated sentences, that try to be varied, elegant, and informative, etc..

For these reasons, they all have a broadly varied vocabulary, but mostly of "core" (not subject-specific) Italian words, and make heavy use of mostly the same function words -- articles, prepositions and conjunctions, pronouns, copulas, auxiliary verbs, etc.  These similarities are not significantly affected by the subject matter, epoch, authorship, or even poetry x prose x theater.

And, again, the frequency distributions D1 and D2 of letters and bigrams are largely determined by the presence or absence of those letters and bigrams in the most common words in the text.  And the character entropies H0 and H1 are just projections of those distributions. 

And, as you note, the Italian (Tuscan) language has changed surprisingly little from the 1300s, and the standard Italian orthography has hardly changed at all.  Thus it is no surprise that D1 and D2 of all those texts cluster tightly together in those character statistics.

That would not necessarily be the case of texts that are not "literary". Like ship logs, accounting books, parish civil registers, catalogs of products, kings,  stars, saits, and churches,  ... or succinct herbals.  Even if they are definitely in Italian.

It also would not be the case if the language went through a significant spelling reform.  Portuguese had one in the 1940s.  Here are three versions of the same novel (Dom Casmurro by Machado de Assis) which are the exactly the same (same words in the same order, thus the same slightly outdated syntax), only in different spellings:
You are not allowed to view links. Register or Login to view. original edition, old spelling.
You are not allowed to view links. Register or Login to view. a recent edition in (almost) current spelling.
You are not allowed to view links. Register or Login to view. a more phonetic transcription of the modern Brazilian reading.
(Beware that these files are in ISO-Latin-1, not Unicode; and this detail matters, because they have accented letters that are incompatible with UTF-8 encoding.)

For comparison, You are not allowed to view links. Register or Login to view. is a Spanish translation (literary, not word-for-word) of that novel.

In my opinion, Portuguese orthography is the second-worst among the languages I know something about.  Worse than the French one, and second only to English.  The reform of the 1940s could have made it at least as phonetic as the Spanish one; but it was half-hearted, and kept many of the worst features of the old script -- like the random and unnecessary use of "ç" for soft "s", the silent word-initial "h", and the four different pronunciations of "x".   But it did remove most doubled consonants, many silent consonants like the "c" in "acto" and "p" in "optimo",  replaced "ph" by "f", and replaced some diphtongs, like "eo" by "eu".  It also added or removed many diacritics.  

So the letter and bigram frequencies of the "1899" and "1999" versions above should be noticeably different.  Maybe not enough to make them "different languages" by your criteria?

The "2099" version is my idea of what a real spelling reform of Portuguese (rather, Brazilian) should have been.  It is as phonetic as I could make it while staying within the Roman alphabet and ISO-Latin accented vowels. The letter and bigram statistics should be quite different from the other two.   I expect that, by your criteria, the differences between "1999" and "2099" will be greater than those between the former and the Spanish translation above.

I still owe you an example of a text in Italian that would not be free prose like those in your list.   Working on that...

All the best, --stolfi
Maybe Libro de arte coquinario would work since it's recipes and such... Library of Congress pdf images here and I imagine there are text or github repos of it in the wild... You are not allowed to view links. Register or Login to view.


(30-06-2026, 04:54 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.I still owe you an example of a text in Italian that would not be free prose like those in your list.   Working on that...

All the best, --stolfi
(30-06-2026, 05:18 PM)JoeyB Wrote: You are not allowed to view links. Register or Login to view.Maybe Libro de arte coquinario would work since it's recipes and such... Library of Congress pdf images here and I imagine there are text or github repos of it in the wild... You are not allowed to view links. Register or Login to view.

Maybe yes, but I cannot process images, I need a textual transcription.
(30-06-2026, 09:28 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.
(30-06-2026, 05:18 PM)JoeyB Wrote: You are not allowed to view links. Register or Login to view.Maybe Libro de arte coquinario would work since it's recipes and such... Library of Congress pdf images here and I imagine there are text or github repos of it in the wild... You are not allowed to view links. Register or Login to view.

Maybe yes, but I cannot process images, I need a textual transcription.

This appears to be another edition of it or at least a big section of it transcribed in plaintext: You are not allowed to view links. Register or Login to view.

I'm not familiar with this edition, sorry.
(30-06-2026, 04:54 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.I still owe you an example of a text in Italian that would not be free prose like those in your list.

You are not allowed to view links. Register or Login to view. is a list of the 366 Catholic patron saints of the day in Italian.  The file main.txt = main.src has one saint per line. The file main.wds has one word per line; use only lines that begin with 'a '.  The character '°' denotes abbreviation or truncation, as in "Sant°"  The character '~'  is hyphen.  Both files are in Unicode UTF-8 encoding.  You should download them; opening in the browser and copy-pasting the contents will not work.

This file is too terse; more succinct than what the VMS parags are likely to be.  It would be better if there was a bit more information about each saint, like "martire" and "secolo XV" and what is its "specialty" (like, here in Brazil Santo Antônio is patron of marriages and the saint to pray to for help finding a misplaced or lost item.)

And I really should tell the story of Santo Expedito, but don't have the time now.

All the best, --stolfi
Pages: 1 2 3 4 5 6