(12-09-2026, 10:22 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.In this podcast Lisa Davis makes several claims about the Voynich text. ...
I agree with several of your criticisms. (And I do not think they are "AI generated"). But I disagree in others.
1. The Voynichese alphabet is "too small for a natural language" only if one assumes that each distinct phoneme is encoded as a distinct single glyph. But this need not be true. Even in English several phonemes are encoded as digraphs, and a single letter or digraph can represent several very distinct phonemes.
The structure of the Voynichese words suggests that
e is a modifier for the preceding glyph, so that
Ch and
Che are different "letters", and so probably are
k and
ke,
ee and
eee,
in and
iin and
iiin. If that possibility is taken into account, the size of the alphabet is easily 20 or more.
2 and 3. If anything, the VMS word structure and statistics
exclude all "European" languages, all Semitic languages, and probably all the languages of Africa, West and Central Asa, Oceania, and the Americas. But they are quite compatible with the monosyllabic languages of East Asia, such as Sino-Tibetan, Vietnamese, etc.
4. The methods used to identify consonants and vowels (like Sukhotin's algorithm, that Guy unsuccessfully tried) are based on the principle that Cs and Vs tend to alternate
in polysyllabic words. But the structure of Voynichese words is definitely not compatible with multiple syllables.
5. Indeed the fact that the word frequency distribution of a text follows Zipf's Law is just a necessary but not sufficient condition for it being natural language, either unencrypted or encrypted one-to-one on words -- that is, in such a way that each word type is mapped to a single word type.
Thus the fact that Voynichese follows the law
does let us rule out is certain ciphers like Vigenère, that are not one-to-one on words. That fact also lets us rule out
some methods to generate random gibberish, like Rugg's "grille" method". But I suspect it would also exclude the "copy and mutate" method, because obtaining a Zipfian distribution with it seems to require a rather nontrivial "mutate" procedure.
6. If the text is a natural language, unencrypted or encrypted one-to-one on words, we expect that each section will have a significantly different word frequency distribution, with rather abrupt changes across section boundaries. Whereas gibberish generated by a random method (including "copy and mutate") would
probably generate either the same word distribution through the whole book, or a gradually drifting one, oblivious of section boundaries.
This has always been a very popular topic of statistical research, but the results seem hard to interpret, because the historical scrambling of the bifolios can both homogenize the distribution within each section and create abrupt transitions between sections.
7. Full-word repeats are indeed rare in the "usual suspect" languages. But not in East Asian languages -- especially in imperfect transcriptions that may omit tones or other phonetic details.
The fact that "the most experienced NSA researchers in the history of Voynich studies" have concluded that the word repeats rule out natural language only reminds us that "all the experts" on anything can often be quite wrong.
(Ironically, someone recently criticized the Chinese theory by pointing out that the Starred Parags section has
too few repeated sequences compared to the Shennong Bencao...)
All the best, --stolfi