(08-09-2026, 09:50 PM)LisaFaginDavis Wrote: You are not allowed to view links. Register or Login to view.This was a really fun interview! You are not allowed to view links. Register or Login to view.
In this podcast Lisa Davis makes several claims about the Voynich text. I would like to ask for the published sources supporting these claims, since I'm only aware of published research arguing against them.
1. Davis states that the Voynich alphabet is "in the Goldilocks Zone" for an alphabet of "about 25 to 30 symbols." However, Koen has shown in his analysis that the functional alphabet reduces to about 13 characters after removing positional variants. (see You are not allowed to view links.
Register or
Login to view.) And Claire Bowern wrote about the same question: "The Voynich Manuscript is highly unusual and non-language-like at the character level." [You are not allowed to view links.
Register or
Login to view.]. Is there a published analysis arguing that the functional alphabet is large enough for a natural language?
2. Davis claims that "it's almost certainly not Asian" and "it's almost certainly not African." What published analysis supports the exclusion of entire continental language groups? "African" is not a language family — Africa contains at least four unrelated language families (Afro-Asiatic, Niger-Congo, Nilo-Saharan, Khoisan). Moreover Asian languages like Chinese and Vietnamese and Semitic languages like Arabic and Hebrew were frequently named as possible languages.
For instance Reddy and Knight wrote in 2011: "A more likely explanation is that the script is an abjad, like the scripts of Semitic languages, where all or most vowels are omitted. Indeed, we find that a 2-state HMM on Arabic without diacritics and English without vowels learns a similar grammar, a∗b+." [You are not allowed to view links.
Register or
Login to view., p. 80]
3. According to Davis "it seems to be most closely related to some kind of Indo-European language." What published analysis supports this claim? Lindemann and Bowern wrote about this question in 2021: "In his 1976 book on computational applications to scientific and engineering problems, Yale physicist William Bennett Jr. used a transcription of Voynichese to illustrate the concept of information entropy in language and its application to cryptography. He found the conditional character entropy of Voynichese to be surprisingly low compared to a sample of European plain texts and ciphers. This means that Voynichese characters are unusually predictable compared to most European languages" [You are not allowed to view links.
Register or
Login to view., p. 22].
4. Davis further argues that "there's a linguist who's doing work trying to figure out which are the vowels and which are the consonants." However, Guy (1991) and Reddy & Knight (2011) both attempted vowel-consonant separation and found it does not work. Reddy & Knight wrote "We find that a curious phenomenon occurs with the VMS – the last character of every word is generated by one of the HMM states, and all other characters by another; i.e., the word grammar is a∗b. There are a few possible interpretations of this. It is possible that the vowels from every word are removed and placed at the end of the word, but this means that even long words have only one vowel, which is unlikely." [You are not allowed to view links.
Register or
Login to view., p. 80] Is there a more recent analysis that contradicts these findings?
5. Davis further states that Voynichese follows Zipf's Law "perfectly" and that "that suggests it's not nonsense... it's an actual natural human language." However, Claire Bowern wrote in 2021: "The fact that Voynich word frequencies follow a Zipfian distribution does not prove that the text is linguistically meaningful" and cited Reddy & Knight's characterization of Zipf's Law as "a necessary (though not sufficient) test of linguistic plausibility." [You are not allowed to view links.
Register or
Login to view.] Moreover, Zipf's Law is also produced by copying processes (see You are not allowed to view links.
Register or
Login to view.) and does not distinguish between meaningful and meaningless text.
6. Davis concludes that the LSA results show "groups that reflect these illustrative topics" and that "there seems to be an actual relationship between the illustrative topics and the actual text." However, the published paper of Layfield and Davis states that "LSA analytics functions by identifying patterns and context, not semantic meaning" (You are not allowed to view links.
Register or
Login to view.). Moreover, the same paper shows that generated meaningless text produces "the same relative shape" as the VMS under LSA.
7. Because of the predictability of Voynichese Davis states that "it's not random, it's not randomly generated text, it's not nonsense." However, You are not allowed to view links.
Register or
Login to view. (2022) wrote: "We cannot and do not attempt here to prove that the VMS is gibberish. Our results do, however, invalidate traditional arguments that the small-scale structure of Voynichese is too non-random to be meaningless." Bowern explicitly leaves the possibility of meaningless text open and states it would be "not surprising" if the VMS were gibberish that is "both language-like and statistically unusual."
Additionally, Tiltman (1967, p. 9) wrote: "Languages simply do not behave in this way." Currier (1976) responded to the question: How do you account for the full-word repeats? with "That's just the point they're not words!" and wrote further "I can think of no linguistic explanation for this sort of phenomenon, not if we are dealing with words or phrases, or the syntax of a language where suffixes are present." And D'Imperio (1978, p. 30) summarized: "The text just doesn't act like natural language." These are not fringe positions — they represent the conclusions of the most experienced NSA researchers in the history of Voynich studies.