08-07-2026, 05:09 AM
@You are not allowed to view links. Register or Login to view.
Since I’m currently working on some analyses that might prove interesting once all the control tests are done, I tested your Core Alphabet to clean up the IVTFF corpus.
The 1-to-1 compression turned out to be very effective in removing the background noise.
Out of curiosity, I ran a script to calculate the Greenberg Synthesis Index on the cleaned text (basically, the algorithm calculates how many "pieces" or morphemes make up an average word). I got a score of 2.86.
If my calculations are correct, that score means the text "could" be compared to a highly agglutinative or even polysynthetic language, where words are built by concatenating numerous small suffixes to one another.
This result made me reflect on your estimate of a 10% error rate. What if lkeede is not a spelling mistake for lkeedy at all? If the language "glues" concepts together, perhaps your core character e and your core character y aren't just simple orthographic variants, but actual grammatical suffixes? For instance, hypothetically speaking, one could indicate the plural, and the other a verb tense.
If we look at the final characters as specific morphological tags rather than transcription errors, maybe the scribe wasn't making mistakes at all.
Furthermore, the hypothesis of an "agglutinative language" would perfectly explain the well-known anomaly of "Word Doubling".
Forgive me if I may, but to me, this data looks very similar to the example you yourself brought up regarding ancient Chinese words, where the repetition "pain pain" means "much pain"… Now, even though I am not a linguist, through a simple web search I noticed that European inflectional languages tend to suppress the immediate repetition of words. But Asian or Austronesian languages (like Chinese or Indonesian) use it constantly to form plurals or to intensify meaning.
To have a counter-proof, I had over 30,000 words from the Indonesian Wikipedia analyzed (removing the hyphens to mimic the raw orthography of the Voynich). The real doubling rate turned out to be 2.06 times higher than chance. This is a value very close to the rate (about 2.2x) that I calculated by running my stochastic models on the IVTFF corpus. In practice, your intuition about Chinese seems to be mathematically supported by the data.
If this were really the case, is it possible that this isn't a text full of distracted scribal errors, but simply a non-European language that makes extensive use of suffixes and reduplication?
I would be very interested to know if you think your alphabet could work by interpreting the word endings as micro-morphemes.
Alfredo
Since I’m currently working on some analyses that might prove interesting once all the control tests are done, I tested your Core Alphabet to clean up the IVTFF corpus.
The 1-to-1 compression turned out to be very effective in removing the background noise.
Out of curiosity, I ran a script to calculate the Greenberg Synthesis Index on the cleaned text (basically, the algorithm calculates how many "pieces" or morphemes make up an average word). I got a score of 2.86.
If my calculations are correct, that score means the text "could" be compared to a highly agglutinative or even polysynthetic language, where words are built by concatenating numerous small suffixes to one another.
This result made me reflect on your estimate of a 10% error rate. What if lkeede is not a spelling mistake for lkeedy at all? If the language "glues" concepts together, perhaps your core character e and your core character y aren't just simple orthographic variants, but actual grammatical suffixes? For instance, hypothetically speaking, one could indicate the plural, and the other a verb tense.
If we look at the final characters as specific morphological tags rather than transcription errors, maybe the scribe wasn't making mistakes at all.
Furthermore, the hypothesis of an "agglutinative language" would perfectly explain the well-known anomaly of "Word Doubling".
Forgive me if I may, but to me, this data looks very similar to the example you yourself brought up regarding ancient Chinese words, where the repetition "pain pain" means "much pain"… Now, even though I am not a linguist, through a simple web search I noticed that European inflectional languages tend to suppress the immediate repetition of words. But Asian or Austronesian languages (like Chinese or Indonesian) use it constantly to form plurals or to intensify meaning.
To have a counter-proof, I had over 30,000 words from the Indonesian Wikipedia analyzed (removing the hyphens to mimic the raw orthography of the Voynich). The real doubling rate turned out to be 2.06 times higher than chance. This is a value very close to the rate (about 2.2x) that I calculated by running my stochastic models on the IVTFF corpus. In practice, your intuition about Chinese seems to be mathematically supported by the data.
If this were really the case, is it possible that this isn't a text full of distracted scribal errors, but simply a non-European language that makes extensive use of suffixes and reduplication?
I would be very interested to know if you think your alphabet could work by interpreting the word endings as micro-morphemes.
Alfredo
. When I talked about prefixes infixes and suffixes I was referring to agglutinative (natural) languages in general, which can use all the three forms, not to the VMS (which I don't even know if it's a language or not).