This might be of interest, since I’ve been working on this and had just put it together in connection with von Stolfi’s thread.
You are not allowed to view links.
Register or
Login to view.
The graphic shows how often a “qo” follows each glyph cluster (not token!)
It’s clearly evident that the glyphs k / t and ch / sh before glyphs of the E-family are followed by “qo” with varying frequency before glyphs of the E-family.
[
attachment=16261]
The glyphs have, so to speak, a long-range effect.
Another interesting aspect of this graph is that they are always arranged in the same order within their respective groups.
That is, t < k < ch < sh
So t always has the weakest attraction to “qo,” followed by k, ch, and sh, which always have the strongest.
(01-07-2026, 04:14 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.You are not allowed to view links. Register or Login to view. is a list of the 366 Catholic patron saints of the day in Italian. The file main.txt = main.src has one saint per line. The file main.wds has one word per line; use only lines that begin with 'a '. The character '°' denotes abbreviation or truncation, as in "Sant°" The character '~' is hyphen. Both files are in Unicode UTF-8 encoding. You should download them; opening in the browser and copy-pasting the contents will not work.
This file is too terse; more succinct than what the VMS parags are likely to be. It would be better if there was a bit more information about each saint, like "martire" and "secolo XV" and what is its "specialty" (like, here in Brazil Santo Antônio is patron of marriages and the saint to pray to for help finding a misplaced or lost item.)
And I really should tell the story of Santo Expedito, but don't have the time now.
All the best, --stolfi
Well, a 'deviant' text if ever there was one

. It's also short, 1812 words, which by itself tends to give unreliable results (~10K words would be better). I removed/replaced the non-italian characters, i.e.
Gesù -> Gesù.
Anyway, the first conclusion is the list of Saints is surely statistically aberrant from other Italian texts and the distances do not reach the (empirical) threshold I use to determine if a text is written in a certain language. So this is good for you!
However it's also true that the most similar texts according to the bigrams distribution are all Italian texts:
[attachment=16277]
The table is ordered by distance from the list of Saints. Texts from 2 to 15 are in Italian, 16 and 17 are Latin, 18 is (old) French, from 19 to 23 English, 24 is German (with diacritics removed) and 25 is Classic Greek (in greek characters, so an outlier).
The list of Saints is not particularly close to Italian (yellowish green colour, while all the other texts in the same language make nice green squares), but surely Italian is closest to the list of Saint than all other languages. So even with an extremely 'deviant' text Italian stands out as the closest relationship. I think this demonstrates that at least the bigrams statistic is mostly determined by the underlying language and not by the content of the text (barring even more extreme cases).
I will add: the vocabulary of the list is extremely aberrant, in practice only variations of the word 'Saint', names of months, numbers from primo (first) to trentuno (thirty-one) and personal names/toponyms. And indeed it shows:
[
attachment=16279]
So by this statistic, indeed, the text is unrecognizable.
Note: the texts in this last table are the same as before (excep the Greek one) but they are numbered differently.
(01-07-2026, 07:05 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.Well, a 'deviant' text if ever there was one
. It's also short, 1812 words, which by itself tends to give unreliable results (~10K words would be better).
Agreed. I am still looking for something better...
Quote:I removed/replaced the non-italian characters, i.e. Gesù -> Gesù.
That is is what happens if you open the file in the browser and an copy-paste it from there. Our server for some reason tells the browser that it is ISO-Latin-1, but in fact it is Unicode UTF-8. That is why I advised to download the file instead. That word is indeed "Gesù" in the file.
Beware that there are saint names with non-Italian characters, namely "ñ" ("San Raimondo di Peñafort.") and "á" ("San Josemaria Escrivá"). Removing those would be cheating...
Quote:
However it's also true that the most similar texts according to the bigrams distribution are all Italian texts:
If I read the first table correctly, it says that the distance d(L,I) from the List to the Italian texts is about the same as the distance d(C,I) from the latter to Caesar's book; but the distance d(L,C) is a bit bigger that both.
[
attachment=16280]
Quote:
I think this demonstrates that at least the bigrams statistic is mostly determined by the underlying language and not by the content of the text (barring even more extreme cases). I will add: the vocabulary of the list is extremely aberrant, in practice only variations of the word 'Saint', names of months, numbers from primo (first) to trentuno (thirty-one) and personal names/toponyms.
But that is my point. This List is "aberrant" because it is not a discursive/narrative text, like all the others. It has a very different nature, purpose, and structure, that causes its word frequency distribution to be extremely different from that of those texts; and that in turn causes the bigram frequencies to be very different too. And yet the language is definitely Italian, and even mostof the words (even most of the names) are specifically Italian.
The "moral of the story" is that, when one tries to identify the language of the VMS (or any other mysterious text) by any set of statistics, one should compare it to reference texts that hopefully have the same nature and vocabulary. That of course is hard since we ignore even the general topic of most sections. We can guess that narrative/discursive prose texts may be adequate references for the Bio section, but maybe not so much for the Herbal and Pharma; and probably very inadequate for Astro/Cosmo/Zodiac and the Starred Parags.
All the best, --stolfi
(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view. Beware that there are saint names with non-Italian characters, namely "ñ" ("San Raimondo di Peñafort.") and "á" ("San Josemaria Escrivá"). Removing those would be cheating...
I have to confess I changed Peñafort to Pegnafort (I can never remember the Alt code for ñ, now I have just copied the ñ from your quote), and Escrivá to Escrivà (I had no idea it was a different 'a'). Two characters won't change much, don't ask me to re-do everything again
(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.If I read the first table correctly, it says that the distance d(L,I) from the List to the Italian texts is about the same as the distance d(C,I) from the latter to Caesar's book; but the distance d(L,C) is a bit bigger that both.
I'm not sure if I understand what you mean. Be careful not to confuse Dante - Divina commedia Inferno (text #15), which is in Italian and clusters with Italian texts, and Dante - De vulgaris eloquentia (text #16) which is written by Dante but in Latin and, indeed, clusters with the De Bello Gallico, as shown by the green squares. You can also see how the colour in the first column (or row) of the tables changes from yellow-green to yellow-orange moving from Dante - Divina Commedia down to Dante - De vulgaris eloquentia.
Then yes: the distance from the list of saints to Italian is comparable to the distance between Italian and Latin. Something like this:
[
attachment=16282]
What one can say is that, after Italian, Latin is the language most similar to the list of saints.
(by the way, I noticed now that there is a text in the table which is duplicated, the two texts by Baricco are both the same text).
(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.But that is my point. This List is "aberrant" because it is not a discursive/narrative text, like all the others. It has a very different nature, purpose, and structure, that causes its word frequency distribution to be extremely different from that of those texts; and that in turn causes the bigram frequencies to be very different too. And yet the language is definitely Italian, and even mostof the words (even most of the names) are specifically Italian.
Of course I agree with you that the statistics of a text depend on the text itself. However in the great majority of cases the statistics (well, of bigrams at least, let's stick to them) are influenced more by the language than by the content, and even in texts very divergent from the 'normal texts' (as the saints list) the influence of the language is still dominant. Ie.: I'd bet that a list of English saints will sit closer to Darwin, Joyce and Milton than to the rest, while a list of German saints will be near Tristan (Thomas Mann).
(01-07-2026, 06:12 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.This might be of interest, since I’ve been working on this and had just put it together in connection with von Stolfi’s thread.
You are not allowed to view links. Register or Login to view.
The graphic shows how often a “qo” follows each glyph cluster (not token!)
It’s clearly evident that the glyphs k / t and ch / sh before glyphs of the E-family are followed by “qo” with varying frequency before glyphs of the E-family.
[attachment=16289]
The glyphs have, so to speak, a long-range effect.
Another interesting aspect of this graph is that they are always arranged in the same order within their respective groups.
That is, t < k < ch < sh
So t always has the weakest attraction to “qo,” followed by k, ch, and sh, which always have the strongest.
Thank you, JoJo_Jost!
Sticking to the thread’s topic is a big thing by itself, these days, but you also add new fascinating evidence to the subject.
I have replicated your results with reasonable accuracy, so these numbers appear to be reliable (with the usual small fluctuations due to transliteration, preprocessing details etc).
Here are some unreliable thoughts, I am not sure that they make sense. Your analysis certainly deserves further investigation:
I guess that come of the consistency in your table might be seen as a consequence of “dialects”: e.g. in Q13 “qo” and “-edy” are very frequent, while “eody” is almost entirely absent. Things like these may help understand why the -edy line gets high numbers and the -eody line gets lower numbers.
I don’t think the entire pattern you mention (t < k < ch < sh) can be explained in a similar way. This seems only possible for some individual cases, e.g. the large difference between t and k for -eody: this suffix is typical of Pharma, where k is exceptionally more frequent than t.
Of course, it is not clear that the dialects are “causing” these numbers, e.g. it could be that there is some phenomenon that causes both the apparent different dialects and the consistency of the table.
I also tried to understand if Feaster’s left-right distributions might provide some perspective into this (e.g. sh and qo having similar line-position preferences), but I couldn’t spot any clear pattern.
It’s also interesting that the percentages in your table span a limited range: they do successfully differentiate ch from sh, but we are still talking about preferences and the table confirms that the two benches behave similarly, leading to similar contacts with the following word. The two gallows behave even more similarly to each other, though the small differences follow the interesting arrangement you pointed out.
(02-07-2026, 07:15 AM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.Sticking to the thread’s topic is a big thing by itself, these days,
I fear I'm the main culprit. Please excuse me.
@ MarcoP
In my opinion, this system cannot be explained as a language phenomenon at the token or word level. But this peculiarity is yet another clear indication that the VMS is not based on substitution or even a simple abbreviation system—especially in languages like Latin, Italian, German, and other European languages of the time. Without a more complex cipher, this cannot be explained.
Probably not in Chinese either, but I’m not familiar enough with it....
I have maybe a logical explanation, but I’m not allowed to write about it here

I hope to post the explanation in my thread in the next few days.
-----
But there’s another piece of “evidence” that, for example, t and k aren’t interchangeable:
k before aiin (20.6 percent) (n=810)
t before aiin (9.0 percent) (n=356)
k before ain (38.1 percent) (n=645)
t before ain (12.4 percent) (n=210)
"t" seems to occur significantly less frequently with the “aiin” family than "k". "t" is more likely to appear with ‘o’ and "ch."
Have Fun
Jost
(02-07-2026, 03:55 PM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.-----
But there’s another piece of “evidence” that, for example, t and k aren’t interchangeable:
k before aiin (20.6 percent) (n=810)
t before aiin (9.0 percent) (n=356)
k before ain (38.1 percent) (n=645)
t before ain (12.4 percent) (n=210)
"t" seems to occur significantly less frequently with the “aiin” family than "k". "t" is more likely to appear with ‘o’ and "ch."
Have Fun
Jost
Except maybe if when the line starts with taiin and has tripled repeats using aiin:
<f40r.P.9;F> taiin.ol.olaiin.or-{plant}dain.okaiin.okaiin.okaiin.daram-
#
<f86v3.P1.3;H> taiin.ytaiin.ytaiin.ytaiin.or.ar.ytar.am-
(02-07-2026, 04:21 PM)Grove Wrote: You are not allowed to view links. Register or Login to view.Except maybe if when the line starts with taiin and has tripled repeats using aiin:
<f40r.P.9;F> taiin.ol.olaiin.or-{plant}dain.okaiin.okaiin.okaiin.daram-
#
<f86v3.P1.3;H> taiin.ytaiin.ytaiin.ytaiin.or.ar.ytar.am-
Sorry, i dont undestand what u mean?
I started simply looking at multiple consecutive repeats and since you mentioned taiin isn’t as popular as kaiin - it just amused me that a couple of lines I happened to be looking at began with taiin and subsequently were followed by three consecutive repetitions involving aiin words.
Sorry for the spuriously shared off topic observation that really doesn’t mean much.