JoJo_Jost > 01-07-2026, 06:12 AM
Mauro > 01-07-2026, 07:05 PM
(01-07-2026, 04:14 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.You are not allowed to view links. Register or Login to view. is a list of the 366 Catholic patron saints of the day in Italian. The file main.txt = main.src has one saint per line. The file main.wds has one word per line; use only lines that begin with 'a '. The character '°' denotes abbreviation or truncation, as in "Sant°" The character '~' is hyphen. Both files are in Unicode UTF-8 encoding. You should download them; opening in the browser and copy-pasting the contents will not work.
This file is too terse; more succinct than what the VMS parags are likely to be. It would be better if there was a bit more information about each saint, like "martire" and "secolo XV" and what is its "specialty" (like, here in Brazil Santo Antônio is patron of marriages and the saint to pray to for help finding a misplaced or lost item.)
And I really should tell the story of Santo Expedito, but don't have the time now.
All the best, --stolfi
. It's also short, 1812 words, which by itself tends to give unreliable results (~10K words would be better). I removed/replaced the non-italian characters, i.e. Gesù -> Gesù.Jorge_Stolfi > 01-07-2026, 08:14 PM
(01-07-2026, 07:05 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.Well, a 'deviant' text if ever there was oneAgreed. I am still looking for something better.... It's also short, 1812 words, which by itself tends to give unreliable results (~10K words would be better).
Quote:I removed/replaced the non-italian characters, i.e. Gesù -> Gesù.
Quote:However it's also true that the most similar texts according to the bigrams distribution are all Italian texts:
Quote:I think this demonstrates that at least the bigrams statistic is mostly determined by the underlying language and not by the content of the text (barring even more extreme cases). I will add: the vocabulary of the list is extremely aberrant, in practice only variations of the word 'Saint', names of months, numbers from primo (first) to trentuno (thirty-one) and personal names/toponyms.
Mauro > 01-07-2026, 09:33 PM
(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view. Beware that there are saint names with non-Italian characters, namely "ñ" ("San Raimondo di Peñafort.") and "á" ("San Josemaria Escrivá"). Removing those would be cheating...

(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.If I read the first table correctly, it says that the distance d(L,I) from the List to the Italian texts is about the same as the distance d(C,I) from the latter to Caesar's book; but the distance d(L,C) is a bit bigger that both.
(01-07-2026, 08:14 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.But that is my point. This List is "aberrant" because it is not a discursive/narrative text, like all the others. It has a very different nature, purpose, and structure, that causes its word frequency distribution to be extremely different from that of those texts; and that in turn causes the bigram frequencies to be very different too. And yet the language is definitely Italian, and even mostof the words (even most of the names) are specifically Italian.
MarcoP > 02-07-2026, 07:15 AM
(01-07-2026, 06:12 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.This might be of interest, since I’ve been working on this and had just put it together in connection with von Stolfi’s thread.
You are not allowed to view links. Register or Login to view.
The graphic shows how often a “qo” follows each glyph cluster (not token!)
It’s clearly evident that the glyphs k / t and ch / sh before glyphs of the E-family are followed by “qo” with varying frequency before glyphs of the E-family.
![]()
The glyphs have, so to speak, a long-range effect.
Another interesting aspect of this graph is that they are always arranged in the same order within their respective groups.
That is, t < k < ch < sh
So t always has the weakest attraction to “qo,” followed by k, ch, and sh, which always have the strongest.
JoJo_Jost > 02-07-2026, 03:55 PM
I hope to post the explanation in my thread in the next few days. Grove > 02-07-2026, 04:21 PM
(02-07-2026, 03:55 PM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.-----
But there’s another piece of “evidence” that, for example, t and k aren’t interchangeable:
k before aiin (20.6 percent) (n=810)
t before aiin (9.0 percent) (n=356)
k before ain (38.1 percent) (n=645)
t before ain (12.4 percent) (n=210)
"t" seems to occur significantly less frequently with the “aiin” family than "k". "t" is more likely to appear with ‘o’ and "ch."
Have Fun
Jost
JoJo_Jost > 02-07-2026, 04:35 PM
(02-07-2026, 04:21 PM)Grove Wrote: You are not allowed to view links. Register or Login to view.Except maybe if when the line starts with taiin and has tripled repeats using aiin:
<f40r.P.9;F> taiin.ol.olaiin.or-{plant}dain.okaiin.okaiin.okaiin.daram-
#
<f86v3.P1.3;H> taiin.ytaiin.ytaiin.ytaiin.or.ar.ytar.am-
Grove > 02-07-2026, 07:08 PM