@ Stolfi, I already replied to you in your chat, but sure, I’m happy to do it again here:
Your argument is: Character statistics are a byproduct of the most common words. “th” pulls “i” because “this is” and “that is” are common. That’s not wrong, but it doesn’t account for the effect of “sh” / “ch” (and others).
To demonstrate this, I’ve separated the occurrences by word and by the frequency of each word.
The bottom row shows the effect for words that occur more than 100 times. There, you might be right—those could be collocations: “shedy” before “qokeedy” and so on.
BUT: The top row consists of words that appear exactly once in the entire manuscript: 543 with “sh,” 1,263 with “ch,” and 730 with “k.” A word that appears only once cannot have a collocation. There’s no pattern like “this is” for a word that exists only once!
And lo and behold: Nevertheless, 20.8% of the one-time “sh” words are followed by a “qo” word, 10.8% of the one-time “ch” words are followed by a “qo” word, and 11.6% of the one-time “k” words are followed by a “qo” word.
The crazy thing is: The ratio of “sh” to “ch” remains nearly the same—between 1.7 and 2.0—regardless of the number of words, that is, in every row, from the one-time words all the way up to the most frequent ones.
That is the question your model must answer: Which word pairs generate this, and the top row?
Next: Your quick checks are correct: “sho,” “shol,” and “shor” are below the baseline rate.
But look at which “sh” words are at the bottom. The five at the bottom are “sho” 9.6%, “shol” 12.1%, “shor” 12.8%, “shar” 12.9%, and ‘shy’ 14.8%, and none of them have an “e.”
Every “sh” word with an “e” is above the base rate: sheo 22.9, sheor 25.5, sheey 36.8, sheol 37.3, sheody 38.3, shey 42.3, shedy 49.0, sheedy 53.0. And your non-sh words that take qo all have an “e” before the end: olkeedy 51.2, cheedy 47.3, otchedy 45.5, qokeedy 41.9. So the rule is not only “sh takes qo.” It is: a sequence of “e” before the end of a word causes “qo,” and “sh” reinforces this effect compared to “ch,” “k,” and “t” (-shedy 49.5%, -chedy 39.3%, -kedy 33.2%, -tedy 31.0%).
And THAT is a rule about a word’s structure. It applies to words that occur only once just as much as to the most common ones.
A word-pair model has no place here.
Second, since you say we should look at word pairs: here are some.
shor and chor: same ending, same length, different initial sound (excluding “e”). In 29% of cases, “shor” is followed by an “sh” word; in 8.4% of cases, “chor” is followed by an “sh” word. In 28,4% of cases, “chor” is followed by a “ch” word; in 9,2% of cases, “shor” is followed by a “ch” word. Here, too, we see this effect.
I calculated this relative to the line in which the word appears—that is, relative to the proportion of “sh” or “ch” words among the other words in the same line: “shor” → “sh” is 2.2 times the line expectation, “chor” → “sh” is 1.1 times; “chor” → “ch” is 1.5 times the line expectation, “shor” → “ch” is 0.7 times the line expectation. So this is not a line effect.
The same pattern holds for keey/sheey with l (20.5% vs. 3.0%), kain/dain with ch (37,9% vs. 21.3%), taiin/kaiin with o (47.1% vs. 30.9%), and keo/sheo with r (20.4 vs. 0.0%).
You might say these are just different words with different neighbors. But here, too, we see the sh/ch effect—
how can all of this be explained using word pairs??? Moreover, the other examples also show that these effects extend even beyond sh and ch and qo...
The point is, to explain this using word pairs, you need a separate collocation history for each pair, and all of these histories would have to happen to coincide in such a way that the initial sound of the preceding word is the same. When you look at the big picture, that’s really quite far-fetched.
So the question really isn’t whether character statistics are a waste of time.
The question is which word-pair model is supposed to generate such a structure—and as hapax, no less!
---
I'm not really concerned with sh and qo here. What matters to me is which rules are tied to a word's form and which to its frequency; the former say something about the cipher, the latter about the text.
You’re right: a single finding of this kind proves nothing. But multiple findings pointing in the same direction are all we have, and that’s why I weigh them against my own assumption just as I would against any other.
And that’s how one decipher texts, too...