28-06-2026, 04:40 PM
(28-06-2026, 04:39 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.If it were just a matter of y and qo, but of course it isn't...
Of course it is not! It is expected that in ANY text in ANY language, the statistics of ANY property of a word will be affected by ANY property of the following word. It would be noteworthy if that was not the case.
Here are statistics for the Starred Parags section of the VMS (my transcription, which is slightly different from Rene's current IVT). The first property (P1) is "the last EVA letter of the word", and the second property (P2) is "the first EVA letter of the next word":
4298 -y. | 1313 -y.q- | 143 -l.k- | 100 -y.t-
2095 -n. | 172 -n.q- | 102 -y.k- | 40 -l.t-
1499 -r. | 136 -l.q- | 25 -r.k- | 12 -n.t-
1416 -l. | 92 -r.q- | 16 -o.k- | 10 -r.t-
317 -o. | 49 -d.q- | 10 -n.k- | 6 -m.t-
The first column is the counts of P1 in the whole section. It says that there are 4298 words that end with "-y", 2095 that end with "-n", etc..
The second column is the counts of P1 before words that begin with "q-". It says that, before words that begin with "q-", there are 1313 words that end in "-y", 172 words that end with "-n", and so on. As you noted before, the ratio of "-y" words to "-n" words is ~2:1 in general, but jumps to ~7:1 before "q-" words.
The third column shows counts for P1 when the next word begins with "k-". In this context, words that end with "-y" are scarce, so much that the most common ending is "-l"; and the ratio of "-y" to "-n" is ~10:1.
The last column shows the situation before words that begin with "t-". In this context, the most common ending is still "-y", but the second most common is "-l", and the ratio of "-y" to "-n" is ~8:1.
But it is more interesting to let P1 be the first EVA letter of the first word, and P2 be the last EVA letter of the next word:
2420 o-. | 350 o-.-r | 314 o-.-l | 945 o-.-y | 486 o-.-n
1835 q-. | 246 q-.-r | 231 q-.-l | 829 q-.-y | 371 c-.-n
1747 c-. | 243 c-.-r | 215 c-.-l | 643 c-.-y | 328 q-.-n
961 s-. | 116 a-.-r | 129 a-.-l | 435 s-.-y | 213 s-.-n
841 a-. | 110 l-.-r | 118 s-.-l | 300 a-.-y | 156 a-.-n
767 l-. | 104 s-.-r | 102 l-.-l | 277 l-.-y | 155 l-.-n
483 d-. | 56 d-.-r | 80 d-.-l | 200 d-.-y | 75 y-.-n
348 y-. | 47 p-.-r | 42 k-.-l | 139 y-.-y | 67 d-.-n
q:c= 1.05 | 1.01 | 0.98 | 1.29 | 0.88
The first column here says that in the SPS there are 2420 words that begin with "o-", 1835 words that begin with "q-", 1747 words that begin with "c-", etc.
The second column says that, before words that end with "-r", there are 350 words that begin with "o-", 246 words that begin with "q-", and so on.
The other columns give the analogous initial-letter statistics for words before words that end with "-l", "-y", and "-n".
Note that the ratio of "q-" words to "c-" words in general is 1.05, and it is more or less the same also before "-r" words and "-l" words.
But before words that end with "-y", there is a significant excess of words that begin with "q-" rather than "c-" (ratio 1.29); whereas, before words that end with "-n", the words that begin with "c-" are more common than those that begin with "q-" (q:r ratio 0.88).
So, from these statistics, the influence of the end of a word on the beginning of the previous word does not seem as dramatic as the "-y.qo-" effect. However, given the size of the counts, it seems statistically significant. (There are more dramatic effects before words that end with "-k" or "-s" or "-t", but these counts are small thus the effect may be just sampling noise.)
Are these influences due to some long-range phonological or grammatical property of the final letter that can attract or repel certain letters almost two tokens away? Quite probably not. The "q-.-y" anomaly is probably due to a few common word pairs that happen to be of the form "q***.***y", like "qokedy.okeey", rather than "c***.***y", like "chedy.lkeedy".
All the best, --stolfi

