(27-06-2026, 04:43 PM)pfeaster Wrote: You are not allowed to view links. Register or Login to view.Here's a playful analogy: In a base-10 system, all multiples of five end with the digit 0 or 5. For example, 12345 and 43210 are both multiples of five. My position: Investigating that final digit looks like it might be worthwhile. ...
But note that, from character-level statistics, you will never learn that numbers that end in '5' are divisible by 5. Looking for numbers that end in '5' may be useful only if you already know that fact, and you suspect that quantities divisible by 5 have a special role in that text.
In fact, given a text with lots of numbers, I cannot imagine any useful
insight one could obtain about those numbers from digit statistics alone.
And note also that the statistics for tokens that end in '5' would be contaminated by counts of tokens like "6.25" and "x^5" and "9-to-5" and "10:35" and "You are not allowed to view links.
Register or
Login to view.".
(By the way, I intended to add to that list a YouTube link ending with '5'. But after scanning a couple hundred items my list of funny videos in vain, I began to suspect that no such thing exists. And indeed Goggle AI explained that YouTube assigns a 64 bit ID number to each video, then encodes that with 11 digits in base-64; and since 11 x 6 = 66, the last two bits of this code will be zero. So there are only 16 possibilities instead of 64 for last base-64 character in the code; and the only digits in that set are 0, 4, or 8.
So I guess that this would be an example of a puzzling fact that would stand out in character-based statistics ... but would hardly ever be explained by them...)
Quote:And another playful analogy: Consider a hypothetical text of unknown character in an unknown script and language. The characters found most often at the ends of words in general are @ (12%), $ (10%), # (8%), & (7%), and ^ (6%). But if we limit our study to the words that appear before one particular word -- ~+*/% -- the character most often found at the end of those words is, by far, ^ (60%), although no specific word or words predominate among them; we just find a lot of different words that end in ^.
My position: Golly, that's remarkable! What kinds of explanation(s) could we find that would be consistent with a pattern like that? Could this give us a clue as to the structure of a language? Or the mechanism of a cipher?
Your position: That statistic alone tells us nothing. Not only that; the type of statistic is all wrong as well. We instead need to be looking at the words that most frequently precede ~+*/%. Anything else is just going to cause confusion.
Yes, and I stand by the latter statement. What insight could we possibly get out of the observation "[^] is more common than usual before [~+*/%]"?
To make further progress, we must look at the
words that have different frequency before [~+*/%]. We will probably find that not just words that end with "^", but other words as well are "attracted" to [~+*/%], and on the other hand some words that end with "^" are "repelled" by it. Then we will realize that the important feature is not "ends with [^]", but may be something else, like "it is a transitive verb" or "it is a word related to horse management"; and the fact that many of those words end with [^] is a mostly meaningless coincidence. In any case, we would have identified a set of words that are somewhat related, possibly semantically or syntactically.
Quote:If we were trying to "decipher" Culpeper (and had no knowledge of the English language), the observation that words ending in [s] are so much more common than usual before [are] would actually be a useful crib in practice -- more useful, I'd say, than any statistics about the specific words that precede [are]. It may reveal nothing in itself, I suppose, but for someone seeking a linguistic solution, I believe it would suggest the hypothesis that [s] is a morphological marker that correlates somehow with the word [are].
Yes, we could make the hypothesis that "-s is a morphological marker that correlates with [are]". But that hypothesis in fact is nothing more than the observation "the word-final -s is more common than usual before the word [are]". It is just stated in a more sophisticated way...
Quote: But noticing the pattern (and pondering possible explanations for it) would be a step in the right direction, and enough steps in the right direction might eventually lead to a solution.
Not necessarily. Character statistics may be just as likely to lead you astray. For instance, in my earlier post I noted that the word final "-s" is also 3 times more common before "and" than in Culpeper's book as a whole. Should we take that as a hint that "and" and "are" have similar meanings or syntactic functions? That would be a mistake, because those numbers for "and" are due mostly to one word ("vertues") and one peculiarity of the structure of the text (that it has the phrase "vertues and uses" in every recipe).
Quote:On the other hand, with an analytical/isolating language like Chinese, I suppose no such clues should exist -- which I suspect may be why you're eager to rule them out in the case of the Voynich Manuscript.
On the contrary. In any text in any language, and for any pair of symbols X and Y, we expect that the frequency of word-final "-X" before words starting with "Y-" will be either higher or lower than the frequency of word-final "-X" in the whole text. That happens because the frequency of word pairs is never just the product of the individual frequencies.
Check You are not allowed to view links.
Register or
Login to view. in modern Mandarin pinyin, for example. (This is the older version I was using six months ago, not the new version that I am still extracting from the Zhenghe Bencao. You must download the file; opening it in the browser and copy-pasting will not work because of UTF-8/ISO-Latin screwup by our WWW server.) Consider these counts:
2676 g. 273 n.g 304 ī.m 201 n.w 511 g.s 103 n.p 436 ǔ.z
2235 n. 85 i.g 109 g.m 122 g.w 204 n.s 62 g.p 340 g.z
1297 ǔ. 61 g.g 62 n.m 66 ì.w 161 ì.s 49 ǔ.p 305 n.z
1276 ì. 35 ì.g 26 ù.m 54 ǔ.w 67 i.s 22 ì.p 94 ú.z
823 i. 28 ú.g 24 ì.m 48 í.w 63 o.s 13 è.p 86 ì.z
615 ī. 18 ā.g 14 ǔ.m 46 o.w 56 ù.s 12 i.p 64 í.z
569 ú. 13 è.g 13 i.m 37 ǐ.w 46 è.s 10 ú.p 62 ù.z
568 ù. 13 ù.g 13 r.m 33 i.w 45 ǔ.s 6 ù.p 50 ǐ.z
479 è. 10 é.g 12 è.m 25 ù.w 44 ú.s 4 ǐ.p 33 i.z
413 ǐ. 9 í.g 7 o.m 24 ī.w 41 í.s 4 o.p 31 ǚ.z
The first column above is the count of words that end with each letter (10 most common only). That is, there are 2676 words that end with "-g" (actually "-ng") , 2235 that end in "-n", 1297 that end with "-ǔ", etc. The second column is the same counts but only for words that occur before a word that begins with "g-". That is, there are 273 word spaces where "-n" faces "g-", 85 where "-i" faces "g-", etc. The other columns are similar for word breaks before "m-" words, "w-" words, etc.
As you can see, there are plenty of strong "anomalous final-letter frequencies", comparable to the anomaly of "-y" before "qo-" in the VMS. In the whole text, the most common word ending is "-g" followed closely by "-n", then "-ǔ" and "-ì" almost tied in distant 3rd and 4rth places. But before a word that begins with "g-", the ending "-n" is more than three times as common as "-ì" and four times as common as "-g". Whereas before "s-" the ending "-g" is 2.5 times as common as "-n". And before "m-" the most common ending is "-ì" (3x the second place), and before "z-" the most common ending is "-ǔ".
The causes of these anomalies are not some phonetic phenomenon of the language or peculiarity of the encoding that causes certain
characters to attract or repel. Like the "-s before and" anomaly of Culpeper, these anomalies are due to certain
words being more common or less common before certain other
words, in that text.
The high frequency of "-ǔ before z-", for example, is due to the common pair "zhǔ zhì" 主治, which, in that version, is what I have identified as the source of
daiin in the Starred Parags section.
Quote:this researcher you posit who gets "forever stuck" strikes me as rather a naive kind of straw man.
It is not a straw man! How many people have spent years tabulating those character frequencies, without getting any useful conclusion out of them?
Quote:I could equally posit a researcher who somehow identifies and collects hundreds of English plurals without ever noticing that they tend overwhelmingly to end in [s] -- remaining convinced a priori that "statistics about characters tell us nothing."
But the statistics for "-s" endings did
not lead to the identification of "plurals" being what what was attracted by "are"! We figured that out that from our knowledge of English grammar. Those "-s" statistics indeed told us nothing besides the statistics themselves.
Quote:If you could have an equally revealing data point about Voynichese -- no more revealing, no less revealing -- would you want to have it?
Well, we have the statistics "-y is more common before qo-" and variations on that. What do they reveal?
All the best, --stolfi