(4 hours ago)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.@ Stolfi, I already replied to you in your chat
Sorry. I saw your post there, but I though you had not really replied to my point. Maybe I did not read it with due care.
Quote:Quote:Me: Character statistics are a byproduct of the most common words.
... The top row consists of words that appear exactly once in the entire manuscript: 543 with “sh,” 1,263 with “ch,” and 730 with “k.” A word that appears only once cannot have a collocation. There’s no pattern like “this is” for a word that exists only once! ... A word-pair model has no place here.
First, many hapax words in the VMS are probably malformed, truncated, or joined versions of words that appear multiple times. So, for example, if "shedy qokedy" is a common word pair, you should see also many pairs like "shoedy qokedy", "shddy qokedy", "shshedy qokedy", and so on -- where the first words of those pairs are hapaxes.
You could try to investigate this possibility by separating "good" words that fit one of the various word models out there from "bad" words that do not fit. IIRC, when I tested my own word model, I found that something like 5-7% of the tokens failed it in some way.
Second, valid hapax words too definitely can have attraction or repulsion for qo-words and contribute to the apparent sh-for-qo attraction.
For example, in English, depending on the topic, it may well happen that the majority of the hapax words are nouns ("chamomille", "shingles", "baldness", etc.), while the verb words are few and mostly occur many times ("cures", "drink", "boil", etc). I am guessing that verbs will rarely be followed by "is", while that will be often the case for nouns. If it happens that the most common verbs do
not begin with with (say) "p", the result will be an
apparent attraction of "p"-words to "i"-words,
including among the hapaxes -- which is actually due to
repulsion of a few common non-hapax non-"p" words by "i"-words.
Quote:The crazy thing is: The ratio of “sh” to “ch” remains nearly the same—between 1.7 and 2.0—regardless of the number of words, that is, in every row, from the one-time words all the way up to the most frequent ones. That is the question your model must answer: Which word pairs generate this, and the top row?
Well, I think that it is you who should answer that question.
But seriously: the first thing I would do is to plot the frequency of each sh
XXX word against the frequency of the corresponding ch
XXX word, for each EVA string
XXX, and see whether those dots -- at least those with higher frequencies -- fall close to a straight line. (A log-log plot would be better, I think)
If they mostly do, it would be evidence (not proof, of course) that the ch-words are equivalent variants of the corresponding sh-words. That is, there is a certain probability that a sh-word will lose the plume and become a ch-word, or vice-versa. Various processes could have this effect: confusion of similar sounds, sandhi, transcription errors, ink fading, poor handwriting on the draft...
(A unique English text that should be used more often in comparisons is the diary of the Lewis and Clark Expedition. Unlike most books and medieval manuscripts, it was written by two Army officers who had gone through the Academy but obviously were not much into literary virtues. So the text has tons of spelling errors and inconsistencies, not only on words of Native origin (I think I saw four different spellings of "buffalo", some in the same parag) but even in very common words like "to-day" and "behaveing" and "prisnair". Some of the errors seem to be stochastic but systematic -- that is, endings like "-ance" and "-ence", "-city" and "-sity" seem to have been often swapped independently of the word.)
If the ch-words are mostly variant spellings of the sh-words, then the two variants should show similar attraction or repulsion for each qo-word (or any other word) that follows it.
On the other hand, if the points on that ch/sh plot do not lie on a straight line, it would be evidence that a ch
XXX word is, at least some of the time,
not a variant of the corresponding sh
XXX word. Then I would expect that the two will have different attraction or repulsion numbers. If they behave the same way in that regard, then I would have to scratch my head harder...
Quote:shor and chor: same ending, same length, different initial sound (excluding “e”). In 29% of cases, “shor” is followed by an “sh” word; in 8.4% of cases, “chor” is followed by an “sh” word. In 28,4% of cases, “chor” is followed by a “ch” word; in 9,2% of cases, “shor” is followed by a “ch” word.
These numbers seem to be consistent with the second possibility above.
Quote:to explain this using word pairs, you need a separate collocation history for each pair, and all of these histories would have to happen to coincide in such a way that the initial sound of the preceding word is the same.
The word pairs do not have to
all agree. The "attraction" that you see is the
sum of individual word attractions, so it is dominated by a few high-frequency words.
Quote:I'm not really concerned with sh and qo here. What matters to me is which rules are tied to a word's form and which to its frequency; the former say something about the cipher, the latter about the text.
Well, again, I am still not convinced that those numbers are due to attraction between
characters or digraphs, rather than words.
Quote:You’re right: a single finding of this kind proves nothing.
It is not that, but rather that this kind of result is unlikely to lead to any useful insight into the Voynichese "grammar".
As I wrote before, trying to decipher a text by computing character and digraph statistics is like trying to understand the ecology of a forest by counting how many animals have tails, how many are brown, and how many have antennas. Each count would mix antelopes and armadillos, bears and beetles, cats and caterpillars, ..., weasels and woodpeckers, ... Those counts may perhaps suggest that the Black Forest is somehow different from the Amazon, and that something changes between summer and winder; but they would never lead to any real understanding of the ecology.
All the best, --stolfi