Jorge_Stolfi > Yesterday, 03:56 PM
(Yesterday, 01:20 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.What hieroglyph might this be similar to?
ololololo > Yesterday, 06:14 PM
Koen G > Yesterday, 08:20 PM
(Yesterday, 03:56 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.I don't know. Like for the other two, there are many Chinese characters that could have been the source for this one...
Jorge_Stolfi > Today, 12:59 AM
(Yesterday, 08:20 PM)Koen G Wrote: You are not allowed to view links. Register or Login to view.Flipped upside down like in your image, the bottom glyph looks exactly like red-initial T was written in some MSS. Bird glyph has been found as red-capital V. The "smoke" thing is like Z. That gives you the last chunk of an alphabetic list, with y missing: t - (uvw) - z
Koen G > 7 hours ago
chenxiang > 5 hours ago
(25-09-2026, 10:25 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(24-09-2026, 04:34 PM)chenxiang Wrote: You are not allowed to view links. Register or Login to view.First, regarding your Pinyin data: it is true that adjacent syllables in Chinese show statistical tendencies (e.g., a syllable ending in 'g' is often followed by one starting with 's'). However, this is a phenomenon of adjacent phonetic assimilation and high-frequency word collocations (词汇搭配), not a grammatical or structural dependency.
In Chinese, the influence of a syllable's final sound strictly applies to the immediately following syllable within the same phrase. It does not extend across word boundaries to mechanically govern the distant structure of an unrelated word ... not a "long-distance effect." In the VMS, as JoJo_Jost and others have shown, the sh/ch at the end of a word strongly affects the qo at the beginning of the next word, ... This kind of mechanical, cross-boundary, long-range statistical dependency is not a feature of natural Chinese syntax or phonology.
What you are discussing is sandhi effects on vowel and consonant quality, not just tone. Yes, sandhi in general, being a phonological process, acts only on the adjacent phonemes of consecutive words.
But sandhi is not the only cause of statistical correlations between phonemes (or letters) in conscecutive words! In fact it is the least important cause. As I explained in my replies to @JoJo_Jost, the main cause is correlations between consecutive words, especially very common ones, that happen to use those characters; and these correlations in turn are due mostly to semantics and to the grammar -- not to the phonemes (or letters) in question.
For instance, in You are not allowed to view links. Register or Login to view. there are 60667 words. Of them, 4157 (6.9%) are i-words -- words that begin with the letter "i". Therefore, if the words were chosen at random, independently from the adjacent words, we expect that any given word, or any given set of words, will be followed by an i-word in 6.9% of the cases.
But if we look at the th-words (words that begin with "th") in that text we find that only 4.6% of them (361 out of 7857) are followed by an i-word.
Should we say then that "there is a long-range repulsion between the "th" at the beginning of a word and the "i" at the beginning of the next word (or vice-versa)"? (And therefore "That text cannot be natural English"? ...?!?)
No. While that statement is true (if we take "repulsion" to mean just "negative correlation"), it is misleading because it suggest that it is the letters "th" that repel the letter "i". But if we look at the individual th-words, this is what we see:
ANY 4167/60667 6.87%
th... 361/7857 4.59%
that 98/793 12.36%
this 23/250 9.20%
the 173/4764 3.63%
then 30/171 17.54%
they 3/342 0.88%
them 11/187 5.88%
Here the notation "M/N P%" means that M of the N occurrences of the word at left, that is P percent of those occurrences, are followed by an i-word.
So we can see that, while there is a definite "i-word repulsion" for some words, including "the", "they", and "them" (which all have P lower than the expected 6.87%), for some other words, including "that", "this" and "then", there is "i-word attraction" instead.
The reason we see a repulsion for th-words as a whole is that their collective P (4.59%) is the average of the Ps of the indvidual th-words, weighted by the frequency of the word; and the th-words with repulsion, like "the", end up dominating this average.
And the same applies to the other side of the word gap. The word "that" is not attracted to i-words in general, but to specific words that happen to begin with "i". Namely
52 that i
22 that it
9 that in
6 that is
1 that idea
1 that if
1 that ill
1 that incredible
1 that intelligent
1 that its
1 that indeed
There are many non-i-words that can follow "that". But it so happens that, on the whole, those pairs with an i-word on the right account for 12.36% of all occurrences of "that", which happen to be more than the frequency of i-words in general.
As you can see, the apparent attraction of "that" for i-words is not a property of the letter "i", but partly a consequence of the grammar (that restricts what follows to certain types of sentences, which happen to often begin with "it", "its", "in", "is", and "if") and mostly due to the semantics and the contents of the text (a first-person narrative, where the subject of many sentences is the pronoun "I").
And the same happens in the VMS. The apparent attraction between Sh-words and the following qo-words is not a property of the characters Sh and qo. It is due to certain words that happen to start with Sh being "attracted" to certain words that happen to start with "qo"; and these cases happening to dominate when the attractions and repulsions are added together.
And anomalies of this sort happen in any meaningful text in any natural language, for the same reasons.
They happen in the Shennong Bencao Jing (SBJ) too. You are not allowed to view links. Register or Login to view. is a digital version of it that I obtained from the net and converted to Mandarin pinyin via Google Translate, several months ago. (The original file had many erros, and differed in many details from the SBJ quoted in the Zhenghe Bencao. My edits and conversion to pinyin must have added many more errros. But the errors should not have a significant impact on what follows.)
Instead of looking for "influences" between adjacent words, let's look at attraction or repulsion between certain initial consonants in words that are two positions apart -- that is, between word t[k] and word t[k+2]. That hopefully will make the results independent of sandhi effects.
You are not allowed to view links. Register or Login to view. of counts and frequencies for the initial consonant pairs of words (sylables) t[k] and t[k+2] from that file.
In the bottom "TT" (totals) row, you can see that 568 out the 12537 syllables t[k+2] (4.5%) begin with the pinyin consonant "g". Therefore, if syllables were generated at random, without regards for the context, we would expect that, for any given word X (or any set of words X), about 4.5% of its occurrences t[k] would be followed, two words ahead, by some word t[k+2] that starts with "g".
But the table shows that, in fact, 254 (21.9%) of all the 1160 tokens t[k] that start with "sh" are followed at t[k+2] by tokens that starts with "g". That is, the initial consonant "sh" strongly attracts the initial consonant "g" even through the intervening word!
On the other hand, only 15 (1.8%) of all 831 tokens t[k] that start with "x" are followed by tokens t[k+2] that start with "g". That is, the initial "x" repels the initial "g" even across the intervening word!
But of course these statements are misleading. What is happening is that there are certain triples of words X Y Z that occur more often than expected by chance, where X is a sh-word and Z is a g-word:
123 shēng shān gǔ
100 shēng chuān gǔ
6 shí wèi gān
2 shēng píng gǔ
1 shēng gāo gǔ
1 shēn wèi gān
1 shāng kǒu gān
1 shāng bǔ gǔ
3 shān shān gǔ
1 shān píng gǔ
3 shān chuān gǔ
1 shā xié guǐ
1 shā wèi gān
1 shā chóng gǔ
1 shā bǎi guǐ
1 shí zhū guǒ
1 shí zhū guān
1 shí zhī gè
1 shí shā gǔ
1 shí chuāng gēn
1 shé shì gǔ
1 shuǐ shā gǔ
1 shuǐ dào gēn
That is, the "attraction" is not between the consonants "sh" and "g", but between the word "shēng" and (two words ahead) the word "gǔ"
Thus the Sh-qo "attraction" reported by @JoJo_Jost is NOT evidence that the VMS is an artificial text generated by algorithm. Quite the opposite, it is another feature that is shared with natural language texts, and is NOT expected in a text that is generated word by word by some manual random process.
And here I should repeat two bits of advice that I posted recently to the net:
The discussion above justifies advice #1: any character- or n-gram count is the sum of word or n-word counts for many words that probably have very different grammatical functions and by meanings, whose frequencies are therefore due to many unrelated causes. A character-based statistic is like counting the number animals in a forest with tail and without tail. The number would be useless because it will mix ants and beetles with yaks and zebras.
- Don't waste time computing character-based statistics (character and n-gram frequencies, character entropy, character correlations as function of distance or position, etc.). Study word statistics instead.
- Never make any claim about statistics of natural languages, no matter how "obviously true", without checking it first. You will often be surprised.
(And as for advice #2, in a recent reply to @JoJo_Jost I claimed without checking that in English there is an attraction between th-words and i-words, because I thought that "obviously" the pairs "that is" and "this is" would dominate the counts. But in fact, as shown above, the repulsive word "the" is so common that the sum of all th/i attractions and repulsions is a net repulsion.)
All the best, --stolfi
chenxiang > 5 hours ago
Quote:There is no classical text that supports the idea of mapping the VMS "things" to ...
chenxiang > 5 hours ago
chenxiang > 4 hours ago
(Yesterday, 06:14 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.(Yesterday, 03:56 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.there are many Chinese characters that could have been the source for this one...We have clearer outlines of the shape and fine details. I think we should ask @chenxiang...
ololololo > 4 hours ago