Sorry, we absolutely have to get bogged down in this for a bit because if this in any way represents your understanding of how to read Chinese it is completely and wildly wrong:
(03-10-2026, 09:13 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.If the local language had been Spanish, the Dictator could have said literally "rojo macho galina", or may have tried to finesse the local grammar and said "roja macha galina" or "galina macha roja" or "rojo galo" or "galo rojo"; and the Author still could have written down what he said.
A reading is a conventionalized way of assigning sound and meaning bundles to text. I am going to engage with this speculatively soon enough, but
there is no convention for reading Chinese characters in Spanish. Claims about how someone would read them "in Spanish" require inventing a literary culture that does not exist, has never existed, and if we speculate existed we have no direct data for. If the local language had been Spanish, there would have been no way to read the text in Spanish because you cannot read Chinese in Spanish anymore than you can read an obscure German dialect in phonetic Irish.
This does explain something that has vexed me for some months on your You are not allowed to view links.
Register or
Login to view.---why you have "readings" for non-Sinosphere languages. Thai, Tibetan, and Burmese do not have character readings because they have never conventionalized a way of reading Chinese. There is simply no way to check if your LLM has returned correct values for those "readings" because there is no data to compare it to and so no way to say what is right or wrong. There is just a simple absence of a way to read the characters. This differs from the Mandarin, Cantonese, first Vietnamese, and Bai examples on the site. Best as I can tell---and they are listed in the order of my confidence here---those are largely correctly done readings. The non-examples seem to be trying to follow this mode, but have pulled a different kind of data. Exactly what is hard for me to say because I'm a Sinosphere guy---but that's the point. This is outside my realm precisely because those are not readings of the characters and I have no resources for engaging with them that way. You might as well ask me how to read Bengali in Greek for all the sense the question makes.
A thing I want to stress here is that the valid examples of readings are not strictly "in Mandarin" and so on. They are lists of Mandarin (and so on) pronunciations of the characters that have been conventionalized for reading Classical Chinese. They do not carry their Mandarin (and so on) meaning because the text is in Classical Chinese but for when the course of the linguistic river of these words has remained unchanged for the last 2500 years. Strategies for disambiguating syllables that work in spoken Mandarin (and so on) will fail here because the Mandarin (and so on) context is absent. Placing "xióng" and "jī" next to each other does not result in the Mandarin word for rooster or suggest the proper meanings for "xióng" and "jī" based on Mandarin (and so on) as outlined above.
You have scoffed at the idea that the way those readings are selected is by etymology rather than by translation, but indeed, that is the case. (I probably muddied the water a bit by saying "cognation" previously, which is unambiguously accurate for Chinese, and not unambiguously wrong for the intense relexification that happened in JKV languages and arguably amounts to creolization. At any rate, if you prefer "etymologically", that's what I mean.) This is the way the Spanish example fails the hardest. The etymological term for 丹 does not exist in Spanish because Spanish did not ingest the Chinese lexicon and develop a literary culture around the characters. Had it, Spanish would have borrowed a word for cinnabar (the better gloss for 丹) in the ballpark of /*dan/. It would probably use it in compounds of Chinese origin, but continue to use /cinabrio/ as the unbound form. This, at least, is the default pattern in JKV languages. Likewise, the expectation would be that both Sino-Spanish /*dan/ and native Spanish /cinabrio/ would be written 丹, but they would be separate words available in separate contexts that could be unambiguously recovered from writing. It is rare, but not vanishingly so, for a proficient reader of a passage of well-formed Japanese to have any doubt if one is looking at the Sinitic or Japonic word, and dictionaries and textbooks overtly track which is which by etymology for learning and reference. Regardless, the passage in question is not in Spanish, it is in Classical Chinese. If the Classical Chinese text were read aloud as your examples, the first word would not be expected to be /rojo/ or /cinabrio/, as they are from the wrong etymology stream, but the etymologically related term, fictitiously posited to be /*dan/. There are practices of semi-translation, especially in Japan and Korea, but they quickly run into the problems of full translation outlined below as they too would result in marked changes to the corresponding text in the VMS and those remarks apply here.
This again all applies directly to the Thai, Burmese, and Tibetan "readings". Your LLM has constructed them very differently because, like Spanish, there is no lexicon to draw from when constructing the reading. In Sinosphere languages you wouldn't usually draw from the native terms, but rather the Sinitic ones, except of course in Sinitic languages, where native terms are Sinitic. This, by the way, is how you decide if a language is the Sinosphere or not. If the characters have local readings, it is a Sinosphere language.
To belabor the point and return to actual, attested linguistic practice there is a bit of a shell game happening here, though I don't believe you are doing it on purpose to deceive. You are recovering the sense from the Classical Chinese text with the characters disambiguating meaning. You then assign that word (not character, word) an English gloss. You then assigning the character a Mandarin phonology. You are then taking the English gloss from the Classical Chinese as previously recovered with the disambiguating characters, and applying it to the Mandarin phonology to argue The Author would recognize it. Less transparent to me is how you are getting the green text in the aforementioned passage; You are not allowed to view links.
Register or
Login to view. is obsolete in most standalone uses in Standard Chinese. But taking all those LLM identifications at face value, just because the character still knocking around in some other use, that doesn't mean it would be readily disambiguated from the other You are not allowed to view links.
Register or
Login to view.; or worse, just because "dan" is common in pinyin transcriptions that doesn't mean they all reflect 丹.
This places your arguments on a fork. On the one hand, you can argue that the characters were read character-by-character, in a conventionalized way. This has the benefit of creating a verifiable text that can be analyzed and commented on, and a lot of my remarks are predicated on the idea that you have found the plain text. This has the downsides that Shira and I have raised---Classical Chinese does not correspond on a character-by-character basis with the spoken languages of the 1400s, to the point it was not comprehensible when read aloud even to people who could read the text. This would result in a VMS with, to paraphrase Sampson lightly, a text that serves no communicative purpose.
On the other hand you can abandon this view and take on translation. This has the upside that it restores one of the features of the Dictation Scenario, that The Author would be able to recognize words. But it wreaks havoc on both the matching process you have proposed and the text you have created for that matching. I've had less to say about this for two reasons. First, you've tended towards the first view; you've said it looks character-by-character and your concrete examples are constructed that way. Second, it is much harder to falsify. "The VMS could match any arbitrarily altered plain text" is trivially true, but without those alterations clearly laid out, we don't learn much from it.
So unwinding back from the Spanish example, we get to this:
(03-10-2026, 09:13 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.In the proposed scenario, the Dictator was asked to read the SBJ in the local language, possibly the only East Asian language which the Author could understand. If that language had been modern Mandarin, the Dictator could just have said "dān xióng jī". The Author would then have written down the sounds "dān xióng jī" in his phonetic notation.
If we suppose the local language was Mandarin, that's not spoken Mandarin Chinese! Far from not understanding the proposed scenario, I followed down the claim you made and checked to see if, in fact, the process that your LLM gave you was indeed "in Mandarin Chinese". It was not! You will find the same true for most major varieties of Chinese with the exception of Wu Chinese and a few backwater dialects of others, mostly under Wu Shanghainese influence. (Positing Shanghainese will not solve this problem in the abstract; it will turn out to be innovative on something the prominent spoken Mandarin languages turn out to conserve because that's how language divergence works.)
A "Dictator" doesn't solve these problems because the problems I'm raising aren't based on recovering phonology from the characters! They are about recovering meaning from the phonology. Even in the patently incorrect view that the characters "directly represent concepts", the phonology does not make those concepts available to The Author so arguments that make recourse to them are invalid! You cannot smuggle the information from the characters through separate English glosses and reunite them and call it "the local language". This becomes completely untenable when you acknowledge that Chinese languages (plural) are not streams of disconnected ideographic symbols, but fully realized, contextually understood languages. The phrase "xióng jī" is absent from colloquial Beijing Mandarin because while the morphemes persist, convention does not allow them to combine to form "rooster". Like a lot of linguistic facts this cannot be known a priori; it simply is the way the language conventionally works; Shanghainese demonstrates the road not taken. The problem here is not endogenous to the dictation scenario; the problem is that predictions made by it (namely: the author could understand the words) are not supported by exogenous linguistic evidence.
To lean a bit on the Latin metaphor, I could posit that the VMS was written by a Chinese man in Europe who learned the local language but could not be bothered to learn the Latin alphabet and invented his own. (I recognize this is a much less tempting thing to do when an alphabet is right there, and the analogy fails here. The other points stand.) That language could be any of the many Latin languages of Europe---Spanish, Italian, English, Basque, You are not allowed to view links.
Register or
Login to view., etc. It would not matter because they You are not allowed to view links.
Register or
Login to view.. By listening to a reading of the Tractatus de Herbis in, say, the Spanish pronunciation of Latin You are not allowed to view links.
Register or
Login to view.mYou are not allowed to view links.
Register or
Login to view., he would be able to write down the phonological values and understand it later. I could then insist failure to understand this is a failure to understand the Dictation Scenario, not a problem with how my thought experiment interfaces with reality.
I would be devoured alive in the comments here if I went down this road!
Not for nothing, when I laid out a You are not allowed to view links.
Register or
Login to view. version of this sort of dictation practice transparently in an alphabet with familiar languages, You are not allowed to view links.
Register or
Login to view. You are not allowed to view links.
Register or
Login to view., You are not allowed to view links.
Register or
Login to view., gave quite harsh reviews. I tend to agree with the criticism for the languages in question, but it reflects normal practices in the Sinosphere! So ironically this reaction is an overcorrection, but a revealing one.
Workhorse features of the kind of Sinosphere dictation you have illustrated to back your claims look wrong to you when you have have the linguistic context to see what is actually happening. When they are Romance languages, you recognize the fork I laid out above and resolve it by insisting on freer translation.
This has been entirely too many words to point out that reading the wrong language out loud and calling it a different language does not transmute it into that other language! It is technically feasible that the VMS is dictation of a dead language in a mode that wasn't comprehensible to anyone, but you cannot then assume it was a living language comprehensible to some traveler and fault me, as your critic, for noticing they are two different things. You can insist all you want that Classical Chinese terms are Mandarin, but actual Chinese lexicographers get the decisive say here and they do not agree with your thought experiment! This has bearing if you want to make claims about what The Author did and did not understand from your proposed text alignment