The Voynich Ninja

Full Version: The 'Chinese' Theory: For and Against
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Stolfi, if you don’t understand why I don’t consider the recipe a serious document, I’ll make a similar comparison.

Let’s take the Carmen Arvale: Enos lases juvate
Neve luae rue, Marma, sins incurrere in pleores
Satur fu fere Mars. Limen sali. Sta. Berber.
Semunis alternei advocapt conctos
Enos, Marmor, juvato
Triumpe triumpe triumpe triumpe triumpe.

Oh my God! It seems I’ve found a comparison for triumpe!
[attachment=17695]
Let’s assume that qokedy is qokeedy written with a mistake.

If you find any differences between your comparison and SBS, be sure to let us know.
(11 hours ago)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.@ Stolfi and chenxiang

Sorry, Stolfi, I actually didn’t want to get involved here anymore because I don’t know enough Chinese. But I have two points regarding your statement:

Stolfi: I now consider the SPS = SBJ claim proven.

1. Your probability calculation isn’t correct as stated, as you yourselves even point out on your page. Because

a. You group several “aiin” forms together as a single unit. This is an unproven assumption, but one you need to establish the “Crip.” At the same time, you exclude other “aiin” variants. The rationale for this was developed after the fact. It was not a prior assumption, and you know how such assumptions—made in hindsight—can be used to distort all sorts of things.

b. If you were to omit this assumption, you would have to include all “aiin” forms, or use only “aiin” (i.e., the raw text); the entire calculation would immediately collapse.

c. You’re comparing two high-frequency words within a limited context; these factors aren’t taken into account in your highly simplified calculation.

d. Even if your calculation were correct, it would still only be a probability, not proof. Even a very low probability leaves both possibilities open. There is NO proof!

2. The strange long-distance effect 

And what you haven’t been able to explain so far- because you don’t speak Chinese, and I understand that - is whether these strange long-range effects exist in Chinese:





If you take the inner glyphs in a word ending in "y" - "sh", "ch" "k" have a significant different effects on how often "q" (=qo) appears as the initial glyph (qo) in the next (!) token.

[*]sh + edy: 51% y->q(o)
[*]ch + edy: 39% y->q(o)
[*]k + edy: 28% y->q(o)


My question for chenxiang

Do such “long-range effects” exist in Chinese?

The question is so relevant for this reason: A Chinese text describing different plants, properties, and active ingredients would have different token sequences for each concept. Instead, the same rule always applies. This means the VMS text is essentially blind to its supposed content—unless it’s a cipher.

And there's more. "sh" generally has a stronger long-distance effect on "qo" than "ch" or "k" do in all tokens. And some of the differences are significant. This does not correspond to any standard Western European language; that much is relatively certain. Of course, I don't know yet how this applies to Chinese.
[*]











"Hi everyone. As a native Chinese speaker, I want to answer the question about 'long-distance effects' in Chinese.

No, Chinese absolutely does not have this kind of long-distance dependency, neither in syntax nor in phonetics. Chinese is an isolating language. Each character stands alone. The strokes, structure, or pronunciation of one character have zero mechanical influence on the next character.

If the Voynich manuscript were a simple phonetic transcription of Chinese (like Pinyin), we still would not see this 51% vs 28% dependency between the end of one word and the start of the next.

The statistical charts shown here (sh+edy -> 51% qo, k+edy -> 28% qo) strongly suggest a mechanical generation process or a complex cipher system, rather than a natural language transcription. Even if Stolfi's 'daiin=主' matching is correct, this long-distance effect is a major problem that cannot be explained by natural Chinese grammar.

I am just a beginner, but I hope this native intuition helps your debate!"
(6 hours ago)chenxiang Wrote: You are not allowed to view links. Register or Login to view.Please don't think I'm just being emotional!

No problem at all, your skepticism is quite understandable.  Shy

Quote:I'm just very mad at myself for missing the word '冬天' (Winter) ... even if I had guessed '冬天', I would have hit the same dead end you did in 2002. ... since you ruled out '冬天' 20 years ago, are you still looking at this You are not allowed to view links. Register or Login to view. symbol,

I am still convinced that those two symbols are upside-down Chinese characters, that went through at least two steps of copying by people who could not read them. The Scribe, in particular, probably never had seen Chinese writing in his life.  I don't see any other plausible explanation for their overall shapes and for  the details like the brush strokes simulated with a quill pen.   

But  don't know which were the original characters.  Maybe they were indeed 冬天.  But I understand why that guess has problems, so I would not bet $10 on that.  Same for the possible connection to the Bai medical tradition (involving homophone confusion or other errors).

Quote:the historical timeline (Ming Dynasty vs. Marco Polo's era)

The Ming xenophobic policies are a constraint that must be taken into account.  

One way to get aroud it is to assume that the Author was in China and got the dictation of the SBJ just before the Ming intensified bans on foreign traders and forced him to return to Europe.  If he was in China in 1375 when he was 20, he would be 75 in 1430, which is about the earliest that the C14 date allows for the transcription of his notes to vellum. 

But it seems that the Ming did not just expel all foreigners from China.  I understand that they only banned foreign traders, mostly in the coastal area; isn't that so? 

And Muslims were not included in the Ming ban, is that correct?  The qo="and" guess suggests that the Author's native language was Arabic or Hebrew...

You are not allowed to view links. Register or Login to view. was a British explorer who disguised himself as a Muslim in order to visit Mecca around 1850.  Maybe the Author was an European who used that same trick to enter China around 1400.

Another possible solution to the Ming problem is that the dictation may have happened outside China, e.g. in Vietnam, Cambodia, Thailand, Burma, Tibet, etc. Or even in Korea or Japan: while their languages are not monosyllabic, when reading the SBJ a doctor in those countries probably would have read each hanzi as one syllable, in some pseudo-Chinese pronunciation.

All the best, --stolfi
(Today, 04:03 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.
(Yesterday, 01:49 PM)chenxiang Wrote: You are not allowed to view links. Register or Login to view.I have seen some of your translations of Classical Chinese sentences and characters, and they are correct. This shows that you also have a good understanding of Chinese.

Thanks for checking, but they are not "my" translations. As I wrote before, I don't speak or read any Chinese or East Asian language, not even enough to say "hello"; except for a tiny bit of Japanese and a couple dozen hanzi.  All the translations that I have posted were provided by internet tools, like Google Translate, ChatGPT, and Gemini (Google's LLM). 

Fortunately, knowledge of Chinese (or whatever is the language of the VMS) is not essential for the work I am doing now.  I need translations of the recipes mainly let me identify parts that the VMS Author apparently left out, like the "flavor" (辛),  "another name" (一名) and "provenance" (生), and some comments that do not make sense outside China, or even inside it; like 东门上者尤良 about chicken heads.  The translations provided by Gemini have been sufficient for that purpose.  But Gemini often makes mistakes or just makes stuff up,so I may need to consult you eventually. 

Quote:what kind of help do you need most? For example, checking Chinese texts, comparing specific passages, finding references, or something else?

My biggest problem is that I don't have an adequate digital file of the Shennong Bencao Jing (SBJ) as would have been available to the Author

As I wrote before, it seems that he transcribed the SBJ as it was quoted in some medical encyclopedia available around 1400 CE, most likely the Zhenghe Bencao (ZHB).  Specifically, the text that was printed in double-size characters, white-on-black, as seen in You are not allowed to view links. Register or Login to view. .

I  started this investigation using two files that claimed to be the SBJ: one from a site called the Chinese Texts Project (CTP), and the other from the Chinese Wikisource archive.  But then I found that those files were modern attempts to reconstruct the book as it would have been in 200 CE or so, and (apart from having quite a few errors) differed from the text quoted in the ZHB in several critical details.  For instance, where the ZHB-quoted text uses 主, those two files used 主治. And 主 = daiin is still the most important "crib" (hanzi-Voynichese pair) that I have at the moment.

So now I have switched to an HTML file that claims to be the whole text of the ZHB (well over 1000 recipes) -- not images, but the hanzi, in Unicode.  In this file, there is markup identifying the large-character parts, but not the white-on-black subset.  Possibly that file was derived from a copy of the ZHB that was printed after ~1700 CE, when printers dropped the white-on-black convention and used black-on-white for everything.

So at present I must take one recipe at a time, extract the large-font text, and then ask Gemini to identify the parts that were white-on-black in the older printed copies.  Gemini has access to literally thousands of articles that discuss the ZHB and the SBJ in great detail, so apparently it knows which are the parts I need.  But it makes mistakes half the time, contradicting itself, sometimes even in the same sentence.  So I must ask almost one word at a time, and and challenge it with the text from the CTP. 

"Which parts of this recipe were white-on-black?" 
"These parts: ..." 
"But the CTP file does not have these two characters ... Is it incorrect" 
"You got me there, yes, indeed those two characters were black-on-white."
"On the other hand, the CTP file has a few more characters after ..."
"Indeed, I was wrong I missed these 5 characters: ..."
"But the last three are harvesting instructions, that are usually black on white, no?"
"Ah, yes, you are so smart! Indeed, that was black-on-white."

And so on.  When I ask ChatGPT to confirm Gemini's answers, they often disagree.  This process is tiresome and takes about half an hour to one hour per recipe.  I did some 50 recipes so far out of the ~365 that are supposed to be in the SBJ.

So one BIG help I need is to locate a similar digital file of the ZHB, extracted by OCR from an older print, that has the white-on-black text marked out.  Do you know if such a thing exists?

Quote:If possible, could you give me one concrete alignment example: one passage from the Shennong Bencao Jing, the corresponding SPS paragraph, and how the symbols/words correspond?

Yesterday I posted to this thread the SBJ entry for 粟米 = "foxtail millet" and the SPS parag that starts at line f107v.20, which I believe is its transcription.   The approximate correspondence is given as text and graphically. 

Quote:I would especially like to know: Is there a stable mapping from VMS symbols to Chinese sounds?

No.  All I know is that each hanzi of the SBJ usually corresponds to one token (word) of the SPS, and about 5 EVA letters on the average.  And the structure of Voynichese words has been studied in detail, and seems to be generally consistent with a phonetic writing like pinyin, jyutping, gwoyeh romatzy, the modern Vietnamese script, or the Burmese, Tibetan, and Thai traditional scripts.  Namely, there are Voynichese symbols that occur only at the beginning, only in the middle, and only at the end of each word, with certain rules about which symbols can follow others.  

Spaces in the Voynichese transcription files are notoriously unreliable, so the correspondence 1 hanzi = 1 word is only very approximate.   The correspondence 1 hanzi = 5 EVA letters seems to be more solid, but only in the average.  It is likely that some hanzi will be found to map to a single EVA letter, while others may map to words of 6-7 letters or more.  

Quote:Were the omission rules fixed before the comparison, or adjusted afterward?

The main omission rules (for "flavor" and "provenance") where deduced when I compared the longest SBJ recipe ("red rooster") to the longest SPS parag (starting at line f105v.32).  I noticed that the spaces between the 主 characters in the SBJ text matched the spaces between the daiin (and variations) in the SPS, considering the 1:5 ratio above; but the initial and final SPS gaps (the EVA strings before the first daiin and after the last one) were way too short.  

Looking at the translation I guessed that the VMS author had omitted the "flavor" field 味甘微温  and the "provenance" field 生平泽 from the SBJ entry.  Deleting the first made the initial gap almost correct (7 EVA letters instead of the expected 3 x 5 = 15).  But the final gap of the SPS parag was still too short, even assuming that "provenance" was skipped.  Checking again the translation I guessed that the VMS author had omitted also the text 鸡白蠹能肥脂 = "chicken white grub can fatten the fat", which, according to Gemini, was incomprehensible even to Chinese doctors in the 1400s.  Deleting that text (which had no 主) made the final gap almost right.  Deleting also the comment 神物 = "[amber is] a divine substance" made the final gap almost perfect -- only 1 EVA letter too long.

Luckily for me, the CTP file that I first used to search for the "rooster" recipe lacked the comment 东门上者尤良 = "those [chicken heads] hung above the East Gate are particularly potent [for killing ghosts]".  When I switched to the ZHB-quoted text, that comment appeared, and then the gap between the first and second daiin became too short.  So I deduced that this comment too had been omitted by the author.

Based on the Rooster and a couple other cases, I decided to assume that the "flavor", "provenance", and "another name" ( 一名) fields had been systematically skipped by the Author on all SBJ entries.  Thus now I always delete those fields before looking for the matching SPS parag.  

Some entries have suspicious non-medical comments, like the "East gate" above, which may or may not have been omitted.  In those cases I must try both variants of the recipe, with and without that comment.  Usually at most one of the variant.

All the best, --stolfi


Hello Professor Stolfi,

I am really sorry to hear about your painful experience with Gemini and ChatGPT! As a Chinese speaker, I completely understand your frustration. AI models often hallucinate when dealing with ancient Chinese texts because they don't truly understand the historical printing formats like "白底黑字" (black characters on a white background) and "黑底白字" (white characters on a black background).

Your request is very clear. Unfortunately, I cannot immediately give you a ready-to-use OCR file with those specific markers, but I can tell you exactly how to find it, and I am willing to help you search on Chinese websites.

Here is my advice for your next step:

Don't rely on AI for the text structure. Instead of asking Gemini "which parts are 白底黑字", you should look for scanned images of the original book. The book you are looking for is likely the "重修政和经史证类备用本草" (the famous 晦明轩 edition of the ZHB).

Where to look: You can ask Chinese scholars or search on Chinese digital libraries like "中国哲学书电子化计划" (CTEXT) or "识典古籍". However, the "白底黑字" format is a specific visual feature of the physical print.

How I can help you: If you can send me screenshots or PDF pages of the specific ZHB recipes you are working on, I can try to help you read the layout, or I can post on Chinese forums asking for a specific digital archive of the "晦明轩" print.

I am still a beginner, but I know how to navigate Chinese internet. Let's work on this together. Please send me a sample page first so I can see exactly what you mean by "black background and white characters"!

Best, Chenxiang
When ancient books were carved and printed, the main text was often black characters on a white background, while important notes or quoted original texts from other ancient books would be printed with white characters on a black background (or vice versa). This is common knowledge in Chinese ancient book studies, but AI has a hard time recognizing it.
(4 hours ago)chenxiang Wrote: You are not allowed to view links. Register or Login to view."Hi everyone. As a native Chinese speaker, I want to answer the question about 'long-distance effects' in Chinese.

No, Chinese absolutely does not have this kind of long-distance dependency, neither in syntax nor in phonetics. Chinese is an isolating language. Each character stands alone. The strokes, structure, or pronunciation of one character have zero mechanical influence on the next character.

@ Stolfi Then you have a serious problem—since these effects occur not only with “sh” but also with other letter combinations. 

Unfortunately, that disproves your theory. 

Under these circumstances, the VMS CANNOT be Chinese - (this long distance effect also occurs in the “Stars” section.)


Sorry... Cry Cry Cry


PS: And certainly not if `qo = "and` were the case. Because `and` has no side effects, since it’s a list, and before you start a list, no one knows that it’s about to start or how long it will last.
(3 hours ago)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.
(4 hours ago)chenxiang Wrote: You are not allowed to view links. Register or Login to view."Hi everyone. As a native Chinese speaker, I want to answer the question about 'long-distance effects' in Chinese.

No, Chinese absolutely does not have this kind of long-distance dependency, neither in syntax nor in phonetics. Chinese is an isolating language. Each character stands alone. The strokes, structure, or pronunciation of one character have zero mechanical influence on the next character.

@ Stolfi Then you have a serious problem—since these effects occur not only with “sh” but also with other letter combinations. 

Unfortunately, that disproves your theory. 

Under these circumstances, the VMS CANNOT be Chinese - (this long distance effect also occurs in the “Stars” section.)


Sorry... Cry Cry Cry


PS: And certainly not if `qo = "and` were the case. Because `and` has no side effects, since it’s a list, and before you start a list, no one knows that it’s about to start or how long it will last.




Specific analysis is as follows: Sad Sad

Natural languages are driven by semantics and syntax, not character statistics: In Chinese (and any natural language), what the first character of the next word is depends mainly on semantics (what one wants to express) and grammar (such as verb-object collocations and subject-predicate relationships). For example, in traditional Chinese medicine botanical texts, the probability of "Honghua" (Safflower) being followed by "Huoxue" (promoting blood circulation) is high because of semantic relevance, not because of some glyph property of the character "hong".

The true manifestation of "remote effect" in Chinese:

Syntactic dependencies (long-distance dependencies): For example, "如果...就..." (if...then...), "因为...所以..." (because...therefore...). Or the "吗" (ma) at the end of a question echoing the interrogative at the beginning. But this is grammatical structure, with clear semantic boundaries.

Prosody and meter: In classical poetry, rhyming (e.g., rhyming at the end of the 1st, 2nd, and 4th lines) and tonal constraints (pingze) do have long-distance constraints, but these are restrictions of specific literary genres, not general rules for ordinary natural language texts (such as botanical monographs).

Tone sandhi: Chinese has tone sandhi for "一" (yi) and "不" (bu), but this is strictly limited to the immediately following syllable (adjacent), and is absolutely never affected by prefixes several words prior (such as the hypothetical affixes "sh", "ch").

Counter-evidence: If a Chinese text describing different plants has character sequence statistics that do not vary with the type of plant (semantic content), but are strongly constrained by prefixes like "sh", "ch", "k", etc., spanning multiple characters to determine the probability of the subsequent "q(o)", then this book is absolutely not a natural language botanical monograph, but can only be an encrypted text or generated pseudo-text. This precisely confirms what you said: "The VMS text is essentially blind to its purported content—unless it is a cipher."

In natural languages, such a strong mechanical dependency is rare (unless it involves some form of grammatical agreement, or highly specific prosodic/phonological rules, but usually not to the precise extent of this puzzle-like statistics. In science, absolutes are rare). But this is very strong counter-evidence. Since Chinese does not have this rule, the manuscript cannot be a direct pinyin transcription of Chinese. However, this absolutely does not mean Stolfi's "Chinese theory" is completely dead. It only means: I feel that if it really is Chinese, then it is absolutely not a "direct transcription," but has gone through a complex codebook or artificial mechanical encoding. I must tell the truth: Chinese indeed does not have this "remote effect." It is an isolating language. If the Voynich manuscript has this 51% vs 28% effect, it absolutely cannot be a natural, direct pinyin transcription of Chinese.

But as a 16-year-old beginner, I don't think this is the doomsday of the "Chinese theory." Perhaps the VMS author was not just transcribing Chinese, but applying a mechanical cipher to a Chinese source text? Or perhaps the manuscript text is a mixed language?

Professor Stolfi, does this "remote effect" mean the clue "daiin = 主" is completely dead too? Or does it only mean there is a layer of encryption on the text? I am still very curious and want to learn!
(2 hours ago)chenxiang Wrote: You are not allowed to view links. Register or Login to view.Professor Stolfi, does this "remote effect" mean the clue "daiin = 主" is completely dead too?
We can’t apply this to anything; for now, it’s just a coincidence. In the VMS text, there are cases where daiin is repeated twice. So, these excerpts should be read as “take for take for?” 
I know that characters can repeat, but these repetitions have certain rules. Does 主 comply with these rules?