The Voynich Ninja

Full Version: The 'Chinese' Theory: For and Against
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
We can actually talk about this:
(Today, 12:38 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.In my estimation, the probability that the matches I have seen (or even just that of the Rooster recipe) are just chance coincidences is negligible. 
because you recently made such an estimation!

I asked about the box width because I didn't think your combinatorial back-of-the-envelope computation bears out, but I wanted to make sure that sense wasn't because I misunderstood the parameters of the thought experiment that generated it. I still don't think it bears out, but I want to front that I suspect we both think it's inadequate, so if you'll permit me to blow through that fairly quickly, we can get to what I think is the actual takeaway we're circling around.

Just for reference, here is the thought experiment.
(08-09-2026, 10:26 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Imagine that you have a row of 30 boxes, and 8 of them have a coin inside,while the others are empty.  You pick 15 of those boxes at random.  You find that 8 of those 15 boxes have a coin inside.

There are comb(15,8) = 6435 subsets of 8 boxes out of the 15.  That is the number of ways you could have got all the 8 coins that way.  That is a big number. Does it mean that this outcome is unremarkable?

Of course not.  That number alone does not mean anything.  To answer that question one must consider also the total number of boxes and how many of them have coins.

The probability P that picking M of N boxes at random will yield J of the K coins they contain is comb(K,J)*comb(N-K,M-J)/comb(N,M).   

For N  = 30, K = 8, M =15, J = 8, this formula gives P = ~0.0011 = ~0.11%; that is, less than one chance in 900.  So getting all 8 coins by picking 15 out of 30 boxes IS a remarkable result.

Simply put, I don't think there is a principled way you can hold J = K = 8 here. Using 30 boxes and absolute positions from the matched entries in the Rooster identification, you end up with somewhere between J = 6 and J = 4 depending on how exactly you do the box assignment, which by the hypergeometric function you offer is 9.0% or 31.8% respectively. You're likely to object to absolute positions---though they are what is implied by a 30 equal boxes model and I think most appropriate on the terms you've set out. (The fact that you might object and I don't think this is decisive model is why I'm blowing through this.) But two of the discrepancies in your methodology where you get to draw the circle around the crib nonrandomly exceed the radius of one of those boxes. That's not exactly J=6, but it lends some credence to the idea that you might be overestimating how exact this apparent match is.

I don't want to dicker over this estimate and the best way to model this too much. I don't sense either of us thinks its a great model, but more to the point, I think we can make a few predictions if you were getting a high rate of false positives implied by the less exact rate of hits without having to commit to a model. I'd expect that you would see many plausible matches and that different inputs would give the same matches. You're manifestly running into these problems. You sort through 7 options you consider plausible You are not allowed to view links. Register or Login to view., and do not even go with the one most suggested by the badness scores. You have double assigned to f115v.11, as noted You are not allowed to view links. Register or Login to view.. I appreciate you being transparent enough that this is showing up in your write-ups, but I think you should take your findings seriously. These findings are not consistent with the claim that you are randomly finding a single close match with a low probability of the match being chance.

I will take another crack at reading your code---I've had access to it for months and I asked for the expressions underlying it because I kept getting lost in the variable chasing. I have some thoughts about the information theory problem implied by it, but they are not done baking yet
(Yesterday, 01:49 PM)chenxiang Wrote: You are not allowed to view links. Register or Login to view.I have seen some of your translations of Classical Chinese sentences and characters, and they are correct. This shows that you also have a good understanding of Chinese.

Thanks for checking, but they are not "my" translations. As I wrote before, I don't speak or read any Chinese or East Asian language, not even enough to say "hello"; except for a tiny bit of Japanese and a couple dozen hanzi.  All the translations that I have posted were provided by internet tools, like Google Translate, ChatGPT, and Gemini (Google's LLM). 

Fortunately, knowledge of Chinese (or whatever is the language of the VMS) is not essential for the work I am doing now.  I need translations of the recipes mainly let me identify parts that the VMS Author apparently left out, like the "flavor" (辛),  "another name" (一名) and "provenance" (生), and some comments that do not make sense outside China, or even inside it; like 东门上者尤良 about chicken heads.  The translations provided by Gemini have been sufficient for that purpose.  But Gemini often makes mistakes or just makes stuff up,so I may need to consult you eventually. 

Quote:what kind of help do you need most? For example, checking Chinese texts, comparing specific passages, finding references, or something else?

My biggest problem is that I don't have an adequate digital file of the Shennong Bencao Jing (SBJ) as would have been available to the Author

As I wrote before, it seems that he transcribed the SBJ as it was quoted in some medical encyclopedia available around 1400 CE, most likely the Zhenghe Bencao (ZHB).  Specifically, the text that was printed in double-size characters, white-on-black, as seen in You are not allowed to view links. Register or Login to view. .

I  started this investigation using two files that claimed to be the SBJ: one from a site called the Chinese Texts Project (CTP), and the other from the Chinese Wikisource archive.  But then I found that those files were modern attempts to reconstruct the book as it would have been in 200 CE or so, and (apart from having quite a few errors) differed from the text quoted in the ZHB in several critical details.  For instance, where the ZHB-quoted text uses 主, those two files used 主治. And 主 = daiin is still the most important "crib" (hanzi-Voynichese pair) that I have at the moment.

So now I have switched to an HTML file that claims to be the whole text of the ZHB (well over 1000 recipes) -- not images, but the hanzi, in Unicode.  In this file, there is markup identifying the large-character parts, but not the white-on-black subset.  Possibly that file was derived from a copy of the ZHB that was printed after ~1700 CE, when printers dropped the white-on-black convention and used black-on-white for everything.

So at present I must take one recipe at a time, extract the large-font text, and then ask Gemini to identify the parts that were white-on-black in the older printed copies.  Gemini has access to literally thousands of articles that discuss the ZHB and the SBJ in great detail, so apparently it knows which are the parts I need.  But it makes mistakes half the time, contradicting itself, sometimes even in the same sentence.  So I must ask almost one word at a time, and and challenge it with the text from the CTP. 

"Which parts of this recipe were white-on-black?" 
"These parts: ..." 
"But the CTP file does not have these two characters ... Is it incorrect" 
"You got me there, yes, indeed those two characters were black-on-white."
"On the other hand, the CTP file has a few more characters after ..."
"Indeed, I was wrong I missed these 5 characters: ..."
"But the last three are harvesting instructions, that are usually black on white, no?"
"Ah, yes, you are so smart! Indeed, that was black-on-white."

And so on.  When I ask ChatGPT to confirm Gemini's answers, they often disagree.  This process is tiresome and takes about half an hour to one hour per recipe.  I did some 50 recipes so far out of the ~365 that are supposed to be in the SBJ.

So one BIG help I need is to locate a similar digital file of the ZHB, extracted by OCR from an older print, that has the white-on-black text marked out.  Do you know if such a thing exists?

Quote:If possible, could you give me one concrete alignment example: one passage from the Shennong Bencao Jing, the corresponding SPS paragraph, and how the symbols/words correspond?

Yesterday I posted to this thread the SBJ entry for 粟米 = "foxtail millet" and the SPS parag that starts at line f107v.20, which I believe is its transcription.   The approximate correspondence is given as text and graphically. 

Quote:I would especially like to know: Is there a stable mapping from VMS symbols to Chinese sounds?

No.  All I know is that each hanzi of the SBJ usually corresponds to one token (word) of the SPS, and about 5 EVA letters on the average.  And the structure of Voynichese words has been studied in detail, and seems to be generally consistent with a phonetic writing like pinyin, jyutping, gwoyeh romatzy, the modern Vietnamese script, or the Burmese, Tibetan, and Thai traditional scripts.  Namely, there are Voynichese symbols that occur only at the beginning, only in the middle, and only at the end of each word, with certain rules about which symbols can follow others.  

Spaces in the Voynichese transcription files are notoriously unreliable, so the correspondence 1 hanzi = 1 word is only very approximate.   The correspondence 1 hanzi = 5 EVA letters seems to be more solid, but only in the average.  It is likely that some hanzi will be found to map to a single EVA letter, while others may map to words of 6-7 letters or more.  

Quote:Were the omission rules fixed before the comparison, or adjusted afterward?

The main omission rules (for "flavor" and "provenance") where deduced when I compared the longest SBJ recipe ("red rooster") to the longest SPS parag (starting at line f105v.32).  I noticed that the spaces between the 主 characters in the SBJ text matched the spaces between the daiin (and variations) in the SPS, considering the 1:5 ratio above; but the initial and final SPS gaps (the EVA strings before the first daiin and after the last one) were way too short.  

Looking at the translation I guessed that the VMS author had omitted the "flavor" field 味甘微温  and the "provenance" field 生平泽 from the SBJ entry.  Deleting the first made the initial gap almost correct (7 EVA letters instead of the expected 3 x 5 = 15).  But the final gap of the SPS parag was still too short, even assuming that "provenance" was skipped.  Checking again the translation I guessed that the VMS author had omitted also the text 鸡白蠹能肥脂 = "chicken white grub can fatten the fat", which, according to Gemini, was incomprehensible even to Chinese doctors in the 1400s.  Deleting that text (which had no 主) made the final gap almost right.  Deleting also the comment 神物 = "[amber is] a divine substance" made the final gap almost perfect -- only 1 EVA letter too long.

Luckily for me, the CTP file that I first used to search for the "rooster" recipe lacked the comment 东门上者尤良 = "those [chicken heads] hung above the East Gate are particularly potent [for killing ghosts]".  When I switched to the ZHB-quoted text, that comment appeared, and then the gap between the first and second daiin became too short.  So I deduced that this comment too had been omitted by the author.

Based on the Rooster and a couple other cases, I decided to assume that the "flavor", "provenance", and "another name" ( 一名) fields had been systematically skipped by the Author on all SBJ entries.  Thus now I always delete those fields before looking for the matching SPS parag.  

Some entries have suspicious non-medical comments, like the "East gate" above, which may or may not have been omitted.  In those cases I must try both variants of the recipe, with and without that comment.  Usually at most one of the variant.

All the best, --stolfi
(Yesterday, 12:37 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.
(Yesterday, 06:45 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.
(16-09-2026, 09:04 AM)rikforto Wrote: You are not allowed to view links. Register or Login to view.The other thing I'm going to want to know is what the badness cutoff is---when is that badness too high to count as a match?
Since the scale of the badness score is somewhat arbitrary, there is no cutoff.  I currently use the score only to rank the candidate parags.

Here is part of the report for the SBJ recipe for 粟米 = "foxtail millet".   The full entry is

  粟米:[味]苦,微寒。无毒。[主]养肾气,去胃脾中热,益气。
  陈者:[味]苦。[主]胃热,消渴,利小便。

The punctuation is a modern addition; the You are not allowed to view links. Register or Login to view. (ZHB) as available in 1400 had no punctuation, so the text as the Author would get it is

  粟米味苦微寒无毒主养肾气去胃脾中热益气陈者味苦主胃热消渴利小便

This entry, atypically, has no "Provenance" field.  After deleting the systematic omissions (味苦微寒 and 味苦) only 25 hanzi remain:

  粟米无毒主养肾气去胃脾中热益气陈者主胃热消渴利小便

Given the average of ~5 EVA letters per hanzi, the corresponding SPS entry is expected to have ~5 x 25 = ~125 letters.   The program was asked to check the occurrence of four "cribs" which I am currently fairly certain about: 主, 气, 气, 主, in that order.  The "canoncal" Voynichese translations should be 主 → daiin and 气 → chedy; but the program will accept some variations like kaiin and cheda,with penalties in the score.  

My program considers only the 243 parags of the SPS that have a single star.  The others are assumed to be two or more parags mashed together by the Scribe, and therefore not suitable for this analysis.

Here is the relevant part of the output:

      243 parags read
      243 were evaluated.
      46 were rejected for having less than 111 letters.
      126 were rejected for having more than 150 letters.
      71 were examined for the requested cribs
      66 had all the cribs -- sizes 114..150
      64 parags had badness 99.0 or less.

To save time, my program rejected right away any SPS parag whose size was too far form that expected size, which left 71 parags.  Of these, 5 were rejected because they did not have the requested cribs nor any of the currently allowed variations thereof.
     
          <b2.6.122> (t1a1)  25
               4   粟米无毒
            1  2 主 养肾
            1  6 气 去胃脾中热益
            1  2 气 陈者
            1  7 主 胃热消渴利小便
   
This is the trimmed SBJ entry above, split at the requested keywords.  The second  column is the counts of hanzis before, between, and after the keywords 

Here is the best (minimum-score) match that the program found:   
   
      <f107v.20>  (q0r0)     0.443 = 0.300+0.143    135(+0)    5.400 e/h


                   18(-3)    0.052 palaroraiinkalkalo
      kaiin 0.150  12(+1)    0.034 chpalcheodyp
      chedy 0.000  34(+1)    0.012 pcholkalopalkaiintaiinolkaiincholk
      chedy 0.000  10(+0)    0.000 kolainaiin
      kaiin 0.150  41(+3)    0.046 okalotarotalaldaralololaiinolkalcholchdar

   
The best-scored parag was f107v.20, with badness 0.443.  The program did not find the two expected daiin at the right places, but it found two kaiin, which caused 2 x 0.150 = 0.300 to be added to the score; but found the two expected chedy.

The note (q0r0) indicates that occurrences of qo were deleted from the parag, and the strings ir, is, and m were expanded to iin, as discussed in previous posts.  After these adjustments, this parag had 135 EVA letters For this parag the program assumed an average of 5.4 EVA letters per hanzi.  With that ratio, the expected sizes of the gaps were 21.6, 10.8, 32.4, 10.8, and 37.8, rounded to 21..22, 10..11, 32..33, 10..11, and 37..38.  The actual sizes of the gaps were 18, 12, 34, 10, and 41, giving errors of -3,+1,+1,0,+3 EVA letters, which correspond to less than 0.56 hanzi.  These errors contributed 0.143 points to the score.

The second-best match was 
   
      <f116r.4>  (q0r0)      1.512 = 0.350+1.162    150(+13)  5.449 e/h


                   17(-4)    0.097 padarsheyosheekyl
      laiin 0.160  12(+1)    0.034 chckhyokaiin
      chedy 0.000  34(+1)    0.012 oteedytararalarydainsheedkchdyotal
      chedy 0.000  14(+3)    0.258 lkainoteedyoto
      raiin 0.190  53(+14)   0.761 otylolrololysainollchedychedyoteychedy?lainotedyoteey


Note that the score is 1.512, of which 0.350 comes from using alternate "spellings" laiin and raiin for 主, and 1.162 comes from the errors in the gap lengths; especially the "tail" gap, that is 14 letters longer than predicted. 

(This parag is also 150 letters long, which would be 6.000 EVA letters per hanzi.  The program considered that value too large and used the ratio 5.449 instead when evaluating the gap errors.  But even if it had used 6.000 the last gap would still be too long and would have given this parag a bad score.) 

Thus I think that f107v.20 is probably the translation of the Foxtail millet entry, and there are no other plausible candidates.

Here is the graphical rendition of the match for that parag:

All the best, --stolfi


I don’t quite understand this yet, and I’m still thinking about it. Could any senior here briefly explain to me the principle behind using this algorithm for analysis?
I have a certain understanding and analysis of the conjectures and deductions Professor Stolfi has provided, and I am drafting a formatted reply to your earlier views. But before that, I still have a very puzzling question.

You said earlier that the hypothesis you roughly discussed is that you think this might not be a cipher, but rather that a European who had been to China listened to Chinese doctors reciting medicinal materials and then used an alphabet he invented to record Chinese pronunciation as dictation notes? But what I really don’t understand is: why did he have to listen to Chinese doctors reciting, and why did he have to use an alphabet he invented to record Chinese pronunciation? Why didn’t he do something simpler: take a copy of an ancient Chinese book back with him, or print one, or even copy a medicinal book word for word without missing a single character? He could have taken it back to his own country and found a relevant diplomat or a scholar who knew both languages to translate it. Wouldn’t the result obtained that way be more accurate? Wouldn’t it be less likely to produce misleading errors?

Take me as an example. My English is not very good, and sometimes I rely on a translator. If I come across an English word I don’t know how to pronounce, I used to use Chinese that I could understand to write one or two words. Reading these words together sounds close to the pronunciation of that English word I couldn’t pronounce. But this can only serve as my method for practicing spoken English and reading words. The problem is, if they tried their best to write down the same pronunciation, why did they still invent their own alphabet? Isn’t that more error-prone or more complicated? Huh
As far as I’m concerned, there isn’t a single indication in the book that in any way points to a Chinese origin.
Basing a translation on a recipe is utter nonsense.
Why? Boil xxx in water and drink it three times a day. Hmmm… it’s called tea.
This sort of recipe applies to every language. That’s how it is with a great many recipes, no matter what language they’re in.

[attachment=17682]
Addendum:
Word repetitions occur time and again.
@ Stolfi and chenxiang

Sorry, Stolfi, I actually didn’t want to get involved here anymore because I don’t know enough Chinese. But I have two points regarding your statement:

Stolfi: I now consider the SPS = SBJ claim proven.

1. Your probability calculation isn’t correct as stated, as you yourselves even point out on your page. Because

a. You group several “aiin” forms together as a single unit. This is an unproven assumption, but one you need to establish the “Crip.” At the same time, you exclude other “aiin” variants. The rationale for this was developed after the fact. It was not a prior assumption, and you know how such assumptions—made in hindsight—can be used to distort all sorts of things.

b. If you were to omit this assumption, you would have to include all “aiin” forms, or use only “aiin” (i.e., the raw text); the entire calculation would immediately collapse.

c. You’re comparing two high-frequency words within a limited context; these factors aren’t taken into account in your highly simplified calculation.

d. Even if your calculation were correct, it would still only be a probability, not proof. Even a very low probability leaves both possibilities open. There is NO proof!

2. The strange long-distance effect 

And what you haven’t been able to explain so far- because you don’t speak Chinese, and I understand that - is whether these strange long-range effects exist in Chinese:

[attachment=17683]
[attachment=17684]
[attachment=17685]

If you take the inner glyphs in a word ending in "y" - "sh", "ch" "k" have a significant different effects on how often "q" (=qo) appears as the initial glyph (qo) in the next (!) token.

[*]sh + edy: 51% y->q(o)
[*]ch + edy: 39% y->q(o)
[*]k + edy: 28% y->q(o)


My question for chenxiang

Do such “long-range effects” exist in Chinese?

The question is so relevant for this reason: A Chinese text describing different plants, properties, and active ingredients would have different token sequences for each concept. Instead, the same rule always applies. This means the VMS text is essentially blind to its supposed content—unless it’s a cipher.

And there's more. "sh" generally has a stronger long-distance effect on "qo" than "ch" or "k" do in all tokens. And some of the differences are significant. This does not correspond to any standard Western European language; that much is relatively certain. Of course, I don't know yet how this applies to Chinese.
[*]
[attachment=17687]
(Today, 04:31 AM)chenxiang Wrote: You are not allowed to view links. Register or Login to view.Why didn’t he do something simpler: take a copy of an ancient Chinese book back with him, or print one, or even copy a medicinal book word for word without missing a single character? He could have taken it back to his own country and found a relevant diplomat or a scholar who knew both languages to translate it. Wouldn’t the result obtained that way be more accurate? Wouldn’t it be less likely to produce misleading errors?

Because he (like Marco Polo and other Medieval Western travelers to China) had learned the spoken language to some extent, but could not read Chinese characters; and no one in Europe could either.  So a printed copy of the ZHB or any other medical book would have been totally useless to him (or anyone else) when he got back to Europe.

Quote:If I come across an English word I don’t know how to pronounce, I used to use Chinese that I could understand to write one or two words. Reading these words together sounds close to the pronunciation of that English word I couldn’t pronounce. But this can only serve as my method for practicing spoken English and reading words.

And that is what I think the Author did, only in reverse. 

Quote:The problem is, if they tried their best to write down the same pronunciation, why did they still invent their own alphabet? Isn’t that more error-prone or more complicated?

For the same reason that people invented many You are not allowed to view links. Register or Login to view. systems: because writing with the Latin alphabet is too slow for taking down dictation (especially for a language that has "foreign" sounds).

On the average, a Voynichese symbol requires fewer pen strokes than a Latin letter.  Most symbols, like a o y d l r s Ch k t can be written with only two strokes of the pen each.  Only Sh requires three strokes.  The "platform gallows" CTh and CKh require four strokes each, but they could be two phonemes each, Ch+t and Ch+k.

All the best, --stolfi
(Today, 06:02 AM)Aga Tentakulus Wrote: You are not allowed to view links. Register or Login to view.As far as I’m concerned, there isn’t a single indication in the book that in any way points to a Chinese origin.

That is true.  

Except for the structure of the words, the absence of identifiable "European" language features (like articles, copulas, inflections, etc.), the many dozens of plants "that no doctor in Bohemia could recognize", the Zodiac divided into strictly 24 x 15 or 12 x 30 "things" (rather than 28/30/31), ...

And a section whose paragraph count and lengths (in tokens) match the count and lengths (in hanzi) of the recipes from the most important Chinese medical book ever...

And two large symbols on page 1 that don't look anything like Western letters or symbols...

Quote:Basing a translation on a recipe is utter nonsense.

I see that you have spent zero effort on trying to understand what I have been doing or writing.  Oh well.

All the best, --stolfi
Reply to ololololo 

"Hi 
I saw your experiment about creating a phonetic script for the Ingush language. I think this is a very interesting and relevant approach to test the 'Chinese Theory'.

As a native Chinese speaker who studies Classical Chinese, I want to point out something about the Shennong Bencao Jing (SBJ). The text is extremely terse. It doesn't use 'and' (和/及) to connect lists.

If we assume the Voynich author was listening to a spoken language he only partly understood, the 'error rate' for a high-frequency character like '主' (daiin/kaiin/laiin) should be very high. But if he wrote it down in a hurry without understanding, how did he know to systematically omit the '味' (flavor) and '生' (provenance) fields?

I think your experiment could be the key to solve this: if we can find a similar phenomenon in a modern listener (e.g., a person writing down a language they don't speak), it will help us understand whether the VMS author was a 'tape recorder' or an 'interpreter'. Keep up the good work!"
Dear Professor Stolfi,

Thank you for the detailed explanation about the Red Rooster and the "qo" prefix. I understand your point that in Chinese, lists are often written without conjunctions, and the author might have added "qo" to make the transcription easier to parse.

However, this brings us back to the paradox that you yourself acknowledged in the previous post. If the author was capable of parsing the Chinese text to the point where he felt the need to add "and" (qo) to clarify the structure, he must have understood the syntax quite well.

If he understood the syntax, why would he spell "主" in four different ways (daiin, kaiin, laiin, dain) in the SAME short recipe?

In my understanding of Classical Chinese, "主" is a very basic, frequently used character. It is not a rare character that would be easily misheard. If it was just a "phonetic transcription" of an unfamiliar sound, why did he always omit "味" and "生", which are also very common characters in these recipes?

I am not denying the statistical alignment you found. But I am genuinely worried that the assumptions we are making (systematic omission, flexible spelling variations, added conjunctions) are becoming too many. When a theory needs this many ad-hoc rules to work, it becomes hard to prove it to the academic community.

I am still reading the 86 pages and will keep looking for Chinese sources that might help.

Best regards,
Chenxiang