The Voynich Ninja

Full Version: The 'Chinese' Theory: For and Against
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
(09-09-2026, 08:14 AM)Aga Tentakulus Wrote: You are not allowed to view links. Register or Login to view.That just made me laugh. Only a Chinese person could write such nonsense. Amber and alchemy. Amber literally floats towards you and you collect it with a net. A European would have known that. Wink 

That is the case only in a few places in the world.  Maybe just in the Baltic Sea.  Amber from the Baltic to Southern Europe is one of the earliest long-distance trade routes known.   Amber can be mined in a few other places, like the Dominican Republic.  That made amber a precious stone everywhere else.  Like in China, where apparently it was considered a "divine substance".

No one knows exactly what alchemical recipe that note in the "Red rooster" entry was referring to.  Presumably it was some process that turned egg yolks into a hard yellow-orange  transparent substance that looked like real amber.  It does not seem impossible...

All the best, --stolfi
(03-07-2026, 06:12 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.The third column is the part of the badness score that is due to the errors in the gap lengths.  Each term is proportional to the square of the error; the first and last gaps have half the weight of the other gaps.  
Can you provide the precise expression you used to arrive at these figures? I can get something like your numbers from this description, but I cannot quite reproduce the results. 

The other thing I'm going to want to know is what the badness cutoff is---when is that badness too high to count as a match?
(10-09-2026, 01:50 PM)eggyk Wrote: You are not allowed to view links. Register or Login to view.Does the idea of the omission of certain sections make sense with the rest of your theory? I'm under the impression that the VMS text is supposed to represent the spoken SBJ written phonetically.

That the "Nature", "Another name", and "Provenance" fields were omitted is pretty clear from the matching results.  

How the SBJ was turned into the SPS is still a matter of conjecture.  My best guess is that "Dictation Scenario" I have described before.  Namely, the Author lived in "China" (short for "some East Asian country") for a few years, enough to learn the local spoken language to a level good enough for practical purposes.

There he also learned about all those books with stuff that was quite different from what was known in Europe.  Like the Zenghe Bencao (ZHB) materia medica, compiled in ~1100 CE -- 26 volumes, over 1000 pages, listing over 1700 remedies, with extensive commentaries from various other medical books going back to 300 BCE. 

Thus my guess is that he decided to copy and bring back to Europe some of that knowledge. Since he could understand the spoken language to some extent, but had no hope of learning the Chinese characters, his only option was to get some literate local, preferably a medical doctor, to read the book aloud; while he recorded what he said in a phonetic aphabet or shorthand that he devised for the purpose.

Copying the whole ZHB was of course out of the question.  I suppose that he discussed his plan with local doctors, and on their advice settled for copying only the Shennong Bencao part of the ZHB.  That was still regarded as the foundation of their medical science. It had only ~365 terse entries, each with a list of diseases but no other commentaries, and thus would be less than 13 pages long.

In that negotiation, of just after he started the project, he may have realized that it would a waste of time and paper to copy certain parts of the SBJ itelf.  Like those three fields above.  Like the comment in the "Five clays" entry,  五石脂各随五色补五脏 = "The five clays, each according to the five colors, tonify the five solid organs."

So he probably told the Dictator to skip those three fields, and he himself omitted any comment that he did understand and felt that it was pointless to copy.

Or maybe he did copy the whole SBJ, with all those fields and comments; but removeed them when, back in Europe, he prepared the final draft for the Scribe.

Therefore, those omissions imply three possibilities (sub possibilities can be "and/or"):

Quote:[Maybe]the source SBJ text was already lacking those sections.

If the copying happened in the area of strong Chinese influence, the source for the dictation was almost certainly a printed edition of the ZHB, which included those fields as part of the original SBJ material.

However, the Dictation Scenario may have happened in some other country, like Vietnam, where the doctors considered the SBJ the foundation of their medicine.  Then they may have been using some local medical encyclopedia, instead of the ZHB; and those fields may have been omitted in that version.  But I don't know whether such alternative encyclopedias even existed. 

Quote:[Maybe]the VMS author strongly understood the spoken language when respoken aloud later, enough to be certain which sections were which.

If the Author lived for several years in the country, he must have been able to understand at least half of the words that the Dictator pronounced -- even if he could not identify most of the plants, medical conditions, or benefits.  He must have understood all three words 味甘平 = "flavor sweet balanced", and must have quickly learned what they meant and that recording them would be pointless.  Ditto for 生山谷 = "grows mountain valley".  

Quote:I don't understand how you know which sections were apparently omitted, or how you know why. I'm getting the impression that these omissions are having to be included because the VMS was not matching to the original SBJ text as hoped.

Yes. 

Again, the first strong evidence that the SPS could be a transcription of the SBJ  was the good match between the distributions of SBJ recipe lengths and SPS parag lengths.  The Red Rooster recipe and the f105v.32 parag were distant outliers in both distributions, so, if the SPS was indeed the SBJ, those two items should correspond to each other.  Eventually I figured out that 主 corresponded to daiin or close variants thereof: their positions matched not only in that recipe but in many others.  

But, while the gaps between consecutive 主 and consecutive daiin-likes matched, the gaps before the first daiin and after the last one were always way too short.  For example, in the full Rooster recipe, there are 7 hanzi (丹雄鸡味甘微温) before the first 主.  At the average rate of ~5 EVA per hanzi, one would expect 35 EVA letters before the first daiin.  But there are only 8 EVA letters there (poar.keeo).  The difference of 27 letters is not negligible.  

But with the assumpton that the ... field is always omitted, there are ont 3 hanzi (丹雄鸡 = "red male chicken") before the first 主.  The expected number of EVAs before the first daiin then drops to 15.  The actual 8 is still too short, but only by 7 EVA letters.

Maybe poar.keeo is actually 3 words shorter than average, or keeo" is a single word like "rooster" that means "male chiken". Or maybe the Author understood enough of the entry to realize that the title should be just "male chicken", since the "red" iwas just Chinese color superstition.

Quote:If so, this may be an issue when trying to validate this theory, because these paragraphs are no longer being matched with the SBJ as written, but instead being matched with a version that has been altered until a match occurs.

Indeed this is an obstacle on the way to convincing others.  

But, again, there are very few ad-hoc (recipe-specific) omissions.  Those three fields are assumed to be omitted from all entries; and obviously it would make no sense for the Author to transcribe them -- if he had even a vague idea of what they meant.  The few ad-hoc omissions often are at the end of the recipe, so they do not affect the gaps between the keys.

Compare this puzzle to the decipherment of ancient languages.  Even when there is a "Rosetta Stone", the two texts are often only approximate translations.  Does that mean that any decipherment based on them is bogus?

All the best, --stolfi
(10-09-2026, 01:38 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Isn't there a standard translation [of 气], using the original word? You are not allowed to view links. Register or Login to view.
There is no English word for the concept. It is like asking to translate "haggis" or "pizza".  Using the Chinese word would be the best option, but only for those readers who know the concept.  It would leave all other readers baffled, and sort of defeat the goal of translating the recipes.

All the best, --stolfi
(11-09-2026, 03:54 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.If the Author lived for several years in the country, he must have been able to understand at least half of the words that the Dictator pronounced -- even if he could not identify most of the plants, medical conditions, or benefits.  He must have understood all three words 味甘平 = "flavor sweet balanced", and must have quickly learned what they meant and that recording them would be pointless.  Ditto for 生山谷 = "grows mountain valley".  

Now here is an example of why I find this problematic. Why is "flavor sweet balanced" or "grows mountain valley" a pointless thing to record? They are quite important pieces of infomation actually; they are the information needed to find and identify the object being described. In fact, if the VMS was ever to be useful to future travellers, its actually amongst the most important pieces of information to record!  

Like, why are those pieces of information any more pointless to a european audience than "this foreign, unknown thing reduces fever"? I would have thought some of the objects described in the SBJ were also found in certain parts of europe too.

(11-09-2026, 03:54 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.So he probably told the Dictator to skip those three fields, and he himself omitted any comment that he did understand and felt that it was pointless to copy.

Or maybe he did copy the whole SBJ, with all those fields and comments; but removeed them when, back in Europe, he prepared the final draft for the Scribe.

These are two very different possibilities with two different sets of assumptions, as I laid out in my message. If the premise is that there were omissions, there should be some reasonable underpinning logic to how and why the omissions occured. To be clear, you are removing parts of the SBJ text based on your interpretation of the exact motivations, frame of mind, and abilities of a theoretical unknown traveller 600 years ago. Those variables are different depending on which scenario above you choose. 

The methodology on this just seems a little shaky, honestly. The entire original foundation of this theory was that there was a word match between the SPS and SBJ, and that the general wordcount and distribution of words of both works correlated in some way. But now that it seems that this is apparently not the case, the original content has to be constantly altered to create matches. Where is the line where this becomes unreasonable? 

If you omit the sections and the matches somewhat improve, but still don't convincingly match, what will change next? Will the order of the text start getting changed to create word matches? Every alteration you make lowers the validity of the match, so they should be done extremely sparingly in my view. I'm concerned that this will reach a conclusion analogous to: "look, the square peg does fit in the triangle hole! (if you cut the corners and shave it down and force it through)". 

And by the way, I say concerned because i've been genuinely interested in this theory, and I think it could have merit in some way (I disagree on many specifics, but find it plausible that the SBJ could have been copied in some way) and I would love to see it be proven if it's true. I just think there is a potential slippery slope here.
(11-09-2026, 10:11 AM)eggyk Wrote: You are not allowed to view links. Register or Login to view.
(11-09-2026, 03:54 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.He must have understood all three words 味甘平 = "flavor sweet balanced", and must have quickly learned what they meant and that recording them would be pointless.  Ditto for 生山谷 = "grows mountain valley".

Now here is an example of why I find this problematic. Why is "flavor sweet balanced" or "grows mountain valley" a pointless thing to record? They are quite important pieces of infomation actually; they are the information needed to find and identify the object being described. In fact, if the VMS was ever to be useful to future travellers, its actually amongst the most important pieces of information to record!

Those two bits of information are less than worthless: they are misleading.

That "flavor" field, introduced by the character 味, has two parts. First the "flavor" proper, which can be 辛 = "acrid" (or "spicy", "pungent"), 甘 = "sweet", 酸 = "sour", 苦 = "bitter", or 咸 = "salty"; sometimes modified by "very","slightly" etc.  However, that is not actually the taste of the substance, but a code for which organs it targets and what general effect it has on the body.  For example, a "sweet" medicine is supposed to to act mainly on the spleen and stomach (the "digestive center" in Traditional Chinese Medicine theory), generally reinvigorating, and stabilizing the qi, etc.  A "sour" medicine is supposed to target mainly the liver and the gallbladder, generally stopping losses and leaks like excessive sweating, diarrhea, etc.  And so on.  

Thus the "flavor" in the SBJ is like the "flavor of Linux".

For example, the "Red rooster" entry has eight sub-entries, for various products derived from the chicken:  meat, head, fat, intestines, gizzard lining, white part of the poo, quills, and eggs.  But since those products all come from the same animal, they are all assumed to have the same "flavor" -- "sweet".

The second part of the "flavor" field could be called the "thermal character", and can be 热 = "hot", 温 = "warm",  平 = "neutral" (or "balanced"), 凉 = "cool", or 寒 = "cold".  Again, these do not refer to actual temperature or anything, but are codes for the general effect on the "yin" vs "yang" state of the body.   Hot and warm medicines are supposed to get the qi moving and shifting the body towards yang; whereas cool and cold ones generally calm down the qi and shift towards yin.  The Red rooster products are all said to be "balanced" with respect to yin and yang, whereas ginger and cinnamon are "hot", Apples are "cool", and Dandelion leaves are "cold".

Thus this part of the 味 field too is pretty useless without knowledge of TCM theory -- which is NOT included in the SBJ.  The interpretation of that field would have been learned by doctors from other sources.  (Hmmm... could perhaps the Bio section of the VMS be a transcription of such a source?) 

Similar considerations apply to the "provenance" field, introduced by 生.  This character has many different senses in general, but in that field of the SBJ entries it means "grows in" for plants, "breeds in" for animals, and "found in" for minerals.  

But when the SBJ authors wrote "grows in mountains", they were not trying to tell the reader where to find the plant, but rather to hint that it was a particularly potent medicine because it grew on mountains.  

The Red rooster is listed as growing in 平泽 ="marshy plains"; which of course is not literal, since chickens were grown everywhere.  What that note is saying is that, since chickens were thought to have originated in low-altitude marshes, any chicken-derived medicine should tend to be good at moving the qi downwards and manage internal fluid accumulations.  Or something like that; I don't think I got the logic right.

Which again is a belief that Chinese doctors had (or may have had centuries before the Author got the SBJ).  So,again, that field would be meaningless and useless to an European doctor in Europe.

Quote:Like, why are those pieces of information any more pointless to a european audience than "this foreign, unknown thing reduces fever"? I would have thought some of the objects described in the SBJ were also found in certain parts of europe too.

Some of the remedies indeed are plants, minerals, or animals that are available in Europe, either natively (like licorice, dandelion, chicken poo) or imported (like ginger, black pepper, cinnamon).  The Author may have known their names and recognized them during dictation.  Perhaps he hoped to identify more of them after the fact, by asking doctors or visiting apothecary shops or farm markets.  "Ah, so this is dà suàn!  We call it "knofl" where I came from!"  If he did not recognize the herb, he may have sketched the leaves and/or roots, hoping that a doctor back home would identify them.  (Hmmm... could that be that what the Pharma section is?)  

Quote:
(11-09-2026, 03:54 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.So he probably told the Dictator to skip those three fields, and he himself omitted any comment that he did understand and felt that it was pointless to copy.  Or maybe he did copy the whole SBJ, with all those fields and comments; but removed them when, back in Europe, he prepared the final draft for the Scribe.
These are two very different possibilities with two different sets of assumptions, as I laid out in my message.

I don't see much difference.  He must have omitted them because he (rightly) realized that they were not worth writing down.  He may have realized that before, during, or after the dictation.  I don't see how we can tell which.

Quote:To be clear, you are removing parts of the SBJ text based on your interpretation of the exact motivations, frame of mind, and abilities of a theoretical unknown traveller 600 years ago.

No, I am removing those three fields because they are clearly NOT transcribed in the SPS.  

If they had been transcribed, they would stand out like sore herrings with red thumbs.  Practically every paragraph would have the same "味" word occurring after the 1-3 words of the title, and that word would be followed by 1-4 strings from a very small set (the five "flavors" and five "temperatures" plus a few intensifiers), and then by the "主" ("mainly for") word, which may occur several times for entries with sub-entries.  

Likewise most recipes would end with a "生" ("grows in")  word followed by two other words, also from a small set ("mountains", "valleys", "plains", etc.).

The pattern of occurrences of daiin in the SPS are generally clear and consistent with the occurrences of 主 ("mainly for") in the SBJ.  On the other hand, there is no trace of patterns in the SPS that would match the patterns of 味, 生, and 一名 in the SBJ.

Quote:The entire original foundation of this theory was that there was a word match between the SPS and SBJ, and that the general wordcount and distribution of words of both works correlated in some way. But now that it seems that this is apparently not the case, the original content has to be constantly altered to create matches.  Where is the line where this becomes unreasonable?

Again, I understand that people find this unconvincing when they just look at those adjustments and alleged "scribal errors", without trying to evaluate how good the keyword gap matches actually are, and how much effect those adjustments could have on them.

Again, I am assuming that the "flavor", "another name", and "provenance" fields were omitted from ALL entries.  Not on a recipe-by-recipe basis.  I don't have that freedom.

For instance, in the Rooster recipe there are eight 主 which divide the recipe in nine parts.  The deletion of the "flavor" field 味甘平 affects only the first gap, before the first 主, while the deletion of the "grows" field 生山谷, the "amber is divine" comment, and the "white grubs" sub-entry affect only the last gap.  

Deletion of the comment 东门上者尤良 = "those heads hung on the East Gate ..." affects the second gap, shortening it from 8 hanzi to 2.  That leaves six gaps that are not affected by these deletions -- which closely match the corresponding gaps in the SBJ text, at the ratio of ~5 EVA letters per hanzi.  These include the 21-hanzi gap between the first and second 主, which matches the distance between the first and third daiin with an error of ~5 EVA letters in 110.

All the best, --stolfi
hello^^

im still new here so please dont take it to heart, but can the voynich manuscript have alphasyllabary? if theres any false infromation that i said you can correct me^^ this is just my opinion.
(12-09-2026, 01:58 AM)A.R.B.C.J Wrote: You are not allowed to view links. Register or Login to view.hello^^

im still new here so please dont take it to heart, but can the voynich manuscript have alphasyllabary? if theres any false infromation that i said you can correct me^^ this is just my opinion.
In the context of Asian theory, I think it’s quite possible that Voynichese works on the same principle. Perhaps the combinations of symbols produce new sounds.
(08-09-2026, 10:26 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.For N  = 30, K = 8, M =15, J = 8, this formula gives P = ~0.0011 = ~0.11%; that is, less than one chance in 900.  So getting all 8 coins by picking 15 out of 30 boxes IS a remarkable result.
Why N = 30 in this post?
(16-09-2026, 09:04 AM)rikforto Wrote: You are not allowed to view links. Register or Login to view.
(08-09-2026, 10:26 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.For N  = 30, K = 8, M =15, J = 8, this formula gives P = ~0.0011 = ~0.11%; that is, less than one chance in 900.  So getting all 8 coins by picking 15 out of 30 boxes IS a remarkable result.
Why N = 30 in this post?

Those numbers were primarily intended to show that the probability computation must take N and K into account, and that the fact that comb(15,8) is a large number does not mean anything by itself.

But indeed the numbers were chosen to be roughly adequate for the Rooster recipe case.

If the claim had been that the mapping was exact, the number N of boxes in the box and coin analogy would be the number of Chinese characters, 72.  

But the gap sizes are allowed to differ from the predicted values, with average deviation of 2-3 EVA letters, which is about half a Chinese character.  Thus, to make the analogy valid, the boxes must be wider than just one character, and then there will be fewer boxes.  Hence the use of N = 30 in that example.

Quote: Can you provide the precise expression you used to arrive at [the gap error scores]?

The formulas I am currently using are given below.  But they are not particularly special.  The only important aspect is that the score for each gap is proportional to the square of the relative error between the actual length in EVA letters and the length predicted from the hanzi count in the corresponding gap of the recipe.  Specifically, the square of the difference of the logs. 

[attachment=17660]

Quote:The other thing I'm going to want to know is what the badness cutoff is---when is that badness too high to count as a match?

Since the scale of the badness score is somewhat arbitrary, there is no cutoff.  I currently use the score only to rank the candidate parags.

The decision of whether the best-ranked (minimum-score) candidate is probably the real translation of the recipe is still done by eye, looking at whether the identified Voynichese keywords were the "canonical" ones (like daiin instead of dain or kaiin) and how big really are the gap errors (they should all be small, mostly 2-3 EVA letters at most).  Thus the mappings between recipes and parags are only tentative.  

That said, with the current formula, when the best candidate has badness score greater than 1.0, it is usually not a real match.  Sometimes I am lucky and there is only parag that matches all the cribs well, with score below 1, and all other parags have mush larger scores.  Sometimes there are a few parags with scores below or just above 1, with no clear gap.  Sometimes the best parag is obviously not it.

Eventually I intend to replace the ad-hoc badness score above with a proper statistical one (log of 1/prob of the pairing being true according to Bayes).  Then one could list the pairs that have 80% or 99% confidence etc.

For the Rooster recipe, there is no need to choose among the ranked candidates.   There is only one candidate paragraph to consider, because both are outliers in the recipe/parag length distributions.  Thus any other parag, being much shorter, would have at least one gap that is way too small.

--stolfi