The Voynich Ninja

Full Version: The 'Chinese' Theory: For and Against
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
(03-09-2026, 10:01 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.
Quote:The hypothesis that a text is written in an awful semi-phonetic, inconsistent version of some Asian language is therefore unfalsifiable.
I would agree. You can butcher the language unknown for you so much that it indistinguishable from gibberish.

I understand that it looks that way -- if you don't bother to look at the details: how much "butchering" is actually allowed, and how likely it is that the matches are just coincidences.

Unfortunately the VMS IS full of errors -- many more than one finds in a typical Medieval manuscript.  Because of what it is, and how it was created. 

But scientists know that, given enough data, it is possible to reliably extract a faint signal from a torrent of loud noise.  That is how they prove that a new sub-atomic particle exists.  That a snow white star a hundred trillion km from us is surrounded by seven invisible dwarf planets.  That the solid core of the Earth, 5500 km deep and under more than 2000km of ultra-hot liquid iron alloy, was rotating about 0.2 degrees per year faster than the crust in the 1970s, but has been slowing down, and since 2010 it has been rotating more slowly than the crust. 

Or that a text which had 95% of its letters replaced at random by other letters is a translation of Hamlet in Albanian -- but with a happy ending.

So I am still extracting recipes from the Zenghe Bencao, running my matching programs on them, and looking for more cribs.  Eventually the mass of results will be un-dismissible...

All the best, --stolfi
(03-09-2026, 10:01 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.A funny example - in Indiana Jones movie an American lady tries to sign in Chinese. From the comments you can learn that it is totally impossible to understand for the real Chinese speakers and someone did a hard, detective work to get what she meant to say  Smile

Well, I remember my wife trying to order 'a cup of tea' in England, using her phonetical interpretation of what English sounds. They just could not understand what she wanted.  She was then pretty amazed  when I asked for the tea saying what, in her mind, sounded exactly the same... and the tea came.
(03-09-2026, 10:01 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.A funny example - in Indiana Jones movie an American lady tries to sign in Chinese. From the comments you can learn that it is totally impossible to understand for the real Chinese speakers and someone did a hard, detective work to get what she meant to say  Smile

Well, I remember my wife trying to order 'a cup of tea' in England, using her phonetical interpretation of what English sounds. They just could not understand what she wanted.  She was then pretty amazed  when I asked for the tea saying what, in her mind, sounded exactly the same... and the tea came to the table.

Morale: a phonetic transcription of a foreign language will be, at best, an incomprehensible mess, unless one knows a lot about phonetics.
(04-09-2026, 12:43 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.
(03-09-2026, 10:01 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.A funny example - in Indiana Jones movie an American lady tries to sign in Chinese. From the comments you can learn that it is totally impossible to understand for the real Chinese speakers and someone did a hard, detective work to get what she meant to say  Smile

Well, I remember my wife trying to order 'a cup of tea' in England, using her phonetical interpretation of what English sounds. They just could not understand what she wanted.  She was then pretty amazed  when I asked for the tea saying what, in her mind, sounded exactly the same... and the tea came to the table.

Morale: a phonetic transcription of a foreign language will be, at best, an incomprehensible mess, unless one knows a lot about phonetics.

I have this with a French friend.. I say something in French, he corrects me and repeats the exact same thing back... It's maddening... Our favourite word is Mile Feuille... He explained that Feuille should be pronounced like a plastic ruler being sprung against a desktop... Big Grin
(04-09-2026, 03:59 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view....

PS. The "SPS = SBJ" claim is quite different from most other "plain natural language" theories.  These typically propose some obscure language and spelling system with arbitrary mangling, abbreviation, dialectal variation etc., and then are free to make up some bizarre recipes or stories.  Like "Boil [the] Sacred Sword under [the] blood of five stones [and] add saltpeter fast [while] grinding [the] strong eagle [as it] flies through [the] cosmic hill".   

The "SPS = SBJ" claim, in contrast, does not specify the language, but specifies the complete plaintext -- even that of the eight missing pages. Not just some classic Chinese medical book, but that specific book. Essentially whole, mostly word for word.  It does not have the freedom to choose between boiling swords or grinding eagles at each step.

The only freedom it has is the order of the SBJ recipes.  That was never fixed, and medical encyclopedias that quoted the whole SBJ would rearrange the recipes to suit their own ordering scheme.  The Dictator must have read them in the order that they appear in his book.  If the dictation happened in China, the book probably was the Zenghe Bencao but if it happened in some other country it may have been a local encyclopedia.  Anyway, the Author may have scrambled the leaves of his draft, or may have reordered the recipes according to his logic, before giving the final draft to the Scribe.  And then the bifolios may have been scrambled and flipped before being bound.  Still, I believe that vestiges of the Zhenghe Bencao order may have survived, and I think I saw hints of that.  More on this later.

On the other hand, the close match between the histograms of SBJ recipe lengths and SPS paragraph lengths implies that each Chinese character became about 5 EVA letters (excluding word spaces).  That is a first constraint on the matching of recipes to paragraphs: I only consider candidates that are close to the size predicted by that ratio. For example, for a middling-sized recipe with 34 original Chinese characters ("heavenly essence herb"), only 21 of the 243 single-star parags pass this first criterion.   

For the longest recipe ("red rooster"), in particular, that criterion already leaves only one possible match, namely the longest single-star paragraph (f105v.32) -- because both are much longer than all the other competitors.  Thus the SPS=SBJ claim necessarily includes the claim that the plaintext of parag f105v.32 is the "red rooster" SBJ recipe.  

This specific match is determined even before looking at the positions of the 主/daiin and other cribs.  Fortunately these do match quite well, as shown before (and recapped below).

As for how much "butchering" is allowed: the SPS=SBJ claim indeed
  • allows/requires deletion of certain parts of the Chinese text, and 
  • allows some variation on the Voynichese side of the cribs.
As for the first kind of "butchering", specifically, there are three fields at the beginning and end of the recipes (the classification in Chinese medical theory, alternative names, and the general provenance) that are assumed to have been omitted by the Author on all recipes. These fields are identified by specific Chinese keywords  (味, 一名, and 生, respectively), and there is no uncertainty about their extent.  In addition, in some recipes there may be some "comment" fields that are not names of diseases or benefits.  In the "red rooster" recipe:
  1. the "rooster head" sub-recipe has a comment 东门上者尤良 = "those [rooster heads] that were hung on the East gate are particularly potent [at killing ghosts] ".
  2. there is a note 女人 = "woman"  before three conditions specific for women
  3. after the note that rooster eggs (!) can be alchemically turned into amber, there is a comment 神物 = "[which is] a divine substance"
  4. the very last sub-entry 鸡白蠹能肥脂 = "chicken white grubs fatten fat", besides lacking a 主 keyword, was already incomprehensible by ~1300, and may be about grubs that grow in chicken manure -- not a part of the "red rooster".
Comment (1) is assumed to have been omitted by the Author.  A priori, the other three comments may or may not have been omitted; so that is a bit of "butchering freedom" that the claim has on the Chinese side.  Trying all combinations indicates that the "white grubs" part (4) was definitely omitted, and the other two comments (2,3) were probably (but not surely) omitted too.

On the Voynichese side, the "butchering freedom" is in the translations of the eight 主 keywords and their positions on the text.  Four of the 主 appear as daiin, and the other four can be matched to kaiin, dain, laiin, and kaiin.  As I have argued before, these variations from daiin are all plausible errors by the Author or the Scribe (if they are not equivalent spellings of the same sound).  As for their positions, the counts of EVA letters between consecutive matched cribs are 8, 100, 15, 22, 34, 20, 43, 23, and 52; which differ from the expected counts by -6, -3, 0, +7, 0, 0, +3, -1, and +2 EVA letters.  Looking at the Chinese side, the gaps are 3, 21, 3, 3, 7, 4, 8, 5, and 10 hanzi, and EVA count errors above are equivalent to -1.2, -0.6, 0, +1.4, 0, 0, +0.6, -0.2, and +0.4 hanzi. So the "butchering freedom" here is a standard deviation of ~4 EVA letters (equivalent to 0.8 hanzi) from the expected spacings of the 主 cribs.

Is that to much "butchering freedom"?  Sorry, but I don't think so...

All the best, --stolfi
(04-09-2026, 12:43 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.Well, I remember my wife trying to order 'a cup of tea' in England, using her phonetical interpretation of what English sounds.

Some years ago I went to a conference in Huddersfield (Yorkshire UK) and stayed in the students dorms.  I asked the girl at the front desk for a towel.  It took a few tries, but she eventually understood: "Ah! The Gentleman needs a taw-eel!"

All the best, --stolfi
(04-09-2026, 02:11 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.As for their positions, the counts of EVA letters between consecutive matched cribs are 8, 100, 15, 22, 34, 20, 43, 23, and 52; which differ from the expected counts by -6, -3, 0, +7, 0, 0, +3, -1, and +2 EVA letters.  Looking at the Chinese side, the gaps are 3, 21, 3, 3, 7, 4, 8, 5, and 10 hanzi, and EVA count errors above are equivalent to -1.2, -0.6, 0, +1.4, 0, 0, +0.6, -0.2, and +0.4 hanzi. So the "butchering freedom" here is a standard deviation of ~4 EVA letters (equivalent to 0.8 hanzi) from the expected positions of the 主 cribs.

Is that to much "butchering freedom"?  Sorry, but I don't think so...
This account leaves out that the "butchering freedom", as we're calling it, over-generates potential Voynichese cribs. The Rooster example has 17 cribs in the SPS and 10 in Shennong's Classic. Selecting 10 items from 17 where the order of the selection doesn't matter (because we are going to keep them in the order they appear in the SPS) is 17C10 = 19,448 potential matches. Because the output of this function is very sensitive to inputs, this must be tracked carefully. A smaller recipe that generates 10 SPS cribs for 7 hanzi has "only" 120 choices, but if there are two fewer hanzi cribs it exceeds double that. This smaller finding is offset somewhat by the fact that middling recipes have more potential matches; you give 21 for a particular recipe, and the sum of them may approach a large fraction of 19,448.

If the size is 2,520  (= 21*(10C7)) or 19,448 is going to have to be pinned down to make a precise evaluation of the odds, but the fact that the "butchering freedom" generates these large search spaces and then you search for one that minimizes the errors has huge consequences for claims about how likely it is that the errors here are small! I did take a look at the problem a few weeks ago, and there is a tendency for your Voynich cribs to significantly exceed the number of Chinese ones, both on average and between proportionally similar entries. I might do this in more detail when you publish a clearer picture of what texts you're using, how you got them, and the parameters of your algorithm. But I feel confident saying the algorithm generates a large match space and then selects a favorable result for your claims.

If I'm following your write-ups correctly, you are also running the matching algorithm repeatedly with different crib sets. Transparently, I am not completely clear on the how and why, so it is less clear to me how much this is skewing things. But I can say the decision to selectively exclude cribs is another analytical choice that makes the search space larger than you may realize. If you exclude some cribs and calculate the odds based on that exclusion, you are not calculating the odds for everything you searched. The fact that you have to make the exclusion is important information!

There are two points here. The first is that these are analytical choices. They are of a different quality from the last-mile arbitrary interpretations some solvers do, and I do find that sincerely laudable. They are, nonetheless, choices you as the analyst are making. They reflect a pair of texts that you have created based on judgements you have made. Second, the way those choices distort the probability of the match is absolutely insensitive to the strength of your arguments for making these analytical choices. (I am purposely not disputing those choices in this post, which is not to say I think they are all bulletproof.) The fact that you can discard thousands if not 10s of thousands of similarly defined matches must be accounted for in your argument of the odds even if---perhaps especially if!---those choices "make sense". You are correct that scientists can carefully process data to make a match, but that work requires extremely careful guardrails to make sure they are not effectively creating the data.

That a number of your choices are invisible to you when evaluating the strength of the match surely means you are underestimating the degree to which you are analyzing changes you have made to the two texts, not the texts themselves!
eggyk Wrote:
"Jorge_Stolfi” Wrote:That's a very interesting experiment.  I should try a similar one: post an audio clip in some language and ask for a volunteer who does not know how to write that language to try to write down the sounds they hear, using whatever spelling system they chooses, and then see how consistent the result was.  Namely, if they always wrote the same word in the same way.

To better approximate the dictation scenario of the Chinese Theory, the volunteer ideally should be somehow familiar with the spoken language, even partly understand it, but they have no idea of its official spelling or even the nuances of its phonetics.  Maybe playing Portuguese to an Italian speaker, Dutch to an English speaker, Czech to a Polish speaker...

All the best, --stolfi

This would have to be done one word at a time and clearly to have any chance. I just brought up a video in portuguese to test (I cannot understand it, but am familiar with its ties to spanish and other romance languages so i get a word here and there) and it's a complete non starter to even try to write it down. 

You are not allowed to view links. Register or Login to view.

I got to about 20 seconds in, and i've been repeating this over and over at 0.5x speed, and its almost completely impossible. Here is my text, does it hold any actual information still? (do not laugh at how obviously horrible this is)

omdilia trente dewes unyos sot brazil de terut sao paolo maz veskiturzian skoviv un paranaon je fustudar uuh furomacum proffesor je portugues espanyol nao apremereves kuv ... uduguay jevinu truz vez, ee kiminkun tenten viriki

and the auto-transcript:
You are not allowed to view links. Register or Login to view.

This is really interesting because, as an Italian who studied Spanish in middle school and only has a basic understanding of it, I can confidently say that I understand roughly 70–80% of what the girl in the video is saying without any problems, even though I’ve never studied or really had any exposure to either Brazilian or European Portuguese.  What’s your native language? What other languages do you speak? Just curious about it  Smile
I went to look at the video and it gave me subtitles, so don't look, just in case you are trying to write it down yourself (it may just be my settings...)
(06-09-2026, 10:16 PM)Yavernoxia Wrote: You are not allowed to view links. Register or Login to view.This is really interesting because, as an Italian who studied Spanish in middle school and only has a basic understanding of it, I can confidently say that I understand roughly 70–80% of what the girl in the video is saying without any problems, even though I’ve never studied or really had any exposure to either Brazilian or European Portuguese.  What’s your native language? What other languages do you speak? Just curious about it  Smile

I speak (British) English as my mother tongue, and Dutch as a second language. I did learn a decent amount of french back in school but i've lost most of it. Lets just say that in cases where i'm trying to work out the instructions on food packaging with no english or dutch, I can just about get almost everything by looking at the french and german instructions and mixing the words I understand together.  Big Grin With other romance languages, I understand a few shared latin words like estudar, informatica etc. 

When I look at that correct transcript, completely honestly without google translating I get something like: 
"... my name is marilia ... (my age is?) 32 years ... in? brazil in the interior (centre?) of sao paulo ... ... 14 years that? live no ... ... ... study ... what professor of portguese and spanish ... is .. primary/first? ... that I go? ... Uruguay ... ... ... other? places? and that something's me ... ... here?"

but with the german, again without auto translate: 
"'Home'? In german 'Home' has basically two (main meanings?), as far as I understand. And heavy? someone can say: 'Home' in the sense 'I am at home', was basically meant as: 'I am ... there?, where? I ... live'. Many are ... not so very therewith? connected, where they ... live. And then that gives ... also 'Home' in the sense of: 'I go home to my ...', or..."

I would also say that I understood more than I could phonetically write with the portuguese example. I wrote a butchered "trente dewes unyos", but did genuinely understand "thirty-two years" when I heard it. With the german example, my knowledge of dutch is basically carrying the entire thing: "viele = veel", "wohnen=wonen", "Noch = nog", "eigentlich = eigenlijk","ich gehe nach = ik ga naar" etc etc. I imagine that every person doing this will have a different experience. Perhaps, if you don't speak german, you will struggle more with that compared to the portuguese?