The Voynich Ninja

Full Version: The 'Chinese' Theory: For and Against
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
(07-09-2026, 10:20 AM)eggyk Wrote: You are not allowed to view links. Register or Login to view.
(06-09-2026, 10:16 PM)Yavernoxia Wrote: You are not allowed to view links. Register or Login to view.This is really interesting because, as an Italian who studied Spanish in middle school and only has a basic understanding of it, I can confidently say that I understand roughly 70–80% of what the girl in the video is saying without any problems, even though I’ve never studied or really had any exposure to either Brazilian or European Portuguese.  What’s your native language? What other languages do you speak? Just curious about it  Smile

I speak (British) English as my mother tongue, and Dutch as a second language. I did learn a decent amount of french back in school but i've lost most of it. Lets just say that in cases where i'm trying to work out the instructions on food packaging with no english or dutch, I can just about get almost everything by looking at the french and german instructions and mixing the words I understand together.  Big Grin With other romance languages, I understand a few shared latin words like estudar, informatica etc. 

When I look at that correct transcript, completely honestly without google translating I get something like: 
"... my name is marilia ... (my age is?) 32 years ... in? brazil in the interior (centre?) of sao paulo ... ... 14 years that? live no ... ... ... study ... what professor of portguese and spanish ... is .. primary/first? ... that I go? ... Uruguay ... ... ... other? places? and that something's me ... ... here?"

but with the german, again without auto translate: 
"'Home'? In german 'Home' has basically two (main meanings?), as far as I understand. And heavy? someone can say: 'Home' in the sense 'I am at home', was basically meant as: 'I am ... there?, where? I ... live'. Many are ... not so very therewith? connected, where they ... live. And then that gives ... also 'Home' in the sense of: 'I go home to my ...', or..."

I would also say that I understood more than I could phonetically write with the portuguese example. I wrote a butchered "trente dewes unyos", but did genuinely understand "thirty-two years" when I heard it. With the german example, my knowledge of dutch is basically carrying the entire thing: "viele = veel", "wohnen=wonen", "Noch = nog", "eigentlich = eigenlijk","ich gehe nach = ik ga naar" etc etc. I imagine that every person doing this will have a different experience. Perhaps, if you don't speak german, you will struggle more with that compared to the portuguese?
I absolutely struggle more with German. I can somewhat understand maybe 15% of the transcription you wrote here, and that’s mainly because some words are similar to Italian, Russian, or French, which are languages I speak.  I also tried to write it down myself, but the result was so bad that I honestly gave up altogether. I’d say it was probably 0–5% accurate at best.
Again for the (still forthcoming) write-up about some of the issues with your analysis of the dictation scenario. I have two questions about this:
[
(07-09-2026, 12:57 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Me: Please give the full reading again as a single line, without the hanzi in parentheses. Isn't there a missing zyu2 in your response above?

Gemini: You are absolutely right. There was a zyu2 (主) missing right after daai6 wong4 fung1 zi2 (大黄蜂子), and the string had a few spacing and character-matching errors.

Here is the fully corrected, syllable-by-syllable Cantonese Jyutping reading of your Bencao text as a single, continuous line without any Hanzi or parentheses:


fung1 zi2 zyu2 fung1 tau4 coi4 gu2 duk6 bou2 heoi1 leoi4 soeng1 zung1 gau2 fuk6 ling6 jan4 gwong1 zaak6 hou2 ngaan4 sik1 bat1 lou5 daai6 wong4 fung1 zi2 zyu2 sam1 fuk1 zoeng3 mun5 tung3 hing1 san1 jik1 hei3 tou2 fung1 zi2 zyu2 jung1 zung2
Questions:
  1. Are you consistently specifying "Cantonese reading" when you prompt for output like this?
  2. Are you consistently asking it to provide it syllable-by-syllable?
(07-09-2026, 01:19 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.Are you consistently specifying "Cantonese reading" when you prompt for output like this? Are you consistently asking it to provide it syllable-by-syllable?
For that sub-thread, that started with You are not allowed to view links. Register or Login to view., yes and yes. As I understood, that is what he transcribed in his original script.

At the moment, I am not trying to map the Voynichese glyphs to phonemes in Cantonese or any other language.  All the matching I do is based on direct mapping of hanzi to Voynichese strings, ignoring punctuation.  In posts and reports I give pinyin transcriptions of Mandarin readings only for the benefit of people who (like me) find it difficult to memorize  hanzi or distinguish 主 from 生 etc.  But those readings are not used in the matching.
(05-09-2026, 03:38 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.This account leaves out that the "butchering freedom", as we're calling it, over-generates potential Voynichese cribs. The Rooster example has 17 cribs in the SPS and 10 in Shennong's Classic.
The longest entry of the SBJ, the "Red Rooster", has 8 (not 10) occurrences of 主.  Excluding the fields that apparently were omitted by the Author in all entries, it is 72 hanzi long:

丹雄鸡崩中漏下赤白沃补虚温中止血通神杀毒辟不祥头杀鬼肪耳聋肠遗溺肶胵裹黄皮泄利屎白消渴伤寒寒热翮羽下血闭鸡子除热火疮痫痉可作虎魄

The longest single-star parag of the SPS is f105v.32.  It has 361 EVA letters in my transcription.  Note that 361/72 = ~5, a ratio that seems to be approximately the same for all recipes.  Ignoring spaces, there are 15 occurrences of daiin = daiin and its allowed variations, marked by brackets below:

<f105v.32>  poarkeeo[daiin]qoairaracphheyqoeedeodyqo[kaiin]qote[dair]aporairapylsheodytairoteeyoteeoolotaiinokeeyqo[kaiin]oraiiraldalsheeo[daiin]chsdqokeeey[dair]o[kaiin]otaiinche[daiin]olkallkl[dain]doeeokcheeoltaiinotcheedychoraiino[daiin]chedyotaiinalkaishd[laiin]sheodokeeodyqoaiinytaiinotairchdaldy[daim]ch[daiin]ockhhyysheyckhysheoqoeeol[kaiin]chsokoltchdysheeeyo[kaiin]araildycheodyoaiirainokshey

That's  5 [daiin], 5 [kaiin], 2 [dair], 1 [daim], 1 [dain], 1 [laiin].

I explained before why I allow these but not [taiin], [raiin], [chaiin] etc. Briefly, m is probably an abbreviation of iin or maybe in, and in sloppy cursive handwriting iin may look like ir, and d may look like an l or k (but not a t or other glyphs).

Besides those 15, there are 13 additional occurrences of the strings *aiin, *ain, *aim and *air, that I do not consider acceptable variations of daiin.

t turns out that, mapping the positions of the 主 in the SBJ recipe to positions in the SPS parag, 4 of these fall close to a  [daiin], and 4 to variations thereof.  (More precisely, the gaps between these 8 nearest keys in the parag are very close to the gaps between the 主, scaled by 5.)

Quote:Selecting 10 items from 17 where the order of the selection doesn't matter (because we are going to keep them in the order they appear in the SPS) is 17C10 = 19,448 potential matches.

Imagine that you have a row of 30 boxes, and 8 of them have a coin inside,while the others are empty.  You pick 15 of those boxes at random.  You find that 8 of those 15 boxes have a coin inside.

There are comb(15,8) = 6435 subsets of 8 boxes out of the 15.  That is the number of ways you could have got all the 8 coins that way.  That is a big number. Does it mean that this outcome is unremarkable?

Of course not.  That number alone does not mean anything.  To answer that question one must consider also the total number of boxes and how many of them have coins.

The probability P that picking M of N boxes at random will yield J of the K coins they contain is comb(K,J)*comb(N-K,M-J)/comb(N,M).   

For N  = 30, K = 8, M =15, J = 8, this formula gives P = ~0.0011 = ~0.11%; that is, less than one chance in 900.  So getting all 8 coins by picking 15 out of 30 boxes IS a remarkable result.

How is this relevant to the "Rooster" matching? Hint: the coins in the boxes have a "主" stamped on them.

Sure, the computation of the "null hypothesis" P for the Rooster=f105r.32 claim is more complicated than for that analogy.  

For one thing, I don't compare the absolute positions of daiin-like strings and 主 characters, but the lengths of the 9 gaps before, between, and after the eight 主 characters and the gaps between the eight assigned daiin-like strings, scaled by the 1:5 ratio.  

The discrepancies seen in these gap sizes are almost all less than 5 EVA letters, which is 1 hanzi; but two of them are -7 letters (in the title of the recipe, which is 3 hanzi but only 8 EVA letters) and +9 letter (in a gap that is 46 letters long).  These two discrepancies are equivalent to -1.4 and +1.8 hanzi.  For comparison, the average gap size is about (361-8*5)/9 = ~35 EVA letters.  This tolerance in the gap lengths affects the number N of "boxes" that are available to be "picked" by the M = 15 daiin-like strings.

But anyway you should be able to see that the null-hypothesis P will be quite small,  even though comb(M,J) is a large number.

All the best, --stolfi
(05-09-2026, 03:38 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.f I'm following your write-ups correctly, you are also running the matching algorithm repeatedly with different crib sets. Transparently, I am not completely clear on the how and why.

Some fields of the SBJ recipes were apparently omitted by the author on all recipes.  Those are the "nature", "another name", and "provenance" fields, introduced by the keywords 味, 一名, and 生, respectively.  Usually the first is right after the entry title, and the other two are at the end of the entry; but when an entry is divided into sub-entries, they may occur at the beginning or end of a sub-entry.  

Anyway, those fields and their extents are trivially identified.  So I always delete them from the SBJ entry before looking for a match.

I also delete unconditionally some comments that would make no sense outside China (or even inside it). Like that note about the heads of chickens that were hung above the East Gate.

But some recipes, including Red Rooster, have some parts that just could have been omitted. In the Rooster recipe, these are the notes 女子 = "[for] women", 可作虎魄="[eggs] can be [alchemically] turned into amber", "神物" = "[amber] is a divine substance", and 鸡白蠹能肥脂 = "chicken white grubs can fatten fat".  Thus, I must try matching the recipe against the SPS parags with and without each of these items.  That is one component of the "variants" parameter in my program.

(It turns out that, in the case of the Rooster recipe, matching is impossible if 可作虎魄神物 and 鸡白蠹能肥脂 are both excluded or both included.  Thus the 16 possibilities reduce to 6: with just one of 可作虎魄, 可作虎魄神物,  or 鸡白蠹能肥脂, and with 女人 either included or excluded.  The three-way choice affects only the length of the gap after the last 主, and the best matchings by far are obtained with 可作虎魄.  The inclusion or exclusion of 女人 affects only the gap between the first and second 主, which is 21 hanzi long; the corresponding EVA gap is 3 letters to long if  女人 is excluded, 6 letters too short if it is included (compared to the expected ~110 letter).  So this choice is basically a meh.)

I am also trying to identify additional cribs.  At this point, the only two pairs I am fairly certain about are USES: 主 = daiin and QI: 气 = chedy, both with a small set of variant spellings or common typos.  So I typically rum my programs on some recipes with just USES, or just USES and QI, to identify the likely matching parags, and then run again with additional tentative cribs like 令 "makes", 血 "blood", 杀 "kill", 除 "eliminate", etc., hoping to identify their Voynichese equivalents.  These trials are a second component of the "variants" parameter.

(I am also considering deleting all qo from the SPS parags before testing for a match, justified by my hunch that this "digraph" is a symbol like "&" added by the Author.  Preliminary tests suggest that this generally improves the matching.)

All the best, --stolfi
(08-09-2026, 11:36 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.But some recipes, including Red Rooster, have some parts that just could have been omitted. In the Rooster recipe, these are the notes 女子 = "[for] women", 可作虎魄="[eggs] can be [alchemically] turned into amber", "神物" = "[amber] is a divine substance", and 鸡白蠹能肥脂

That just made me laugh. Only a Chinese person could write such nonsense.
Amber and alchemy. Amber literally floats towards you and you collect it with a net.
A European would have known that. Wink
You are not allowed to view links. Register or Login to view.
You are not allowed to view links. Register or Login to view.
You are not allowed to view links. Register or Login to view.
How reliable are the definitions of the characters for daiin and chedy? I mean that these words are among the most common (daiin is the most frequent) words in VMS. They could just as easily be compared to any other common word from another text. And why can’t we match other characters from the recipe, knowing the arrangement of the characters daiin and chedy?
(09-09-2026, 10:26 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.How reliable are the definitions of the characters for daiin and chedy? I mean that these words are among the most common (daiin is the most frequent) words in VMS. They could just as easily be compared to any other common word from another text. And why can’t we match other characters from the recipe, knowing the arrangement of the characters daiin and chedy?

Not all daiin are 主, and not all chedy are 气.  But indeed 主 and 气 are very common characters in the SBJ.  

The character 主 = "mainly [for]" is used to introduce the list of conditions that the remedy can treat; thus there should be at least one 主 in every entry.  Some entries contain two or more sub-entries, and each should have its own 主.  The "Red rooster" recipe has eight 主; the matching SPS parag f105v.32 has five daiin (not counting the variants like dair, kaiin, etc.), four of which match the positions of four 主.  (The modern Mandarin reading of that character is zhǔ in pinyin, which, according to my faithful lalamo, would sound close to "joo" or "ju-oo" in English spelling.  But Mandarin itself has changed a lot in the last 600 years.) 

The character 气 stands for a theoretical concept of traditional Chinese medicine (TCM) that has no adequate translation in English.  It is some impalpable fluid or essence that was thought to be produced by the "digestive center" (stomach and spleen) and flowed to other organs and muscles, enabling them work.  I have been translating it as "vital energy", but there are other equally unsatisfactory translations.  Many pathological conditions are attributed to anomalies in the flow of 气; for example, whooping cough seems to be explained as the 气 flowing up towards the lungs instead of down.  Many tonics are said to "boost the 气" or otherwise benefit it.  About 2/3 of the recipes in my file have at least one 气; the "Red rooster" recipe is one of the 1/3 that lack it. (The character is read in modern Mandarin as pinyin qì, which should sound like "chee" in English spelling.)

All the best, --stolfi
Quote:The character 气 stands for a theoretical concept of traditional Chinese medicine (TCM) that has no adequate translation in English.

Isn't there an standard translation, using the original word?
You are not allowed to view links. Register or Login to view.
(08-09-2026, 11:36 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Some fields of the SBJ recipes were apparently omitted by the author on all recipes.  Those are the "nature", "another name", and "provenance" fields, introduced by the keywords 味, 一名, and 生, respectively.  Usually the first is right after the entry title, and the other two are at the end of the entry; but when an entry is divided into sub-entries, they may occur at the beginning or end of a sub-entry.  

Anyway, those fields and their extents are trivially identified.  So I always delete them from the SBJ entry before looking for a match.

I also delete unconditionally some comments that would make no sense outside China (or even inside it). Like that note about the heads of chickens that were hung above the East Gate.

So, this topic of omissions has been mentioned a few times. Does the idea of the omission of certain sections make sense with the rest of your theory? I'm under the impression that the VMS text is supposed to represent the spoken SBJ written phonetically. Therefore, those omissions imply three possibilities (sub possibilities can be "and/or"):

1) That the source SBJ text was already lacking those sections. 
or
2) That the dictator omitted certain sections not relevant to the transcriber when dictating
  2b) That the transcriber was aware of the sections that exist in the SBJ despite not being able to read it
  2c) That the transcriber was able to communicate exactly which types of information they wanted to hear, and that the dictator should skip other parts
or
3) That the VMS author later omitted sections from the voynichese draft
  3b) That the VMS author strongly understood the spoken language when respoken aloud later, enough to be certain which sections were which.

I don't understand how you know which sections were apparently omitted, or how you know why. I'm getting the impression that these omissions are having to be included because the VMS was not matching to the original SBJ text as hoped. If so, this may be an issue when trying to validate this theory, because these paragraphs are no longer being matched with the SBJ as written, but instead being matched with a version that has been altered until a match occurs. 

Of course, it's entirely possible that you are correct about what exactly was omitted, but it does require explanation and should be extemely consistent in methodology if the results are to be taken seriously. For example, "they always omitted this common section" can be consistently applied and reasonably argued, but "this probably wouldn't have meant anything to people outside China" sounds reasonable on the whole, but is fundamentally subjective and open to interpretation.

If I missed a post where you addressed this already, I apologise.  Smile