03-07-2026, 06:12 PM
(02-07-2026, 12:10 PM)Ruby Novacna Wrote: You are not allowed to view links. Register or Login to view....Continuing the previous post: the main part of those pages is the "Chosen Match" section, which lists the SPS parag that seems to be the best match for the entry -- usually the one that got the lowest "badness score" while using the most cribs. For that "bee larva" (BLAR) entry iy is
SBJ entry parsing:
<b1.4.100> (trim) 44
2 蜂子
1 10 主 风头除蛊毒补虚羸伤中
2 0 久服
1 12 令 人光泽好颜色不老大黄蜂子
1 8 主 心腹胀满痛轻身益
1 3 气 土蜂子
1 2 主 痈肿
SPS entry parsing:
<f114r.4> 0.833 237(+14) 0.350
13(+2) 0.044 fdeechdyopche
daiin 58(+7) 0.250 ypchedyodalychedyqopcheokaiinshedypodaiinochedalloiinchedy
qokaiin 0(+0) 0.000
chdy 63(+2) 0.019 daiindchdoseedolchdolkchedychokaiinchdyqoeedyokchedydaiiinchedy
daiin 40(+0) 0.000 ykeedyokeeedychedolchdaiinykardarycheold
chedy 17(+1) 0.017 dkchsaiinchdedyqo
daiin 15(+4) 0.153 okchedaiinchain
The top part is the SBJ entry, minus the standard omitted fields ([nature], [another name], [provenance]), divided into "hits" (the cribs used in the matching, bold) and "gaps" (the hanzi between those cribs). The cribs are 主 = "mainly for" (3x), 久服 = "extended use", 令 = "makes", and 气 = "qi", sort of vital energy or essence. The number 44 at the top is the length of the entry. The numbers in the following lines are the lengths of the "hits" (left column, bold) and and the following "gaps" (right column), all in hanzi.
The bottom part is the parag that seemed to be the best match. IThe number 0.833 is the "badness score" of the match. The second-best match (shown further up in the page) had a score of 5.032, which basically means "not a match". The next number 237 is the length of the parag in EVA letters (so Ch and Sh count as 2, CTh as 3, etc). The "(+14)" means that 237 is 14 EVA letters longer than expected given the length of the SBJ recipe, which was 44 x 5.06 = 223 (rounded). The 0.350 is the part of the badness score 0.833 that is due to imperfect matching of the hanzi and EVA cribs.
The following lines show the partition of the parag into "hits" and"gaps" that was assumed to correspond to the partition of the SBJ recipe. Thus the first 主 was matched to a substring "daiin" of the parag, the 久服 was matched to a "qokaiin", the 令 was matched to a "chdy", and so on. The integers in the second column are the lengths of the gaps, with prediction errors in parentesis. Thus the gap before the first matched "daiin" is 13 EVA letters, which is 3 letters more than the 2*5.06 = 10 letters expected given the 2 hanzi gap in the recipe before the first 主. This error is equivalent to 0.6 hanzi. The 58(+7) means that the gap from that first "daiin" to the "qokaiin" is 57 EVA letters, which is 7 more than expected given the 10 hanzi between 主 and 久服. And so on.
The third column is the part of the badness score that is due to the errors in the gap lengths. Each term is proportional to the square of the error; the first and last gaps have half the weight of the other gaps.
The extra penalty term 0.350 listed on the head line is the sum of 0.300, because my current guess for the correct Voynichese translation of 令 is "cheody" or "cheydy", but the program found only a "chdy" at the right place; plus 0.050, because my current guess for the correct translation of 久服 is the pair "(q)okeedy.(q)okaiin", but the program found only a "qokaiin" at the right place.
I am still trying to find more cribs, and figure out what is really the correct Voynichese translation of each crib,what are the most common "errors" in those strings, and what should be the penalties for imperfect matches. Thus the penalties for imperfect matches may change, and that may change the results of the matching algorithm. But in ths case the distance between f114r.4 and the second-best match is so large that it is unlikely to lose its title.
All the best, --stolfi
