(Yesterday, 06:45 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view. (16-09-2026, 09:04 AM)rikforto Wrote: You are not allowed to view links. Register or Login to view.The other thing I'm going to want to know is what the badness cutoff is---when is that badness too high to count as a match?
Since the scale of the badness score is somewhat arbitrary, there is no cutoff. I currently use the score only to rank the candidate parags.
Here is part of the report for the SBJ recipe for 粟米 = "foxtail millet". The full entry is
粟米:[味]苦,微寒。无毒。[主]养肾气,去胃脾中热,益气。
陈者:[味]苦。[主]胃热,消渴,利小便。
The punctuation is a modern addition; the You are not allowed to view links.
Register or
Login to view. (ZHB) as available in 1400 had no punctuation, so the text as the Author would get it is
粟米味苦微寒无毒主养肾气去胃脾中热益气陈者味苦主胃热消渴利小便
This entry, atypically, has no "Provenance" field. After deleting the systematic omissions (味苦微寒 and 味苦) only 25 hanzi remain:
粟米无毒主养肾气去胃脾中热益气陈者主胃热消渴利小便
Given the average of ~5 EVA letters per hanzi, the corresponding SPS entry is expected to have ~5 x 25 = ~125 letters. The program was asked to check the occurrence of four "cribs" which I am currently fairly certain about: 主, 气, 气, 主, in that order. The "canoncal" Voynichese translations should be 主 →
daiin and 气 →
chedy; but the program will accept some variations like
kaiin and
cheda,with penalties in the score.
My program considers only the 243 parags of the SPS that have a single star. The others are assumed to be two or more parags mashed together by the Scribe, and therefore not suitable for this analysis.
Here is the relevant part of the output:
243 parags read
243 were evaluated.
46 were rejected for having less than 111 letters.
126 were rejected for having more than 150 letters.
71 were examined for the requested cribs
66 had all the cribs -- sizes 114..150
64 parags had badness 99.0 or less.
To save time, my program rejected right away any SPS parag whose size was too far form that expected size, which left 71 parags. Of these, 5 were rejected because they did not have the requested cribs nor any of the currently allowed variations thereof.
<b2.6.122> (t1a1) 25
4 粟米无毒
1 2 主 养肾
1 6 气 去胃脾中热益
1 2 气 陈者
1 7 主 胃热消渴利小便
This is the trimmed SBJ entry above, split at the requested keywords. The second column is the counts of hanzis before, between, and after the keywords
Here is the best (minimum-score) match that the program found:
<f107v.20> (q0r0) 0.443 = 0.300+0.143 135(+0) 5.400 e/h
18(-3) 0.052 palaroraiinkalkalo
kaiin 0.150 12(+1) 0.034 chpalcheodyp
chedy 0.000 34(+1) 0.012 pcholkalopalkaiintaiinolkaiincholk
chedy 0.000 10(+0) 0.000 kolainaiin
kaiin 0.150 41(+3) 0.046 okalotarotalaldaralololaiinolkalcholchdar
The best-scored parag was f107v.20, with badness 0.443. The program did not find the two expected
daiin at the right places, but it found two
kaiin, which caused 2 x 0.150 = 0.300 to be added to the score; but found the two expected
chedy.
The note (q0r0) indicates that occurrences of
qo were deleted from the parag, and the strings
ir,
is, and
m were expanded to
iin, as discussed in previous posts. After these adjustments, this parag had 135 EVA letters For this parag the program assumed an average of 5.4 EVA letters per hanzi. With that ratio, the expected sizes of the gaps were 21.6, 10.8, 32.4, 10.8, and 37.8, rounded to 21..22, 10..11, 32..33, 10..11, and 37..38. The actual sizes of the gaps were 18, 12, 34, 10, and 41, giving errors of -3,+1,+1,0,+3 EVA letters, which correspond to less than 0.56 hanzi. These errors contributed 0.143 points to the score.
The second-best match was
<f116r.4> (q0r0) 1.512 = 0.350+1.162 150(+13) 5.449 e/h
17(-4) 0.097 padarsheyosheekyl
laiin 0.160 12(+1) 0.034 chckhyokaiin
chedy 0.000 34(+1) 0.012 oteedytararalarydainsheedkchdyotal
chedy 0.000 14(+3) 0.258 lkainoteedyoto
raiin 0.190 53(+14) 0.761 otylolrololysainollchedychedyoteychedy?lainotedyoteey
Note that the score is 1.512, of which 0.350 comes from using alternate "spellings"
laiin and
raiin for 主, and 1.162 comes from the errors in the gap lengths; especially the "tail" gap, that is 14 letters longer than predicted.
(This parag is also 150 letters long, which would be 6.000 EVA letters per hanzi. The program considered that value too large and used the ratio 5.449 instead when evaluating the gap errors. But even if it had used 6.000 the last gap would still be too long and would have given this parag a bad score.)
Thus I think that f107v.20
is probably the translation of the Foxtail millet entry, and there are no other plausible candidates.
Here is the graphical rendition of the match for that parag:
[
attachment=17668]
All the best, --stolfi