(02-07-2026, 08:13 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.Thank you for your work, because if today I can run these algorithms (which were unthinkable until a few years ago) it is thanks to the foresight of your work. Without that fundamental dataset, none of this would be possible.
Thanks for the compliment, but my contributions to the interlinear transcription file were minor. Most of the merit should go to Gabriel Landini (who recently passed away) and René Zandbergen (who kept maintaining the file after Gabriel and I left the Voynich scene).
Quote:do you believe there might be edge cases in the transcription markers that my cleaning does not intercept?
If you are using Rene's current version of the transcription file, he has a very detailed description of the format at his website.
The only thing that I would add to that spec is my conviction the manuscript itself (not just the transcription file) has many errors or quirks that could hide doublets. In particular, I am convinced that
m is an abbreviation, most likely for
iin; that the ending -
ir is a scribal error for -
iin (not -
in), and -
iir is an error for -
iiin; that
Ih,
ITh,
IKh are just malformed versions of
Ch,
CTh,
CKh; and that the position of the plume on
Sh (on the first
e, on the second
e, or midway) is not significant. I would "fix" these errors before any analysis.
There must be also many bogus or omitted word spaces,
d replaced by
k, confusion between
r and
s,
a and
y,
ain and
aiin etc.; but at present I do not know how to detect and correct these other "errors".
However, be aware that the possibility of error is disputed by many, and anyway the effect of such corrections on the number of doublets would not be great.
Quote:On an empirical level, I performed a visual cross-check on the high-resolution facsimile (examining several samples, for example f108r) and the doublets counted by my parser find an exact physical correspondence in the ink written on the parchment...
I went to look for doublets on my version of the transcription file (which has many small differences from Rene's version). I did NOT do the "error fixes" above, except map Ih ITh, IKh to Ch, CTh, CKh.
I counted 292 doublets on 66 different pages. Most pages had 3 or fewer doublets. The exceptions were:
6 <f75r> Page 1 of Bio
6 <f75v> Page 2 of Bio
7 <f76r> Page 3 of Bio(?) Text only
6 <f77r> Page 5 of Bio
4 <f78r> Page 7 of Bio
5 <f80v> Page 12 of Bio
4 <f85v2> Big fold-out
7 <f103v> Starred Parags
6 <f104v> Starred Parags
6 <f107r> Starred Parags
5 <f107v> Starred Parags
7 <f108r> Starred Parags
11 <f108v> Starred Parags
7 <f111r> Starred Parags
4 <f112r> Starred Parags
7 <f116r> Starred Parags + Unknown
These counts more or less match the plot on your PDF file, but I do not see signs of the peaks at section boundaries that were so prominent in the You are not allowed to view links.
Register or
Login to view.. But that was
density of doublets per page, not
count. So perhaps there was something wrong with the denominator (number of words in the page)?
From the above counts it seems that Starred and Bio do have a higher number of doublets per page than the other sections. But they also have more words per page. On the other hand they seem rather uniform with respect to this statistic.
The higher numbers for f108r, f108v, and You are not allowed to view links.
Register or
Login to view. can be explained by the fact that, on each of those pages, the Scribe mashed 10-12 parags in a single giant parag, thus boosting the number of words on those pages.
The higher number for 116r may have something to do with the fact that the last half of that page is a block of 20 lines of dense text with no stars.
I should have counted also the number of words per page and compute the densities of doublets, not just counts. But I don't have the time now, sorry.
Doublets are common in monosyllabic languages compared to other languages. These are examples from a Chinese medical book:
太
山山谷 tài shān shān gǔ.
瘀
血血瘕欲死 yū xiě xiě jiǎ yù sǐ
劳极
洒洒 láo jí xiǎn xiǎn
伤
寒寒热 shāng hán hán rè
明
目目痛 míng mù mù tòng
洗洗酸痛 xiǎn xiǎn suān tòng
...
Some of these are just accidents - a compound term that ends in a syllable followed by another compound that start with the same syllable. Like 太
山山谷 = "high mountains and mountain valleys". Some are indeed compounds with two identical syllables. Like
洗洗 xiǎn xiǎn = "chills".
In non-monosyllabic languages, one could look instead for repeated syllables within or across words. English seems to be rather scarce on those, while they seem to be more common in Romance.
If the doublets in the VMS are mostly two-syllable compounds, like xiǎn xiǎn, then you finding would be that, within a page, certain words are positively correlated with those words, while other words are negatively correlated with them. Which would not be unusual in a natiral language.
All the best, --stolfi