(07-07-2026, 03:18 PM)tavie Wrote: You are not allowed to view links. Register or Login to view.So if you're using a chatbot/LLM like ChatGPT, Claude, or Gemini for your posts here, please use it for strict translation (I presume from Italian to English) of the words you have written, and don't permit it to alter your text any more than that.
Unfortunately such a Ludditian policy is a losing battle. AI-assisted rewriting of emails, articles, forum posts, and general prose is already becoming the norm, and it isn't going away. The tools for it are built-in to most editors at this point.
The real issue isn't whether AI helped polish someone's wording. The issue is whether the post contains clear, honest, substantive reasoning. Bad AI writing should be criticized as bad writing. Empty “AI slop” should called out as empty slop. But treating any AI-assisted wording as unacceptable is an outdated standard that will be impossible to enforce and increasingly detached from how people actually starting to write things now.
Whether we like it or not, the onus is being forced back onto readers, who will have to judge an argument for itself: Is it coherent? Is it supported? Can the person defend its claims? That matters far more than whether a language tool helped clean up the prose, regardless of whether any language translating involved.
(07-07-2026, 04:03 PM)asteckley Wrote: You are not allowed to view links. Register or Login to view. (07-07-2026, 03:18 PM)tavie Wrote: You are not allowed to view links. Register or Login to view.So if you're using a chatbot/LLM like ChatGPT, Claude, or Gemini for your posts here, please use it for strict translation (I presume from Italian to English) of the words you have written, and don't permit it to alter your text any more than that.
Unfortunately such a Ludditian policy is a losing battle. AI-assisted rewriting of emails, articles, forum posts, and general prose is already becoming the norm, and it isn't going away. The tools for it are built-in to most editors at this point.
The real issue isn't whether AI helped polish someone's wording. The issue is whether the post contains clear, honest, substantive reasoning. Bad AI writing should be criticized as bad writing. Empty “AI slop” should called out as empty slop. But treating any AI-assisted wording as unacceptable is an outdated standard that will be impossible to enforce and increasingly detached from how people actually starting to write things now.
Whether we like it or not, the onus is being forced back onto readers, who will have to judge an argument for itself: Is it coherent? Is it supported? Can the person defend its claims? That matters far more than whether a language tool helped clean up the prose, regardless of whether any language translating involved.
There's nothing wrong with using AI as an auxiliary tool. But if that's the only thing you use, it's a big problem

But I don't think this research is dependent on AI, or if it is, it's not a significant factor.
(07-07-2026, 12:43 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view. (07-07-2026, 11:00 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.In my transcription of the Starred Parags section ("Quire 20"), these are the string doublets and their frequencies, ignoring spaces and line breaks but not parag breaks: [...]
I re-counted the string doublets in the Starred Parags section after "correcting" m -> iin, ir -> iin, hh -> he. The only significant changes were aiinaiin rising from (4) to (29), and daiindaiin rising from (4) to (10).
The corrected file has N = 56817 EVA letters (including '?'). The string daiin occurs 415 times, and aiin occurs 1774 times (including 415 as part of daiin). By the same reasoning as before, the string doublet aiinaiin should occur ~55 times, so the actual count (29) is a bit too low. The doublet daiindaiin should occur only ~3 times, so the actual count (10) is somewhat notable.
All the best, --stolfi
I'm sure you noticed we got 10 occurrences of
daiindaiin. Since basic probability says we should only see about 3, this clearly isn't just random noise.
Doesn't this suggest the text has internal rules forcing these strings together way more often than chance allows?
On the flip side, the exact opposite happens with
aiinaiin, which actually makes perfect sense from a linguistic perspective.
When a highly frequent string like
aiin actively avoids repeating (we see 29 but expect 55), it usually acts like a connector, article, or preposition. Think of the word 'the' in English: it's everywhere, but you never write 'the the' (unless it’s hidden inside adjacent words, which easily explains those 29 hits).
For me, this backs up the conditional variance I saw in the Python script regarding the active suppression or attraction caused by specific surrounding words.
The point raised by @dashstofks in You are not allowed to view links.
Register or
Login to view. and later @nablator in You are not allowed to view links.
Register or
Login to view. (on which word is more often duplicated in the text) is yet to be adressed.
(07-07-2026, 04:47 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.The point raised by @dashstofks in You are not allowed to view links. Register or Login to view. and later @nablator in You are not allowed to view links. Register or Login to view. (on which word is more often duplicated in the text) is yet to be adressed.
I haven't checked the whole VMS, but in the Starred Parags section chol occurs 131 times, cholchol occurs 2 times (expected ~0). As noted above, daiin occurs 415 times and daiindaiin occurs 10 times (expected ~3). Looking for them as strings, ignoring spaces and line breaks but keeping parag breaks, with the three "corrections" mentioned previously.
I think this discrepancy about chol and daiin shows the importance of analyzing each section separately. For all we know, the VMS is a compilation of several separate ur-books, each with specific topic, style, format, etc. As far as we can tell, the only thing in common among all sections is the script and (partly) the language/spelling/encryption, and the fact that they were compiled by the same Author and penned by the same Scribe.
We cannot even assume that the ur-books were composed by the same person or in the same country and century. Herbal-A may be a Basque herbal from the 12th century, while herbal-B may may be a Mingrelian one from the 10th, both translated into Voynichese and merged by the Author into a single herbal.
Anyway, the great differences in topic and format alone should cause most statistics to differ from section to section. Statistics computed over the whole VMS are likely to be more confusing than illuminating.
All the best, --stolfi
[
attachment=16397]
(07-07-2026, 05:35 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.and the fact that they were compiled by the same Author and penned by the same Scribe.
If that were actually the case, the colour of the ink and the hand might change, but not the style of writing, as there is, after all, only one author.
Some things can be seen but not calculated.
Keep your eyes off the computer screen.
(07-07-2026, 06:12 PM)Aga Tentakulus Wrote: You are not allowed to view links. Register or Login to view. (07-07-2026, 05:35 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.and the fact that they were compiled by the same Author and penned by the same Scribe.
If that were actually the case, the colour of the ink and the hand might change, but not the style of writing, as there is, after all, only one author.
Some things can be seen but not calculated.
Keep your eyes off the computer screen.
@Aga Tentakulus,
The thing is the human eye is just an analytical tool, exactly like statistics or algorithms.
One doesn't necessarily exclude the other.
In fact, I could argue that the eye can be easily fooled too, especially since we aren't looking at the real manuscript, we are looking at a digital photograph.
The pixels making up that high-resolution image are just another form of standardization, exactly like the Takahashi transcription or any other text file. They are both simply tools that digitize and standardize the original information.
If we take the 'purist' route of keeping computers out of it, we'd have to toss out almost every visual analysis out there unless the researcher is physically sitting in the Beinecke Library, next to VMS.
Algorithms don't replace human intuition, they just help us quantify the visual patterns you're pointing out, to make sure they aren't just optical illusions or scribal noise.
At the end of the day, they're all just tools.
Once you’ve calculated the contrast with the handwriting and have the result for the writing style, please let me know.
I can tell at a glance.
(07-07-2026, 06:12 PM)Aga Tentakulus Wrote: You are not allowed to view links. Register or Login to view.[images of f81v]
Do you mean that, in your opinion, those two halves of You are not allowed to view links.
Register or
Login to view. are by different Scribes?
To my eyes, it is definitely the same hand on the whole page. Between the two parags the Scribe just shifted on the chair, had a beer for lunch, switched to a new quill, whatever.
There are more dramatic abrupt changes in the appearance of the script, like between You are not allowed to view links.
Register or
Login to view. and f105r, or between lines 13 and 14 of f105r. But even there I don't see definite evidence of a change of hand.
There are differences in the shapes of some glyphs on different pages, sure; but by that criterion we should conclude that each token in the VMS was written by a different Scribe. Maybe even each glyph of each word.
And then one must consider the BEEP sorry.
Eve if we accept that there were two or more Scribes, it is still more likely than not that there was only one Author. The mere fact that the bifolios were of the same size and quality and gathered together gives that probability a significant handicap. Then there is the originality of the script, the consistent style and (poor) quality of the drawings, the uniform format through the entire Herbal section...
All the best, --stolfi
(08-07-2026, 07:34 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.There are differences in the shapes of some glyphs on different pages, sure; but by that criterion we should conclude that each token in the VMS was written by a different Scribe. Maybe even each glyph of each word.
[
attachment=16408]
Your theory seems to be correct.