(11-07-2026, 04:41 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Not really. It is enough that the mapping from plaintext to Voynichese maps each word type to the the same word type or word types, in a sufficiently large fraction of the cases. Like, for example, 70% of the daiin come from the plaintext word "herb", 70% of the Chedy come from "serpent", etc. That would be enough to make two pages with the same topic have a higher-than-expected number of shared words.
But there is another problem with the idea of using word usage similarity to order the pages. It may work for the Bio section, where a topic like "digestive system" may span several pages. It is not expected to work on Herbal and Recipes, because each paragraph is likely to be an independent text, and we do not expect parags in adjacent pages to use more common words than parags in two distant pages.
Thus, if one orders the pages by that criterion, one would get a seemingly continuous sequence, where word usage changes gradually along the sequence and is increasingly different as the pages gert further apart. However, that gradient would be spurious, created by the ordering process itself.
As an analogy, suppose that one orders the books on a bookshelf by size. The result will be a nice smooth gradient of sizes. But of course that does not mean anything: we cannot infer that the books were printed in that order...
All the best, --stolfi
This is correct: ordering the pages by word usage similarity is not necessarily going to give you the original page ordering. Consider a novel like
The Fifth Season or
A Game of Thrones, where narrators and settings vary chapter by chapter. Grouping by word usage probably would have a better chance of grouping all chapters with the same narrator together than reconstructing the author's intended chapter order.
But in principle, an ordering based on word usage similarity could be treated as an input for other analytical techniques. For example, it's worth seeing how page- or folio-level trends in handwriting—perhaps average glyph sizes, average glyph proportions, or the presence/absence of specific glyph features—change across a given ordering independently of word usage. Maybe you could start to evaluate hypotheses on the compositional sequence of the bifolia.
(12-07-2026, 07:05 PM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.It is so simple. Confoliate and conjoint correlation compliments this scenario nicely and the conclusion of this investigation should not surprise any hoax hypothesist.
Indeed, the bifolium-level production pattern — the scribe writing sheet by sheet, with inside facing pages most similar because they were visible together — is expected for the self-citation method. Every time the author starts a new (empty) page it is useful to use another page as source. There are indications that the scribe preferred the last completed sheet for this purpose. One peculiarity for the words used seven or eight times is that they often appear on subsequent pages (see Timm 2014, p. 16).
The objection that the scribe would need linguistic sophistication and a complicated method doesn't hold. Mary D'Imperio of the NSA described the mechanism in 1978: "The scribe, faced with the task of thinking up a large number of dummy sequences, would naturally tend to repeat parts of neighboring strings with various small changes and additions." Gaskell and Bowern (2022) confirmed this experimentally: when they asked volunteers to produce meaningless text, many independently adopted exactly this approach.
And also a recent YouTube experiment (You are not allowed to view links.
Register or
Login to view.) demonstrates the same point from a different angle. Nearly 200 volunteers were asked to produce pure gibberish — either by keyboard smashing or by inventing fake words. Both strategies produced Zipfian frequency distributions. The presenter's conclusion: "passing Zipf's law doesn't prove it's a real language. It just proves a human wrote it." The volunteers weren't trying to mimic language — their cognitive constraints (repetition bias, mental fatigue) result in the statistical patterns we see in the Voynich text.
No linguistic knowledge or deliberate engineering is needed. Self-citation is the natural human response to the task of filling pages with text-like content.
(13-07-2026, 09:47 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view. (12-07-2026, 07:05 PM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.It is so simple. Confoliate and conjoint correlation compliments this scenario nicely and the conclusion of this investigation should not surprise any hoax hypothesist.
Gaskell and Bowern (2022) confirmed this experimentally: when they asked volunteers to produce meaningless text, many independently adopted exactly this approach.
Gaskell and Bowern (2022) also noted explicitly that their 42 writing samples were short. They acknowledged that their data is insufficient to prove whether the "higher-level structure" and long-range patterns seen across the 200+ pages of the Voynich Manuscript could be sustained purely by intuitive gibberish:
"We recruited 42 volunteers to write short “gibberish” documents and statistically compared them to several transcriptions of the VMS and a large corpus of linguistically meaningful texts...
However, our writing samples are too short to test whether the higher-level structure of VMS pages and quires could also be produced by gibberish"
You are not allowed to view links.
Register or
Login to view.
Morover, the Voynich text adheres to highly reststrictive rules regarding which characters can be placed next to each other, so while we humans naturally tend to repeat words to fill a page, sustaining such strict unnatural structural rules over hundreds of pages implies a deliberate, conscious generation system or algorithm, rather than just an intuitive, passive reflex.
Quote:Indeed, the bifolium-level production pattern — the scribe writing sheet by sheet, with inside facing pages most similar because they were visible together — is expected for the self-citation method.
I just would like to add that some bifolio patterns are really striking.
Here is the usage of
m (m) in some of first pages, coming from
voynichese.com website:
You are not allowed to view links.
Register or
Login to view.
[
attachment=16583]
Can you see these clusters? They are from
3r, 3v, 6r, 6v pages. And you know what? They all make a single bifolio
How would you explain it if you claim that Voynichese is a natural text?
(15-07-2026, 12:10 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.How would you explain it if you claim that Voynichese is a natural text?
One of the more common conjectures among people who think there is a meaningful text is that
m is a positional variant of a more common character. Why it should be more available here and why a scribe might have preferred it for a single bifolio are open questions, sure, but nothing there strikes me as out of line with long-held views that
m is likely a positional variant
(13-07-2026, 09:47 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.Indeed, the bifolium-level production pattern — the scribe writing sheet by sheet, with inside facing pages most similar because they were visible together — is expected for the self-citation method.
I am not sure that I understand this point correctly. When writing a bifolium, also the outside pages would be visible together.
(15-07-2026, 12:10 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Can you see these clusters? They are from 3r, 3v, 6r, 6v pages.
Also 8r, 8v, 17r, 17v, 22r-24v... This pattern of abundance of
m on consecutive pages but NOT on facing pages contradicts Torsten's "with inside facing pages most similar because they were visible together".
[
attachment=16592]
Interesting!
Here's my chart showing bifolia, quires, scribes, and sections. This will be published in our forthcoming book, but since that's still a year away from being published, I'm happy to share the diagram here. Those of you who are interested in following up on our findings may find it useful.
(15-07-2026, 02:09 PM)nablator Wrote: You are not allowed to view links. Register or Login to view. (15-07-2026, 12:10 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Can you see these clusters? They are from 3r, 3v, 6r, 6v pages.
Also 8r, 8v, 17r, 17v, 22r-24v... This pattern of abundance of m on consecutive pages but NOT on facing pages contradicts Torsten's "with inside facing pages most similar because they were visible together".
I have a couple of resources that I created for my own research and, since they are relevant to the singulion ideas discussed by Lisa and Colin in their paper, I have packaged them up for distribution in case anyone else might find them useful.
They include:
- Singulion Reconstructed Images of the Voynich Manuscript (Beinecke MS 408)
- Voynich Manuscript Transliterations in Singulion (Unbound-Sheet) Order
Both are described in more detail and can be obtained here:
You are not allowed to view links.
Register or
Login to view.
(I recently reconfigured my servers and website applications, so I apologize if you hit any downloading issues. You can PM me if you experience a problem.)
(15-07-2026, 12:10 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Can you see these clusters? They are from 3r, 3v, 6r, 6v pages. And you know what? They all make a single bifolio 
How would you explain it if you claim that Voynichese is a natural text?
For instance, a change of the encryption 'key' could generate these differences (notice I'm agnostic on the VMS being meanigless or meaningful, I'm just considering the possibilities).