The Voynich Ninja
[split] Arguments for the text being generated - Printable Version

+- The Voynich Ninja (https://www.voynich.ninja)
+-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html)
+--- Forum: Theories & Solutions (https://www.voynich.ninja/forum-58.html)
+--- Thread: [split] Arguments for the text being generated (/thread-6160.html)



[split] Arguments for the text being generated - DG97EEB - 03-10-2026

I'm 95% convinced it's generated, but that might only be the top layer.. I have a feeling that there's plain text underneath that was encoded and that was then normalised through a generative layer, and all the tests that are run on the stats are only reflective of that layer which is why no one can quite get all the way to the bottom.. but given that doesn't reflect a single working practice I can find in 15th Century Germany, it may also be wrong...


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - JoJo_Jost - 03-10-2026

I'm 72.4 percent Wink Big Grin sure that they're tables—probably four of them (one for vowels, two for consonants and consonant clusters, and one for miscellaneous)—with around 200 table cells...but it may also be wrong, too...


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - DG97EEB - 03-10-2026

Deleted


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - Pierre Dumont Himself - 03-10-2026

You know you can use a real map, right? Milano should be just slightly to the east of Pavia, and Bellinzona to the west.


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - DG97EEB - 03-10-2026

(03-10-2026, 05:01 PM)Pierre Dumont Himself Wrote: You are not allowed to view links. Register or Login to view.You know you can use a real map, right? Milano should be just slightly to the east of Pavia, and Bellinzona to the west.

Yeah, thought it looked wrong...will delete


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - hatoncat - 03-10-2026

(03-10-2026, 07:29 AM)DG97EEB Wrote: You are not allowed to view links. Register or Login to view.I'm 95% convinced it's generated, but that might only be the top layer.. I have a feeling that there's plain text underneath that was encoded and that was then normalised through a generative layer, and all the tests that are run on the stats are only reflective of that layer which is why no one can quite get all the way to the bottom.. but given that doesn't reflect a single working practice I can find in 15th Century Germany, it may also be wrong...


Can you clarify what you mean by "95% convinced it's generated"? What is "it"?


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - DG97EEB - 03-10-2026

(03-10-2026, 07:10 PM)hatoncat Wrote: You are not allowed to view links. Register or Login to view.
(03-10-2026, 07:29 AM)DG97EEB Wrote: You are not allowed to view links. Register or Login to view.I'm 95% convinced it's generated, but that might only be the top layer.. I have a feeling that there's plain text underneath that was encoded and that was then normalised through a generative layer, and all the tests that are run on the stats are only reflective of that layer which is why no one can quite get all the way to the bottom.. but given that doesn't reflect a single working practice I can find in 15th Century Germany, it may also be wrong...


Can you clarify what you mean by "95% convinced it's generated"? What is "it"?
This is GPT verbatim but it summarises my research, and of course very liberally borrows from everyone before... standing on the shoulders of giants. If and when I publish this formally, everyone will be acknowledged.. remember, this is a forum post not a published paper....most of this is known and I've simply joined all the dots.

"Observed Voynich Grammar

By “grammar” here, I mean repeatable statistical rules that govern how Voynich tokens are formed, where they appear, and what tends to follow what. This does not assume that the manuscript encodes a natural language.

Important limits and retractions

- A fixed folio or page “codebook” does not work.
Proof: when we removed the explicit page-family codebook from the frozen generator, mean joint error improved from 3.061 to 1.140, a 62.8% reduction.
Corollary: folios do have local vocabulary structure, but it is not well described as selecting words from a fixed page-specific list.

- Third-order within-token context is not a secure core rule.
Proof: the extra predictive value of third-order context changes sign or drops below our resolution threshold under different smoothing settings. One representative result was only 1.45 SD.
Corollary: the safe conclusion is that Voynich token formation needs at least second-order context. We should not insist on third-order context.

- Exact identity of the previous word is not completely irrelevant, but it is much weaker than we once thought.
Proof: most transcription systems show little remaining exact-token information after morphology and local structure are included, although one transcription retains a 2.72 SD residual.
Corollary: the shape and class of the previous token matter reliably. Its exact identity may sometimes add a small extra effect.

- A literal single wheel is not supported.
Proof: the strongest proposed opener cycle is not unique, and literal wheel models fail held-out data.
Corollary: the transitions between line openings are real. The physical mechanism behind them is still unknown.

1. Voynich words have a strong internal grammar

- Voynich words are not arbitrary strings of glyphs.
Proof: held-out prediction requires at least second-order glyph context.
Corollary: what glyph can come next depends on more than the immediately previous glyph.

- Some glyph sequences are effectively forbidden.
Proof: we found 19 graphotactic zero constraints that replicate across six EVA-family transliterations and survive an alternative atomic representation.
Corollary: only a restricted subset of possible glyph sequences is actually allowed.

- Some apparently dramatic glyph rules turn out to depend on transcription conventions.
Proof: several very large EVA effects shrink when composite glyphs are decomposed. One drops from about 32 SD to 2.9 SD.
Corollary: the internal grammar is real, but we have to distinguish between genuine writing-system rules and artefacts of the EVA transcription scheme.

- "q" changes what can follow it.
Proof: after controlling for the exact continuation, "q" context still predicts the k/t choice. The primary effect is 4.77 SD, and it is positive in 4 of 5 alternative transcriptions.
Corollary: "q" is not simply an independent prefix. It participates in the internal grammar of the word.

- The length of an "i" run helps determine how the run ends.
Proof: this survives short-context controls in all six transcription layers. The weakest result is still 3.52 SD.
Corollary: repeated "i" strokes form a structured construction rather than unconstrained repetition.

2. Words belong to structured families

- Voynich words form dense edit-distance families.
Proof: words that differ by only one to three glyph edits occur in systematic neighbourhoods rather than as randomly scattered strings.
Corollary: the vocabulary is organised into families of related forms.

- One-edit relationships are especially important.
Proof: ED1 neighbours show the strongest local structural concentration.
Corollary: changing a single glyph is one of the main ways the system creates related forms.

- These families are productive.
Proof: several transformations generalise to forms that were not part of the original fitted pairs.
Corollary: the system is not just memorising a list of words. It appears to reuse form-building rules.

- Members of the same family are not interchangeable.
Proof: exact surface form still predicts context even after family membership is supplied to the model.
Corollary: related forms are genuinely related, but they do not all do the same job.

3. Some recurring changes behave like operators

- Pre-final "d" is a reproducible transformation.
Proof: pairs roughly of the form "Xy ↔ Xdy" behave differently in matched contexts.
Corollary: adding or removing "d" is associated with a systematic change in how the form behaves.

- Other one-edit changes also behave systematically.
Proof: recurring changes involving k/t, a/o, d/e, and gallows forms show repeatable effects across families.
Corollary: the vocabulary is better understood as a network of transformations than as a flat list of unrelated words.

- Different operators move together.
Proof: across 36 tested operators, mean absolute correlation is 0.559, compared with a null of 0.174. The difference is 38.9 SD.
Corollary: these transformations are coordinated rather than independent.

- But they do not collapse neatly into one hidden master state.
Proof: attempts to explain the operator covariance with one portable latent dimension fail transfer tests.
Corollary: the system seems to have several interacting grammatical dimensions.

4. Currier A and B use much of the same machinery

- Currier A and B share many of the same formal families.
Proof: collapsing eight formal families reduces A/B divergence from JSD 0.4606 to 0.3246, a 29.5% reduction.
Corollary: a large part of the Currier difference comes from choosing different variants from related families.

- Gallows behaviour explains more of the difference.
Proof: adding gallows structure reduces JSD further to 0.3062, a 33.5% total reduction from baseline.
Corollary: A and B appear to use partly shared grammar with different settings or probabilities.

- Currier cannot yet be cleanly separated from section and scribal hand.
Proof: those factors are heavily confounded in the surviving manuscript.
Corollary: Currier is a real distributional distinction, but we cannot yet say with confidence whether it represents dialect, author, cipher setting, or something else.

5. Section also changes the grammar

- Different sections favour different variants from the same formal families.
Proof: section information improves prediction even after global token frequency and family structure are included.
Corollary: Herbal, Biological, Pharmaceutical, Stars and other sections do not simply use different vocabularies. They alter the probabilities within a shared system.

- Some operator behaviour transfers between sections.
Proof: held-out transfer into Biological and Stars material is roughly 2.2 to 3 SD, depending on the test.
Corollary: part of the formal grammar is manuscript-wide.

- The transfer is not universal.
Proof: several sections do not resolve the same effects.
Corollary: the manuscript seems to use a shared grammar that is locally adjusted by section.

6. Position inside a physical line matters

- Line beginnings, middles and endings behave differently.
Proof: 33 of 35 previously identified positional effects survive the main control battery.
Corollary: physical line position is part of the grammar.

- Some positional slots remain even after simplifying the vocabulary.
Proof: P1, P3 and E−2 type effects survive family and morphology controls.
Corollary: line-position effects are not just caused by a handful of common words.

- Line openings and line endings favour different forms.
Proof: opener and terminal glyph distributions differ systematically from medial positions.
Corollary: the physical line has its own boundary rules.

7. Adjacent words are linked

- The end of one word helps predict the beginning of the next.
Proof: previous-token edge and morphology features improve held-out prediction, with a primary residual of about 3.84 SD.
Corollary: neighbouring Voynich words are not sampled independently.

- Useful information includes word shape as well as identity.
Proof: first glyph, ending structure and token length retain predictive value after frequency controls.
Corollary: the next word depends partly on the form of the previous word.

- This effect is mainly immediate.
Proof: lag 1 is much stronger than lags 2 to 4.
Corollary: this is mostly a local word-to-word rule.

8. A physical line break changes the system

- A normal space and a line break are not equivalent.
Proof: within-line edge dependence is stronger across a SPACE than across a physical LINE BREAK. This survives matched-volume testing in the five full-coverage transcription layers.
Corollary: Voynich lines are genuine structural units, not arbitrary wrapping of continuous text.

- The strongest version is bounded rather than universal.
Proof: one smaller transcription does not resolve the SPACE versus LINE BREAK contrast.
Corollary: the line-reset effect is strong in the main representations, but we should not claim perfect universality.

- The ED1 drop at line boundaries is not proven to be a separate avoidance rule.
Proof: preserving token frequencies and broad vertical geometry can reproduce much of it.
Corollary: the ED1 boundary effect is likely part of the wider line grammar rather than a standalone instruction to avoid similar words across lines.

9. Repetition has its own structure

- Exact repetition is not distributed randomly.
Proof: recurrence at distance two is enriched, while immediate exact repetition is not clearly suppressed under matched controls.
Corollary: repetition follows a structured distance profile.

- Recent reuse on the same page improves prediction.
Proof: a banded lag-2-to-64 recurrence channel improves the target statistic by 2.23 SD.
Corollary: recently used material remains locally available for reuse.

- Distance matters.
Proof: the banded recurrence model beats a uniform recurrence model by 2.22 SD.
Corollary: reuse probability changes with distance.

- Whether lag 1 should be excluded remains unresolved.
Proof: the exclusion test gives only 1.55 SD and succeeds in 2 of 5 folds.
Corollary: local memory is real, but its exact lower boundary is not settled.

10. Folios have local vocabulary structure, but not a fixed codebook

- Folios preferentially reuse particular word families and neighbourhoods.
Proof: distinct-word and family locality survives held-out tests at roughly 6 to 13 SD, depending on transcription and fold.
Corollary: local repertoire is a genuine feature of the manuscript.

- This is not simple copying from nearby text.
Proof: folio-order shuffles and explicit copy-and-modify models fail while locality remains.
Corollary: local reuse is structural rather than just a copying habit.

- A fixed folio palette also fails.
Proof: palette-style generators produce too much exact repetition while still missing the real pattern of distinct-word reuse.
Corollary: local vocabulary availability must change dynamically or depend on other hidden conditions.

11. Paragraph position is a major factor

- The first line of a paragraph is special.
Proof: paragraph-line position adds about 0.323 bits per opener, around 8.9 SD above the matched null. About 70% of that signal comes from line 1.
Corollary: paragraph beginnings represent a distinct grammatical state.

- Paragraph boundaries partly reset the system.
Proof: opener-transition behaviour changes strongly at paragraph starts.
Corollary: the manuscript has structure above the level of the individual line.

12. There is a separate vertical grammar of line openings

- The opening form of one line helps predict the opening form of the next line.
Proof: this survives controls for section, Currier, Davis hand and exact paragraph-line position.
Corollary: successive physical lines are linked through a specific left-margin process.

- This remains after accounting for the strong paragraph-position effect.
Proof: residual gain is about +0.0163 bits per opener, or 3.48 SD.
Corollary: paragraph position alone cannot explain the line-to-line opener pattern.

- Repeating the same opener on consecutive lines is strongly disfavoured.
Proof: identical opener transitions occur well below conditioned expectation.
Corollary: the next opener is selected partly in relation to the previous opener.

- The effect is mainly first-order.
Proof: the immediately previous opener matters, while corresponding lag-2 directional contrasts do not resolve.
Corollary: the relevant state is short-lived.

13. Some opener transitions are directional

- A "d → q → y → d" pattern is reproducible.
Proof: those transition asymmetries survive conditioning in ordinary paragraph text.
Corollary: opener changes contain directional information, not just frequency differences.

- But "d/q/y" is not the only important cycle.
Proof: in the full candidate search it ranks only 13th of 165 comparable structures.
Corollary: the vertical grammar is richer than a single three-state cycle.

- Finer analysis finds several specific transitions.
Proof: "q→ch" survives in both Currier A and B, "y→o" survives in A, and "d→q" survives in B. The apparent A "o→q" effect disappears after conditioning.
Corollary: Currier A and B share parts of the opener grammar but weight them differently.

14. The opener does not control the whole line

- Previous opener predicts the next opener better than the previous line ending does.
Proof: after conditioning on opener, the ending-to-next-opener effect does not resolve.
Corollary: this is specifically an opener-to-opener process, not just ordinary text continuing across the line break.

- The opener has only a very short influence on the rest of its own line.
Proof: there may be a weak token-2 effect, but by token 3 and beyond the body-level difference disappears.
Corollary: there is no evidence for one hidden state controlling an entire line.

- Lines with different major openers quickly become statistically similar downstream.
Proof: after removing the first token, "q", "d", "y" and "o" opener classes do not show a resolved whole-line body difference.
Corollary: the vertical opener process and the horizontal word-generation process are partly separate.

15. Simple wheel models fail

- The Currier-B "y→o→q→y" cycle is not enough to generate the text.
Proof: the held-out miss is 3.90 SD, rising to 5.24 SD under the whole-quire bound.
Corollary: seeing a local cycle does not mean that cycle is the generative mechanism.

- Allowing several transition components still does not solve it.
Proof: multi-component versions continue to miss the relevant dependence by around 5 SD in the main test.
Corollary: the evidence does not currently justify a “multiple wheels” explanation either.

- What survives is the transition rule, not the hardware.
Corollary: the scribe could have produced it mentally, procedurally, mechanically, through a table, or as a side effect of some deeper encoding. The statistics do not choose between these possibilities.

16. Exact previous-word identity adds surprisingly little once the main grammar is known

- Most next-token predictability comes from structural features rather than remembering the exact previous word.
Proof: in the selector model, adding exact previous-token identity after family, section, recency and edge compatibility improves prediction by only about 0.0093 bits per token.
Corollary: the system seems to operate mainly on classes and features rather than on ordinary lexical bigrams.

- A small exact residual may still exist.
Proof: one alternate transcription resolves an exact-identity effect at 2.72 SD.
Corollary: saying that exact word identity never matters would be too strong.

17. Simple generators still fail the manuscript as a whole

- The best compact joint model does not reproduce Voynich well enough.
Proof: held-out discrepancy is +0.0618 bits per token, with null SD 0.00772, or about 8.0 SD.
Corollary: reproducing several individual rules is not the same as reproducing the manuscript.

- The failure is not confined to one metric.
Proof: some boundary and vocabulary diagnostics miss by roughly 35 SD.
Corollary: the missing mechanism is not just one parameter that needs tuning.

- Large synthetic generators show the same problem.
Proof: they can produce text that looks superficially Voynich-like while failing the strongest held-out structural tests.
Corollary: visual resemblance, word-frequency resemblance, or lots of near-neighbour words are weak tests of a real solution.

Proof statement

The claim that Voynich has a reproducible surface grammar does not depend on one striking statistic.

It comes from a collection of independent findings:

1. Words obey strong internal sequencing rules.
2. Words form productive families of related forms.
3. Some recurring changes behave like grammatical operators.
4. Currier A/B and manuscript section change the probabilities within that system.
5. Position inside a physical line matters.
6. Adjacent words influence one another.
7. A physical line break weakens that local dependence.
8. Repetition follows a short-range distance structure.
9. Paragraph beginnings create a distinct state.
10. Successive line openings influence one another.
11. Some opener transitions are directional.
12. Simple copying, fixed codebooks, Markov models and wheel models fail to reproduce all of these features together.

These findings have been tested using held-out prediction, matched randomisations, alternative transcription systems, shuffled controls, representation changes, ablation tests and explicit generative models.

The simple null model, in which Voynich words are sampled more or less independently from a frequency distribution, is therefore untenable.

That establishes a surface grammar.

It does not establish what the text means, what language may lie underneath it, whether it is a cipher, or what historical mechanism produced it.

Corollary

Any serious explanation of the Voynich text has to reproduce all of these things at the same time:

restricted word formation; productive families of related forms; recurring operator-like transformations; Currier and section effects; line-position rules; local word-to-word dependence; a physical line reset; short-range reuse; paragraph structure; and a separate short-memory system governing successive line openings.

A model that reproduces only Voynich-looking words, edit-distance families, Currier differences, line effects, or opener cycles is not enough.

The simplest architecture consistent with the evidence so far is:

a global form grammar
plus locally conditioned choice among related forms
plus section and Currier effects
plus within-line word-to-word and recurrence rules
plus physical line-position constraints
plus paragraph state
plus a mostly separate vertical line-opener process.

That is the grammar we can currently demonstrate. The mechanism underneath it remains open."


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - Jorge_Stolfi - 04-10-2026

(03-10-2026, 07:46 PM)DG97EEB Wrote: You are not allowed to view links. Register or Login to view.This is GPT verbatim but it summarises my research

Thanks for posting this summary!  But let me summarize my objections too:

Quote:Corollary: folios do have local vocabulary structure, but it is not well described as selecting words from a fixed page-specific list.
Not sure what this means, but we expect that a herbal (which is ~1/3 of all parags text in the VMS) would have a more or less fixed core vocabulary V0 that occurs uniformly over all pages (apart from sampling error), plus maybe some vocabularies T1, T2,.. (like "tea" or "poultice")  that depends on the type of plant and occurs only in  certain subset of the pages, plus some vocabulary  P1, P2,... (like the plant name) that occurs on only in one page.  Except that the vocabularies V0 and Tk may be different for Herbal A and B. 

Whereas a materia medica (which is at last a strong possibility for the Starred Parags section) should not have well-defined Tk, and the vocabularies Pk (the names of the remedies) should be specific for each parag.

Isn't this the case?

Quote:Voynich words have a strong internal grammar. They are not arbitrary strings of glyphs. Some glyph sequences are effectively forbidden.
No kidding... Dodgy

Quote:Some apparently dramatic glyph rules turn out to depend on transcription conventions.
Who would have guesses this... Dodgy 

Quote:The extra predictive value of third-order context changes sign or drops below our resolution threshold ... Voynich token formation needs at least second-order context.

Word structure models seem to imply this result.  But they also imply that much longer contexts are important.  For instance, to check or enforce the rule "at most one gallows per word" one needs one bit of information that depends on all preceding glyphs, possibly 6 or more.

Quote:The shape and class of the previous token matter reliably. Its exact identity may sometimes add a small extra effect. Adjacent words are linked ... Corollary: neighbouring Voynich words are not sampled independently. the next word depends partly on the form of the previous word. Most next-token predictability comes from structural features rather than remembering the exact previous word. lag 1 is much stronger than lags 2 to 4.
Doing the analysis this way may be misleading.  It may well be that it is in fact the exact identity of the token that matters, but the class and shape alone give good predictions because all statistics will be dominated by the most frequently occurring tokens. 

For example, in English narrative or discursive prose the most common words by far are "the", "and", and "of".  If one builds a next-token predictor based on the first two letters of the current word, one should get quite good results; and knowing the rest of the word should not add much. But that is because the first two letters already identify the whole word with significant accuracy. 

Could this be happening in the tests discussed above?

Quote:The end of one word helps predict the beginning of the next.
You mean, like a word ending with "we" in English is  followed by the word "are" more often than expected by chance?

Quote:"q" is not simply an independent prefix. It participates in the internal grammar of the word.
The first part is true.  The second part could be stated more informatively as "if q (or qo) is assumed to be a prefix, it can be attached only to certain words, not to any word at random"

Quote:The length of an "i" run helps determine how the run ends.
Yes.  But more, 
  • the statistics of "ir" are more similar to those of "iin" than to that of "in",
  • the statistics of "iir" are more similar to those of "iiin" than to those of "iin"
  • the statistics of "iiir" are more similar to those of "iiiin" than to those of "iiin"
the latter is just saying that "iiir" and "iiiin" are both essentially absent.
And this pattern can be explained by the "r" in these endings being a misreading of "in" by the Scribe, considering that the shapes can be quite similar in hurried handwriting.

Quote:Voynich words form dense edit-distance families. One-edit relationships are especially important. changing a single glyph is one of the main ways the system creates related forms. These families are productive. Corollary: the system is not just memorising a list of words. It appears to reuse form-building rules.
Again, the observation is correct but the corollary is an unwarranted conclusion.  Those results can be explained by variations in the spelling of words, which can be due to several causes including inherent spelling flexibility (like one could have written "been" as "bin" or "bene" in 1600's English, in the same text), scribal errors, fading ink (that could easily turn Sh into Ch, Ch into ee, y into a or o, etc.)

Quote:pairs roughly of the form "Xy ↔ Xdy" behave differently in matched contexts.  Corollary: adding or removing "d" is associated with a systematic change in how the form behaves.
How strange.  Dodgy Like adding or removing final "-s" in English?  Or adding and removing final "-g" in Mandarin?

Quote:Corollary: the vocabulary is better understood as a network of transformations than as a flat list of unrelated words... Different operators move together.... But they do not collapse neatly into one hidden master state...  Corollary: the system seems to have several interacting grammatical dimensions.
Can't these observations be explained as the result of spelling variations, as proposed above?

Quote:Currier [A/B languages] is a real distributional distinction, but we cannot yet say with confidence whether it represents dialect, author, cipher setting, or something else.
That is: we have not made any progress on this question in the last 50 years. Sad

Quote:Section also changes the grammar... Herbal, Biological, Pharmaceutical, Stars and other sections do not simply use different vocabularies. They alter the probabilities within a shared system. ... the manuscript seems to use a shared grammar that is locally adjusted by section.
So there is a non-trivial grammar common to the whole VMS, but sections with vastly different topics and text structure use some distinctive phrases and sentence structures?  Who would have guessed...  Dodgy

(Seriously, isn't this a strong argument against "gibberish"?)

Quote:Position inside a physical line matters  - Line beginnings, middles and endings behave differently. Line openings and line endings favour different forms.  Corollary the physical line has its own boundary rules.
The differences are real, but the corollary is misleading as stated.  The encryption camp seems to have  taken it as meaning that the encryption method is somehow reset or modified around line breaks (the LAAFU theory).  But recently it was pointed out that the "anomalous" statistics around line breaks can have several other explanations, including 
  • the fact that line breaks in any running text have a clear tendency to occur before longer words;
  • the high likelihood that m is an abbreviation for iin (and maybe other endings) that the scribe could use when space was tight;
  • the tendency for word spaces to be omitted in the final part of a line;
AFAIK it has not been shown that these effects cannot suffice to explain the observed anomalies around line breaks. 

Quote:there is no evidence for one hidden state controlling an entire line.
Which kinda goes against the LAAFU theory, no?

Quote:Recent reuse [of a word] on the same page improves prediction. local repertoire is a genuine feature of the manuscript. reuse probability changes with distance.
Again, what a surprise... Dodgy

Quote:Paragraph position is a major factor  The first line of a paragraph is special.  paragraph beginnings represent a distinct grammatical state. the manuscript has structure above the level of the individual line.
And this too?  Dodgy

Quote:There is a separate vertical grammar of line openings ... the vertical grammar is richer than a single three-state cycle.
This claim should be more detailed.  What are the transition probabilities? Do they depends on distance from parag start? How do they change with section?  Do they extend across parag  breaks in the Stars section? How accurate is the prediction?

Quote:A model that reproduces only Voynich-looking words, edit-distance families, Currier differences, line effects, or opener cycles is not enough.The simplest architecture consistent with the evidence so far is:
  • a global form grammar
  • plus locally conditioned choice among related forms
  • plus section and Currier effects
  • plus within-line word-to-word and recurrence rules
  • plus physical line-position constraints
  • plus paragraph state
  • plus a mostly separate vertical line-opener process.

As you know, I have a simpler "architecture" that seems to explain all these effects.  But that goes in another thread... Cool

All the best, --stolfi


RE: AI Cracks 217-Year-Old Napoleonic Cipher in Six Hours, Revealing Pre-War Troop Deploy - DG97EEB - 04-10-2026

I have a lot more Jorge... And I can cover all your points...  Will try to post something non GPT later..