The Voynich Ninja
Deconstructing the VMS : A positionless ledger (part 1) - Printable Version

+- The Voynich Ninja (https://www.voynich.ninja)
+-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html)
+--- Forum: Analysis of the text (https://www.voynich.ninja/forum-41.html)
+--- Thread: Deconstructing the VMS : A positionless ledger (part 1) (/thread-6040.html)

Pages: 1 2 3 4 5


Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

So, some of you may remember my work where I created a 'statistical' generator using what I called a ledger.  You can find it here: You are not allowed to view links. Register or Login to view.

Well, I have discovered why it didn't exactly inspire anyone. My logic was way off.  I have since been improving on my methods and digging deeper into the possibilities of the ledger and I believe I have something that works MUCH better.  Before I bury you in data I'll have to state that I have a LOT of information to impart so it will take several posts to explain all of this.  I'll start with the big news first then the explanations.

Before I go any further, I want to give credit where it is due. This work follows and expands on the work of Torsten Timm and Andreas Schinner.  Their self-citation model proposes that Voynich words were produced by copying existing words and modifying them. This is the foundation of what I have been calling “copy and mutate.”  So if you're hoping for a decoding, TLDR; it's not here.


Positionless ledger

PrefixCore midfixRare midfixSuffix
AL, R, ID, ML, S, R, M, N
OA, O, D, L, CH, SH, E, S, R, ID, L, Y, S, R, M, G
DA, O, CH, SH, ED, L, YO, L, Y, R
LA, O, D, CH, SHL, E, Y, SO, D, L, Y, S, R, G
CHA, O, D, EL, CH, SH, SA, O, D, L, E, Y, S, R, G
SHA, O, D, CH, EL, SO, D, L, E, Y, S, R
EA, O, D, ECH, Y, SO, D, E, Y, S, R, N, G
YA, D, CH, SHO, L, SS, R
SA, O, CH, SH, EDO, E, Y
RA, O, CHD, SH, E, IO, L, Y
QOE
H
MO
NDO, Y
G
XA
CS
ID, R, IL, NL, S, R, M, N

This simple ledger is capable of creating and verifying any word in the Voynich, written by any scribe, with the exception of gallows words and hapax. I have intentionally removed those two categories in order to hopefully simplify and explain how this ledger works.

The first thing you'll notice in that table is likely that I have combined CH and SH into one atomic unit.  There's also reasons for that which I'll explain later.

How to use the ledger:

Let's assume you want to create the word daiin.

  1. Look at the prefix column for the row that contains D.
  2. In the core midfix column you'll see you can choose A.  You now have D→A
  3. Now look in the prefix column for the row that contains A.
  4. Look in the core midfix column and you'll see you can choose I. You now have D→A→I
  5. Look in the prefix column for the row that contains I. 
  6. Look in the core midfix column and you'll see that you can also choose I as a midfix. You now have D→A→I→I
  7. You'd again look in the I prefix row and in the suffix column you can choose N
  8. You now have the word D→A→I→I→N

The ledger contains everything needed to reproduce:
  • 393 core word types
  • 412 rare word types
  • 805 total word types
  • 12,564 manuscript tokens

That is around 36% of the total Voynich tokens.

Prefix Stacking.

Consider these Voynich words
  • aiin
  • daiin
  • odaiin
  • chodaiin

All four appear in the manuscript.

What immediately stands out is that they share the same ending: aiin. The difference is what has been added in front of it:
  • aiin
  • d + aiin
  • od + aiin
  • chod + aiin

This is what I mean by prefix stacking. Instead of creating five completely unrelated words, the system appears to reuse the same base while building a longer structure to its left.

The longest word can be written as:

Code:
CH → O → D → A → I → I → N

The shorter words reuse parts of that same chain. This gives us transitions such as:
  • CH→O 
  • O→D
  • D→A 
  • A→I 
  • I→I 
  • I→N

If you look how the ledger works, each time you pick a new midfix, you go back to the prefix column.  It is essentially stacking prefixes in front of each new midfix you choose. 

The ledger does not need a separate rule for every complete word. The same transitions are reused as prefixes are added to familiar words.  This does not prove that the words were created in this exact order. It shows that the words form a clear family built around the same ending, with progressively larger prefixes.

And for those of you who suspect that the Voynich has Hebrew roots, I have good news: the ledger also works from right to left.

Start with a suffix. Find a prefix row that permits that suffix, then continue working backward by adding permitted midfixes until you reach the prefix where you want to stop.

For example:
  • N
  • I → N
  • I → I → N
  • A → I → I → N
  • D → A → I → I → N

You have now built daiin backward.  Every word covered by this ledger can be reconstructed in this direction. Naturally, this proves absolutely nothing about Hebrew. But the joke was sitting there, and I was not going to waste it.

Another example of prefix stacking

Consider this family:
  • chedy 500 occurrences
  • lchedy 108 occurrences
  • olchedy 34 occurrences
  • dolchedy 3 occurrences

All four words appear in the manuscript. None is a hapax or gallows word.  At each step, one new atom is placed in front of the previous word:
  • chedy
  • L + chedy = lchedy
  • O + lchedy = olchedy
  • D + olchedy = dolchedy

In ledger form:
  • CH → E → D → Y
  • L → CH → E → D → Y
  • O → L → CH → E → D → Y
  • D → O → L → CH → E → D → Y

The complete chain requires only:
  • D→O 
  • O→L 
  • L→CH 
  • CH→E 
  • E→D 
  • D→Y

Every one of those transitions appears in the core ledger.  This is a particularly clear example of prefix stacking. The word chedy remains intact while L, then O, then D are stacked in front of it.

Thats is exactly the kind of nested word family that prefix stacking would produce and the word families everyone keeps seeing in the Voynich.

Missing Letters?

You may have noticed some dashes in the ledger and some letters with no midfix or suffix and, that is correct.  H for example. It has no midfix as it never appears as a prefix.  And it has no suffix as it's never the letter before the suffix.  Now you would think that choldy would give it a midfix of o.  However, H never appears without a C or an H in front of it (keep in mind, no gallows and no hapax tokens).  But, it's included because in the full ledger, which I'll describe in the next post, that is going to change.

Which brings me to the next section:

Why I treat CH and SH as atoms

In my earlier ledgers, I treated every transcription character separately. That meant CH was represented as C followed by H, and SH was represented as S followed by H.


For example:

  • chedy = C → H → E → D → Y
  • lchedy = L → C → H → E → D → Y
  • olchedy = O → L → C → H → E → D → Y
  • dolchedy = D → O → L → C → H → E → D → Y

This created a problem in my original table. Every time another prefix was stacked onto the word, C and H were pushed one position farther into the midfix:

  • chedy H in the first midfix position
  • lchedy H in the second midfix position
  • olchedy H in the third midfix position
  • dolchedy H in the fourth midfix position

I therefore had to keep track of first midfix, second midfix, third midfix and so forth. The ledger became larger and more complicated as words became longer.  Eventually I realized that much of this apparent midfix depth was being created by the way I had divided CH and SH. I was splitting units that behaved more cleanly when kept together.


I changed the atomization so that CH and SH are each handled as one ledger atom:

  • chedy = CH → E → D → Y
  • lchedy = L → CH → E → D → Y
  • olchedy = O → L → CH → E → D → Y
  • dolchedy = D → O → L → CH → E → D → Y

Now the same word family can be described with a few reusable transitions:

  • D→O 
  • O→L 
  • L→CH 
  • CH→E
  • E→D 
  • D→Y

The ledger no longer needs to know whether CH is in the first, second, third or fourth midfix position. It only needs to know which atoms can follow CH and which atoms can lead into it.

SH works the same way.


This does not mean that C, S and H have been removed. They remain separate atoms when they occur separately. Only the exact sequences CH and SH are kept together.  “Atom” here simply means that the ledger handles the sequence as one unit. The important result is that once CH and SH are treated this way, the positional midfix columns are no longer necessary. The same information can be represented by a much smaller positionless ledger. 
 What had looked like complicated midfix depth was, in large part, the same units being pushed farther to the right by added prefixes.

Eliminating the first argument: Wouldn't any language look like this?

Yes, a positionless ledger can be constructed from any language. For comparison, I chose Finnish. Finnish is an agglutinative language: it commonly builds words by joining smaller elements together. That makes it a useful control for the kind of word-building process I am proposing for Voynichese.

I used the first 12,564 non-hapax Finnish tokens, the same number of tokens represented by the Voynich ledger above.


LedgerWord typesCore midfix transitionsRare midfix transitionsSuffix transitions
Voynich805533163
Finnish3,91927720140


So yes, Finnish can be placed into the same kind of ledger. But it requires far more word types and far more transitions to describe the same number of tokens. The ledger format itself is not unique to the Voynich Manuscript. What is unusual is how much of the Voynich vocabulary can be described by such a small set of reusable transitions.

Finnish positionless ledger: Seitsemän veljestä

PrefixCore midfixRare midfixSuffix
AA, D, E, H, I, J, K, L, M, N, P, R, S, T, U, VA, I, N, R, S, T
DA, E, I, O, ÄA, E, O, U, Ä
EA, D, E, H, I, K, L, M, N, O, P, R, S, T, U, V, Y, ÄJA, E, H, I, N, R, S, T, Ä
FRL, Ö
GA, I, OE
HA, D, E, H, I, J, K, L, M, N, O, T, U, V, Y, Ä, ÖA, E, I, O, T, U, Ä
IA, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, V, ÄA, E, H, I, N, O, S, T, Ä
JA, E, I, O, U, Y, ÄA, U, Ä
KA, E, I, K, M, O, R, S, U, Y, Ä, ÖNA, E, I, K, O, S, U, Y, Ä, Ö
LA, E, H, I, J, K, L, M, O, P, S, T, U, V, Y, Ä, ÖA, E, I, L, O, U, Y, Ä
MA, E, I, M, O, P, S, U, Y, Ä, ÖA, E, H, I, O, U, Ä
NA, E, G, H, I, K, L, N, O, P, S, T, U, Y, ÄJ, M, VA, E, I, O, Ä
OA, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, VA, H, I, N, O, S, T
PA, E, I, O, P, S, U, Y, Ä, ÖL, RA, I, O, U, Ä
RA, E, H, I, J, K, M, N, O, P, R, S, T, U, V, Y, ÄÖA, I, O, S, Y, Ä
SA, E, I, K, O, P, S, T, U, V, Y, Ä, ÖM, NA, E, I, O, S, T, U, Y, Ä
TA, E, H, I, K, O, P, R, S, T, U, Y, Ä, ÖVA, E, I, O, U, Y, Ä, Ö
UA, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, VA, E, I, N, O, S, T, U
VA, E, I, O, U, ÄY, ÖA, E, I, O, U, Ä, Ö
YD, E, H, I, K, L, M, N, P, R, S, T, V, Y, Ä, ÖI, N, S, T, Y, Ä, Ö
ÄD, E, H, I, J, K, L, M, N, P, R, S, T, V, Y, ÄE, H, I, N, R, S, T, Y, Ä
ÖH, I, J, K, M, N, R, S, T, Y, ÖD, L, P, VI, N, S, T, Ä

The hook to read the next post: Gallows are lipstick!

And, I'll also present the full ledger, explain why gallows words and hapax words are omitted from the ledger above and hopefully, give some convincing evidence as to why they were omitted.

Thanks. I have donned my flame retardant clothing, fire away.


RE: Deconstructing the VMS : A positionless ledger (part 1) - oshfdk - 24-08-2026

How is this different from a slot grammar? If treated as a slot grammar, how does this ledger score against the existing slot grammars in Mauro's metric? Most likely this question is about the full version of the ledger with gallows, etc, because the metric measures the efficiency in reconstructing the full text, which requires a nearly 100% coverage.


RE: Deconstructing the VMS : A positionless ledger (part 1) - JoJo_Jost - 24-08-2026

Hmm, but actually, all this is already well known? it’s just presented slightly differently. The very limited selection of bigrams, the high level of structure. And of course, if you leave out the ‘gallows’, you end up sticking to the basic structures of the familiar families: the ‘e’ family, the ‘aiin’ families, the ‘air’ family, and the ‘or/ar’ and ‘ol/al’ bigrams. In that respect, I’m not quite sure what to make of it yet, or perhaps I’ve missed something (which is always quite possible Wink Big Grin ).


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 07:12 AM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.How is this different from a slot grammar? If treated as a slot grammar, how does this ledger score against the existing slot grammars in Mauro's metric? Most likely this question is about the full version of the ledger with gallows, etc, because the metric measures the efficiency in reconstructing the full text, which requires a nearly 100% coverage.

Ok, couldn't sleep last nite so I dug into your question. I was vaguely aware of Mauro's work but hadn't really examined it.  So here's some preliminary results as I understand his work:

This ledger is basically a conditional slot grammar rather than a conventional slot grammar: the the next letter in a word depends on the glyph immediately before it. I had to force it into a finite slot-style implementation so that Mauro's coverage/efficiency metric could be applied. 

With gallows retained it covered 100% of the 8,424 RF1a-n types tested (the transcription Maro used) and generated about 7.82 billion possible forms, giving an efficiency of 1.08×10⁻⁶. For comparison, Mauro's EXTENDED-12 covers 88.45% while generating 6.12 billion forms, with efficiency 1.16×10⁻⁶. So the two have almost identical efficiency, but the ledger reaches full coverage. 

Stripping gallows (which I will explain my reasonings for this in the next post) makes a huge difference. The ledger still gives 100% coverage of the stripped vocabulary, but the generated space falls from 7.82 billion to about 337 million. Efficiency improves about 15-fold, to 1.65×10⁻⁵. Mauro’s EXTENDED-12 is 1.16×10⁻⁶ at 88.45% coverage.

So the stripped ledger has full coverage and roughly 14 times Mauro’s efficiency.

In my work, I use the Takahashi and the Zandbergen/Landini transcriptions.  Like Mauro's work, I remove any word that contains a splat and I did the same for the RF1a transcription just so I could compare apples to apples.

Disclaimer:  To answer your question rather quickly, ChatGPT was used to read Mauro's methodology and PDF and create a set of instructions for Codex to clean the RF1a transcription and produce the python code to run the test.  I couldn't find any code repository for Mauro.  And yes, I will be creating a repository which I'll link in the final post to all of this.


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 07:39 AM)JoJo_Jost Wrote: You are not allowed to view links. Register or Login to view.Hmm, but actually, all this is already well known? it’s just presented slightly differently. The very limited selection of bigrams, the high level of structure. And of course, if you leave out the ‘gallows’, you end up sticking to the basic structures of the familiar families: the ‘e’ family, the ‘aiin’ families, the ‘air’ family, and the ‘or/ar’ and ‘ol/al’ bigrams. In that respect, I’m not quite sure what to make of it yet, or perhaps I’ve missed something (which is always quite possible Wink Big Grin ).

The difference is that I’m not starting from those families or defining them as units. I’m deriving a small conditional transition ledger from the text itself and then asking whether that same transition system can actually reconstruct the vocabulary. So far, it reproduces the tested core, hapax and gallows populations at 100% across all five scribes. The common families fall out of the ledger because they are heavily used transition paths, rather than being built into the model beforehand. And, I believe the reason we see these 'families' is that a copy/mutate production system would create them.  Copy a word. Copy the word again, change 1 letter.  daiin / dain / dair / etc.  You just created a family.

The other part I think is potentially new is the separation between a very small common transition system and a much larger rare-transition layer. Most ordinary words use the common transitions; rarer or more unusual forms draw on the additional transitions.

I agree that the underlying regularity of Voynichese is not new. What I’m trying to determine is whether those known regularities can be reduced to a concrete, falsifiable construction rule rather than just described as word families.  

So, The whole purpose of this ledger experiment/theory is to find the smallest, easiest and most accurate construction method possible, one that was 15th century scribe doable.


RE: Deconstructing the VMS : A positionless ledger (part 1) - rikforto - 24-08-2026

(24-08-2026, 01:45 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.What I’m trying to determine is whether those known regularities can be reduced to a concrete, falsifiable construction rule rather than just described as word families.  
Sorry, I think this gets to the heart of the matter here. What is the null hypothesis you're hoping to reject? And perhaps this will be clearer from your answer to that, but how does it inform your experiment design?


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 01:51 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.Sorry, I think this gets to the heart of the matter here. What is the null hypothesis you're hoping to reject? And perhaps this will be clearer from your answer to that, but how does it inform your experiment design?

The null hypothesis is that the ledger is just another way of describing the known local patterns in Voynichese and has no real predictive power.

The test is simple: build the ledger from one body of text, freeze it, and test it on material that did not create it. In my testing, I built the ledger using only Scribe 1 and then tested how well that frozen ledger could reproduce text attributed to Scribes 2–5.  If it can reproduce unseen material to a relatively high degree, then the null starts to fail.

Foreshadowing: it does fail. It's not a spectacular failure but I think I can explain why.

I’ll cover that in detail in a later post. As I said, I’ve done a ridiculous amount of work on this, and there’s far too much to cram into one post.


RE: Deconstructing the VMS : A positionless ledger (part 1) - Rafal - 24-08-2026

Out of curiosity, why do you call it a ledger?
I am not an English native speaker but a ledger is an accounting book and this doesn't deal with accounting.

Wouldn't "word generator" or "word template table" be a better name?


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 02:40 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Out of curiosity, why do you call it a ledger?
I am not an English native speaker but a ledger is an accounting book and this doesn't deal with accounting.

Wouldn't "word generator" or "word template table" be a better name?

Interesting question.  I am old enough that I have used accounting ledgers myself, many years ago. I just started thinking of it as a ledger because it has the rows and columns a ledger or spreadsheet would have and was just easy for my brain to remember that one word.  Technically it's not in itself a generator as it can create substantially more words that the Voynich has. And because there is no 'rule' as to when to end a word (scribal preference is my best guess) it could in theory generate an infinitely long word.  As a template table, I suppose that would work but just too many syllables for my small brain.

The simplest way to think of it is as a Voynich word verifier.  You create what you think is a Voynich word and it can tell you if that word could exist.  Not does, but could.

Another way to look at it is as a transition verifier.  You can take any letter in the left column and then see if there is a midfix or suffix that matches it.  If so, then that transition exists in the Voynich.

In order to be used in a generator or a template I'd have to include letter counts (weighting), which I do have.  In that form, it could be used to create a generator and I'm still working on that.

Now, if you want the explicit details, I store it as a JSON which has no bearing whatsoever on what a 15th century scribe would use.

So for me, the easy, dumb ass word was simply... a ledger.  And, when you look at how I think it was used, it was a row/column lookup, just like you would in an accounting ledger.


RE: Deconstructing the VMS : A positionless ledger (part 1) - rikforto - 24-08-2026

(24-08-2026, 02:18 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.If it can reproduce unseen material to a relatively high degree, then the null starts to fail.
These do not seem like mutually exclusive conditions to me---and perhaps more to the point, you need to show that they are mutually exclusive---and so I don't believe you can hinge falsifiability on them. There is no inherent tradeoff in the ability of one model "to reproduce unseen material" and another to do the same. In fact, a significant problem with Voynichese studies is that many models capture at least some of the features reasonably well, so tests distinguishing them must be carefully undertaken and the tests for that thoroughly justified.

What it seems like you've tested is whether or not this model manages to reproduce key features of texts not in its sample set without reference to any other model. This is fine to its limits, but from what I can see, I agree with the point that it has mostly reproduced features from other models rather than ruling them out.