The Voynich Ninja
Deconstructing the VMS : A positionless ledger (part 1) - Printable Version

+- The Voynich Ninja (https://www.voynich.ninja)
+-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html)
+--- Forum: Analysis of the text (https://www.voynich.ninja/forum-41.html)
+--- Thread: Deconstructing the VMS : A positionless ledger (part 1) (/thread-6040.html)

Pages: 1 2 3 4 5


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 03:13 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.These do not seem like mutually exclusive conditions to me---and perhaps more to the point, you need to show that they are mutually exclusive---and so I don't believe you can hinge falsifiability on them. There is no inherent tradeoff in the ability of one model "to reproduce unseen material" and another to do the same. In fact, a significant problem with Voynichese studies is that many models capture at least some of the features reasonably well, so tests distinguishing them must be carefully undertaken and the tests for that thoroughly justified.

What it seems like you've tested is whether or not this model manages to reproduce key features of texts not in its sample set without reference to any other model. This is fine to its limits, but from what I can see, I agree with the point that it has mostly reproduced features from other models rather than ruling them out.

There are really four basic claims that I could make:

  1. Does the ledger predict material it was not built from? Yes. A ledger built from scribe 1 core material reproduces a high proportion of core text from Scribes 2–5.  I'll provide that result in a later post.
  2. Does the organization of the ledger capture something beyond bigram statistics? I'm not sure yet. When I preserve the actual scribe 1 bigrams and randomize them, most of the predictive power remains. So I can't show that and won't make that claim.
  3. Did the scribes actually use something like this to construct words? The copy and mutate results are consistent with that idea, but they don't prove it. So, I won't make that claim and I think it would be difficult to prove it if I did. What I can say is that, if a ledger like this were being used, the procedure itself would be mechanically quite simple: copy a word, change one atom according to the available transitions in the ledger, and continue.
  4. Does the ledger represent a physical device? Other than a basic sheet of paper (or vellum), I have no evidence for that either. I have considered that some wheel or grille type device may have been used and there's some very crude evidence but I've yet to construct anything solid.  So, I feel safe in claiming that it could have been constructed and used as a single sheet of paper. Anything beyond that, I won't make any claim.

So, I don't think the ledger itself needs to “beat” every other model simply to be a valid. At the moment, I can show that it it has strong similarity to the scribes it was not built from, and I can show that the system itself is compact enough to fit on a single sheet of paper (vellum). If I were to claim anything from 2 or 3 above or if I claimed anything more than a piece of paper, then yea, I'd definitely need some solid proof and falsification.  I'm not saying that 2-4 are excluded from possibilities.  I just haven't gone down that road yet.


RE: Deconstructing the VMS : A positionless ledger (part 1) - Jorge_Stolfi - 24-08-2026

(24-08-2026, 04:04 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Does the ledger predict material it was not built from? Yes. A ledger built from scribe 1 core material reproduces a high proportion of core text from Scribes 2–5. 

I didn't look at your "ledger" model in detail, but from what I read it seems that it is what mathematicians call a finite state automaton (FSA), whose language (in the mathematical sense, meaning the set of generated words) is the Voynichese lexicon.  Specifically, it is an acyclic automaton, meaning that it has no closed loops, and therefore it generates only a finite set of words. 

For any finite set of words there is an FSA that generates all those words and no others.  That automaton can be constructed by a procedure that seems to be similar to your ledger construction method.

In particular, any "slot grammar" or any other model that generates a finite set of words has an equivalent FSA that generates the same words.  The opposite may not be true; that is, there may be no slot grammar that generates the same words as a given FSA.  That is why slot grammars are always imperfect: they either generate only a subset of the VMS lexicon or generate many words that are not in the lexicon.  On the other hand, for that same reason a slot grammar is more informative: if the VMS lexicon can be closely approximated by a slot grammar (say, with only 5% of errors either way), that is a very significant fact about the language or encoding.

Quote:Did the scribes actually use something like this to construct words?

Even before I found the SPS=SBJ match, I had no respect for all the "gibberish" theories, for common sense reasons.  The fact that the lexicon is described by a slot grammar or some other mathematical model does not mean that the text was created using that model, and does not rule out a meaningful text.  Not even a natural language text in the plain.

All the best, --stolfi


RE: Deconstructing the VMS : A positionless ledger (part 1) - rikforto - 24-08-2026

You're very focused on the model, saying things like, "I don't think the ledger itself needs to 'beat' every other model simply to be a valid." But a model isn't valid or invalid, instead claims about it are. And the claim that caused me to bring up the other models was, "The null hypothesis is that the ledger is just another way of describing the known local patterns in Voynichese and has no real predictive power." This is an obvious comparative to other models, which are the other ways of describing known local patterns in Voynichese. To distinguish this model and reject the null hypothesis you lay out here, it does in fact have to beat other models in some well-defined, mutually exclusive sense. And notice that this is not a condition that I imposed, this is something you took on by how you answered a particular question I asked about falsifiability claims.

What I would suggest for future iterations of this line of research is that you begin with a well-defined question and a model and analysis tailored to answer it. You might re-read Timm and Schinner for how they structure their argument because you are working in that tradition and the paper is quite well-argued whether one is ultimately persuaded to autocitation or not. It begins by laying out a number of claims researchers have made based on statistical features of the VMS and how those imply (but do not conclusively prove) meaningfulness. It then lays out a process for generating text that is both plausible for the technology of the time and is meaningless, trying to make as few assumptions as possible. It then shows that a meaningless text generated that way has the features that imply meaningfulness while being meaningless. It concludes by arguing that this counterexample provides a plausible alternative to the hypothesis that the text is meaningful based on those statistics. While the model is obviously crucial, not to mention interesting, notice it does not logically begin with the model, but instead begins with a set of claims made by proponents of the meaningfulness hypothesis and the model parameters fall out of trying to find a way to test those claims.

So my questions are these: What question is your model built to answer? Why is that a question worth asking? Why should we be persuaded it tests that? Then, is it is comparative to other models? (You have said both "yes" and "no" on this thread!) If it is comparative, the same questions repeat; what comparison is the model built to make; why is that a comparison worth checking; why should we be persuaded by those comparisons? This will focus your analysis and how you present it, and make it easier for people to see what you're contributing here


RE: Deconstructing the VMS : A positionless ledger (part 1) - oshfdk - 24-08-2026

(24-08-2026, 01:31 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.So the stripped ledger has full coverage and roughly 14 times Mauro’s efficiency.

As far as I remember, Mauro's metric is measured in bits. That's what I was interested in.


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 06:28 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.
(24-08-2026, 01:31 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.So the stripped ledger has full coverage and roughly 14 times Mauro’s efficiency.

As far as I remember, Mauro's metric is measured in bits. That's what I was interested in.

As a positionless conditional-ledger it scores 479,627 Nbits with gallows retained and 416,094 with gallows stripped. Mauro's published best is 480,985 Nbits, or 481,947 after his coverage correction.

So, with gallows, real close.  Without, a considerable difference.

Other comparisons from Mauro's work. This is with gallows.

GrammarCoveragePossible words generatedEfficiencyF1
Zattera SLOT41.26%4.66 million7.10×10⁻⁴1.43×10⁻³
BASIC-1350.58%4.56 million8.90×10⁻⁴1.78×10⁻³
BASIC-1148.70%2.28 million1.72×10⁻³3.40×10⁻³
COMPACT-747.86%670,0005.75×10⁻³1.13×10⁻²
EXTENDED-1288.45%6.12 billion1.16×10⁻⁶2.33×10⁻⁶
Finite conditional ledger100.00%7.82 billion1.08×10⁻⁶2.15×10⁻⁶



RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 05:26 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.
(24-08-2026, 04:04 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Does the ledger predict material it was not built from? Yes. A ledger built from scribe 1 core material reproduces a high proportion of core text from Scribes 2–5. 

I didn't look at your "ledger" model in detail, but from what I read it seems that it is what mathematicians call a finite state automaton (FSA), whose language (in the mathematical sense, meaning the set of generated words) is the Voynichese lexicon.  Specifically, it is an acyclic automaton, meaning that it has no closed loops, and therefore it generates only a finite set of words. 

For any finite set of words there is an FSA that generates all those words and no others.  That automaton can be constructed by a procedure that seems to be similar to your ledger construction method.

In particular, any "slot grammar" or any other model that generates a finite set of words has an equivalent FSA that generates the same words.  The opposite may not be true; that is, there may be no slot grammar that generates the same words as a given FSA.  That is why slot grammars are always imperfect: they either generate only a subset of the VMS lexicon or generate many words that are not in the lexicon.  On the other hand, for that same reason a slot grammar is more informative: if the VMS lexicon can be closely approximated by a slot grammar (say, with only 5% of errors either way), that is a very significant fact about the language or encoding.

Quote:Did the scribes actually use something like this to construct words?

Even before I found the SPS=SBJ match, I had no respect for all the "gibberish" theories, for common sense reasons.  The fact that the lexicon is described by a slot grammar or some other mathematical model does not mean that the text was created using that model, and does not rule out a meaningful text.  Not even a natural language text in the plain.

All the best, --stolfi

I think you are answering a different test from the one I described.  The ledger was not built from the whole Voynich lexicon and then tested against the same lexicon. Of course that would prove nothing. It was built from Scribe 1 core material, frozen, and then tested on core material from Scribes 2–5. It accepted 89.3% of their tokens and 77.6% of their types. The point is that those four scribes were not used to build it.  The ledger also does not generate only the attested lexicon. Its transitions can be recombined into unattested words. So it is not simply an automaton that memorizes the finite set of words it was built from.

I am also not sure how you concluded that it is acyclic after saying you had not looked at it in detail. In the ledger, whenever a glyph that can begin a word appears inside a word, the positional count begins again from that glyph. For example, suppose D and Y can both begin words, and DY and YD are both permitted first-position transitions. DYDYDY can then be read as D→Y, Y→D, D→Y, with each pair occupying the same first position. The walk repeatedly returns to the same D and Y states. That is a loop, not a one-way path through a fixed series of slots.  And yes, that word exists and yes, the ledger validates it.

I agree that none of this proves that the scribes actually used a ledger, that the text is meaningless, or that it could not encode natural language. I did not claim any of those things. I would love to discover the Voynich has content. I hope it does. And I even briefly looked at the possibility of using the ledger as an encryption method.  

Consider this: There are 77 core transitions, 107 rare-only transitions, and 137 suffix transitions in just the scribe 1 ledger.  That's more than enough to contain content.

I am NOT saying this is how the Voynich works.  And this is a VERY...  for those who missed that... VERY crude example. But, it does show that encoding content with the ledger is possible and with some work, it could be made to look like Voynich.

Consider that transitions can be used to represent Latin characters.  And assume that certain bigrams are "anchors".  In this example: ch | sh | qo | yd | ol are those anchors.

T → char
H → shar
E → she

And I encode that as charsharshe

And then a new word

S → ydai 
Y → ydar 
S → ydai 
T → char 
E → she 
M → yda

And I encode this: ydaiydarydaicharsheyda

So next, I create this:

W → old 
O → ydar 
R → char 
K → olar 
S → ydai

When I'm done I have this huge string

charsharsheqoydaiydarydaicharsheydaqooldydarcharolarydai

Now I can add spaces anywhere

char shar she qo ydai ydar ydai char sheyda qo oldydar char olar ydai

And I have a simple substitution encryption that kinda looks like Voynich.

And, it can be decoded back into : the system works.

And I can improve it further by adding in some high entropy and low entropy transitions to obscure things a bit and create this

chail shoey sheo qo ydar yda ydar chala sheod dyqo olod ydar chala olod yda

Which is still the same encoded string and still decodable.

And, if I want to get very modern, I can assign transitions a binary bit. Depending on which transition I use, I can encode an 8 bit binary number which can then be decoded back to an ascii character.

01011001 01100101 01110011 = Yes

So can the Voynich still contain content even if the ledger is used to create it?  Oh yes, very much so.  That method is similar to what Michael Greshko presented with his Naibbe cypher.  I believe he used multiple unigrams and bigrams to encode single latin letters. This ledger can use the same method, just transitions instead of bigrams.

What I'm hoping to demonstrate is simply a viable method by which it MIGHT be produced and some of the choices that MIGHT have been made if it was constructed with this method.  Do I believe that copy and mutate was how the Voynich was created?  Yes.  Does that exclude content?  No.  But if any content is in the Voynich, someone has a LOT of explaining to do.

The ledger only describes constraints on the visible Voynichese.  Any language or encryption decoding is going to have to follow those constraints.

Thanks for making me think about this reply. Big Grin

Edit:  I had to have some help with this but yes, I would say that it is an FSA as described.  However, it's not probabilistic, like yours, nor is it acyclic.  And here's where I had to have help.  This would be the formula that it uses... I think.  GPT came up with this so take it for what it's worth.

   

Now, I supposed if I used the weighting data I have stored with it and used it to create a generator then, I believe you could call it probabilistic.  I'm not quite there yet.



RE: Deconstructing the VMS : A positionless ledger (part 1) - oshfdk - 24-08-2026

(24-08-2026, 06:54 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.As a positionless conditional-ledger it scores 479,627 Nbits with gallows retained and 416,094 with gallows stripped. Mauro's published best is 480,985 Nbits, or 481,947 after his coverage correction.

How exactly was the ledger converted to a chunk dictionary to compute Nbits?


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

With the questions I'm getting I've realized that perhaps I didn't explain the ledger well enough. Let's see if this helps

Think of every non-suffix letter in the Voynich as being a parent.  Each parent can have children. Suffixes are the exception, they are only children.

So when I look at the word daiin.  D is the prefix and it immediately becomes a parent.  A is a child of D.  But then A becomes a parent of I so I is a child of A. But, I then becomes a parent and it's child is also named I.  That I then becomes a parent of the suffix N. 

When I look at the word dshody, again, D is the prefix and becomes a parent of the child SH. Because I looked at daiin first, A and SH are now siblings as both have D as a parent.  SH becomes a parent and has the child O.  O becomes a parent and has D as a child. I then end the word with the D suffix child, Y.

Now for those into genealogy you notice that in dshody that the same D that started that word also becomes a child of O.  And in daiin, the I becomes it's own parent and it's own child. I did not say the Voynich was without it's own special case of inbreeding.

So, when building the ledger using every word and this parent, child, sibling relationship, every word will be attested.  The ledger will always be capable of creating 100% of Voynich words or any other language fed into it.  So it's not exactly a slot model per-se.

Using the ledger, here's the complete ancestry tree for the letter D using just scribe 1 herbal.  I hope this helps explain how it works.

   


RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 24-08-2026

(24-08-2026, 09:11 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.How exactly was the ledger converted to a chunk dictionary to compute Nbits?

Good question.  As I mentioned I had to enlist codex to create the python as I'm not that familiar with his work and this particular math is a bit beyond me. But I wanted to try to give you an answer.  So, here's what it says:

"It's a lossless conditional Huffman codec directly from the cleaned RF1a-n text, using the ledger’s general idea that the available next glyph depends on the current glyph. It would be more accurate to call this a ledger-style conditional transition codec."

Here's the code it used to create it.  It would appear you have a better understanding of it that I do so if you see anything odd, let me know and I'll try to fix it.


.txt   Mauro.txt (Size: 3.29 KB / Downloads: 3)


RE: Deconstructing the VMS : A positionless ledger (part 1) - Jorge_Stolfi - 25-08-2026

(24-08-2026, 08:20 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.The ledger also does not generate only the attested lexicon. Its transitions can be recombined into unattested words. So it is not simply an automaton that memorizes the finite set of words it was built from. ... I am also not sure how you concluded that it is acyclic after saying you had not looked at it in detail. In the ledger, whenever a glyph that can begin a word appears inside a word, the positional count begins again from that glyph.

Apologies for misunderstanding the nature of your ledger.  

But if the ledger has cycles, it accepts an infinite set of word types, no?  So you have a separate limit on the length of the words?  Or you generate text without spaces, and insert spaces by a separate method?

Does your ledger accept words with two gallows, or with an EVA [dlrs] between two benches {ch,sh}?

Quote:The ledger was not built from the whole Voynich lexicon and then tested against the same lexicon. Of course that would prove nothing. It was built from Scribe 1 core material, frozen, and then tested on core material from Scribes 2–5.

Ok.  But I don't think there was ever a question of whether the hypothetical five scribes used different vocabularies.  The claim was that there were only two "languages", and Herbal was divided among them.  

It would make more sense to investigate vocabulary sharing among the sections, treating Herbal-A and Herbal-B as two sections.  It is expected that there will be differences in vocabulary due to different subject matters.

Quote:It accepted 89.3% of [scribes 2-5] tokens and 77.6% of their types.

Those numbers seem rather low.  But low numbers are expected because of the small size of scribe-1 sample and the variety of topics.