Dunsel > Yesterday, 04:04 PM
(Yesterday, 03:13 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.These do not seem like mutually exclusive conditions to me---and perhaps more to the point, you need to show that they are mutually exclusive---and so I don't believe you can hinge falsifiability on them. There is no inherent tradeoff in the ability of one model "to reproduce unseen material" and another to do the same. In fact, a significant problem with Voynichese studies is that many models capture at least some of the features reasonably well, so tests distinguishing them must be carefully undertaken and the tests for that thoroughly justified.
What it seems like you've tested is whether or not this model manages to reproduce key features of texts not in its sample set without reference to any other model. This is fine to its limits, but from what I can see, I agree with the point that it has mostly reproduced features from other models rather than ruling them out.
Jorge_Stolfi > Yesterday, 05:26 PM
(Yesterday, 04:04 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Does the ledger predict material it was not built from? Yes. A ledger built from scribe 1 core material reproduces a high proportion of core text from Scribes 2–5.
Quote:Did the scribes actually use something like this to construct words?
rikforto > Yesterday, 05:29 PM
oshfdk > Yesterday, 06:28 PM
Dunsel > Yesterday, 06:54 PM
(Yesterday, 06:28 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.(Yesterday, 01:31 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.So the stripped ledger has full coverage and roughly 14 times Mauro’s efficiency.
As far as I remember, Mauro's metric is measured in bits. That's what I was interested in.
| Grammar | Coverage | Possible words generated | Efficiency | F1 |
|---|---|---|---|---|
| Zattera SLOT | 41.26% | 4.66 million | 7.10×10⁻⁴ | 1.43×10⁻³ |
| BASIC-13 | 50.58% | 4.56 million | 8.90×10⁻⁴ | 1.78×10⁻³ |
| BASIC-11 | 48.70% | 2.28 million | 1.72×10⁻³ | 3.40×10⁻³ |
| COMPACT-7 | 47.86% | 670,000 | 5.75×10⁻³ | 1.13×10⁻² |
| EXTENDED-12 | 88.45% | 6.12 billion | 1.16×10⁻⁶ | 2.33×10⁻⁶ |
| Finite conditional ledger | 100.00% | 7.82 billion | 1.08×10⁻⁶ | 2.15×10⁻⁶ |
Dunsel > Yesterday, 08:20 PM
(Yesterday, 05:26 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(Yesterday, 04:04 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Does the ledger predict material it was not built from? Yes. A ledger built from scribe 1 core material reproduces a high proportion of core text from Scribes 2–5.
I didn't look at your "ledger" model in detail, but from what I read it seems that it is what mathematicians call a finite state automaton (FSA), whose language (in the mathematical sense, meaning the set of generated words) is the Voynichese lexicon. Specifically, it is an acyclic automaton, meaning that it has no closed loops, and therefore it generates only a finite set of words.
For any finite set of words there is an FSA that generates all those words and no others. That automaton can be constructed by a procedure that seems to be similar to your ledger construction method.
In particular, any "slot grammar" or any other model that generates a finite set of words has an equivalent FSA that generates the same words. The opposite may not be true; that is, there may be no slot grammar that generates the same words as a given FSA. That is why slot grammars are always imperfect: they either generate only a subset of the VMS lexicon or generate many words that are not in the lexicon. On the other hand, for that same reason a slot grammar is more informative: if the VMS lexicon can be closely approximated by a slot grammar (say, with only 5% of errors either way), that is a very significant fact about the language or encoding.
Quote:Did the scribes actually use something like this to construct words?
Even before I found the SPS=SBJ match, I had no respect for all the "gibberish" theories, for common sense reasons. The fact that the lexicon is described by a slot grammar or some other mathematical model does not mean that the text was created using that model, and does not rule out a meaningful text. Not even a natural language text in the plain.
All the best, --stolfi

oshfdk > Yesterday, 09:11 PM
(Yesterday, 06:54 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.As a positionless conditional-ledger it scores 479,627 Nbits with gallows retained and 416,094 with gallows stripped. Mauro's published best is 480,985 Nbits, or 481,947 after his coverage correction.
Dunsel > Yesterday, 09:47 PM
Dunsel > Yesterday, 10:16 PM
(Yesterday, 09:11 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.How exactly was the ledger converted to a chunk dictionary to compute Nbits?
Mauro.txt (Size: 3.29 KB / Downloads: 2)
Jorge_Stolfi > 10 hours ago
(Yesterday, 08:20 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.The ledger also does not generate only the attested lexicon. Its transitions can be recombined into unattested words. So it is not simply an automaton that memorizes the finite set of words it was built from. ... I am also not sure how you concluded that it is acyclic after saying you had not looked at it in detail. In the ledger, whenever a glyph that can begin a word appears inside a word, the positional count begins again from that glyph.
Quote:The ledger was not built from the whole Voynich lexicon and then tested against the same lexicon. Of course that would prove nothing. It was built from Scribe 1 core material, frozen, and then tested on core material from Scribes 2–5.
Quote:It accepted 89.3% of [scribes 2-5] tokens and 77.6% of their types.