Hi everyone,
First of all, I am an amateur when it comes to both the Voynich Manuscript and historical linguistics. English is also not my first language, so apologies in advance if I misuse some terminology.
I am
not proposing a decipherment, and I am not claiming that Voynichese represents any particular kind of language.
I would like to ask whether the following kind of controlled experiment has already been attempted, and whether there is an obvious methodological problem with it that I am overlooking.
The basic question is:
Quote:If a Herbal plant and a Pharmaceutical plant fragment are matched independently from the illustrations, does the associated text preserve a statistically detectable relationship between the two?
More specifically, instead of asking whether exactly the same complete Voynich "word" occurs in both places, I am wondering whether one could compare recurring
sub-word glyph structures — for example bigrams or trigrams — while giving relatively little weight to very common sequences and more weight to comparatively rare ones.
Why I thought this might be worth testing
As I understand it, one unusual feature of Voynichese is its relatively low conditional entropy. Within a token, the preceding glyphs strongly restrict which glyphs or glyph sequences are likely to follow.
There also appears to be considerable positional structure within Voynich "words". Glyphs and short recurring glyph sequences have preferred positions: some occur disproportionately near the beginning of a token, others near the middle or end, and many tokens appear to be constructed from a relatively limited repertoire of recurring components.
I am
not assuming that these components are morphemes or that this proves Voynichese represents a synthetic natural language.
The same surface structure might potentially arise from several very different mechanisms:
- genuine morphology
- phonological or syllabic restrictions
- a cipher or another encoding system
- a system of abbreviations
- an artificial procedure generating word-like forms from a limited set of components
For the experiment I have in mind, the actual meaning of these components would not matter. The only requirement would be that recurring sub-token structures can be identified statistically.
Herbal–Pharmaceutical correspondences
What made me think about this is the limited number of proposed correspondences between plants in the Herbal section and plant fragments in the Pharmaceutical section.
Some of these proposed matches seem considerably stronger to me than simply saying that two drawings happen to look vaguely similar.
One particularly interesting case is You are not allowed to view links.
Register or
Login to view. and You are not allowed to view links.
Register or
Login to view. in relation to f102r2.
The two Herbal plants occur on the same physical bifolium, while the plant fragments proposed to correspond to them appear next to each other again on f102r2.
There are several other proposed Herbal–Pharma correspondences, particularly in the f102 material.
This also creates an important methodological problem. Several proposed Pharmaceutical matches occur on the same page or panel, so they cannot simply be treated as independent pairs of complete texts.
For that reason I was thinking of carrying out the experiment on two separate levels.
1. Individual Pharmaceutical labels
First, the strongest Herbal–Pharma correspondences would have to be selected using
only the illustrations, before looking at textual similarity.
Where a particular Pharmaceutical fragment has a sufficiently localized label that can reasonably be associated with that object, that label could then be compared with the text on the corresponding Herbal page.
Instead of searching only for an identical complete word, the label could be decomposed into shorter recurring glyph sequences — for example sequences of two or three glyphs.
By this I mean sequences of actual transcribed glyphs, rather than arbitrary groups of two or three ASCII characters produced as an accidental consequence of the EVA transcription system.
For each such label, one could then ask:
- How strongly are its characteristic short glyph sequences represented on the visually corresponding Herbal page?
- How frequently do the same sequences occur on other comparable Herbal pages?
A sequence that is extremely common throughout Voynichese would tell us relatively little. Potentially more informative would be comparatively rare sequences that occur unusually often on the Herbal page selected
in advance from the illustration.
So the test would not simply be:
Quote:Does the exact same word occur in both places?
It would be closer to:
Quote:Does the visually corresponding Herbal page contain an unusually high concentration of the same relatively rare sub-word elements that occur in this Pharmaceutical label?
This would provide an object-level test wherever sufficiently well-localized Pharmaceutical labels exist.
2. The f102 Pharmaceutical material as a larger textual unit
The second level would use the larger textual context rather than individual labels.
Several of the most interesting proposed Herbal–Pharma correspondences occur together in the f102 Pharmaceutical material.
Since several matching plant fragments may share the same Pharmaceutical page or panel, it may be more useful to compare the complete text of a relevant f102 panel — or eventually the f102 material as a whole — with the group of Herbal pages whose plants are visually represented there.
For example, if one part of f102 contains plant fragments apparently corresponding to several particular Herbal pages, the combined textual profile of those Herbal pages could be compared with the text of the relevant f102 panel.
The question would then be:
Quote:Is the text in this part of f102 statistically more similar to the group of Herbal pages visually associated with its plant fragments than it is to other comparable Herbal pages or groups of pages?
The same principle could eventually be applied at a broader level to the f102 Pharmaceutical material as a whole.
In simplified form:
Code:
textual profile of the relevant f102 material
vs.
combined textual profile of the Herbal pages
selected independently from the illustrations
vs.
comparable alternative groups of Herbal pages
Again, I would expect extremely common Voynichese components to carry relatively little information, while comparatively rare and informative glyph sequences would be more useful.
Why use both levels?
The two tests seem to me to answer slightly different questions.
The label-level test asks whether a particular illustrated object has a detectable textual fingerprint.
The f102-level test asks whether a whole group of visually associated Herbal entries leaves a detectable statistical fingerprint in the Pharmaceutical material where those objects reappear.
Several of the correspondences that interest me also appear to belong to Currier A and to the same proposed scribal hand. I would hope that restricting comparisons appropriately might reduce the risk of simply rediscovering already known differences between Currier A and B, or between different scribes.
Suitable control pages would presumably also need to be matched as closely as possible for these factors, as well as for text length and section.
What would a positive result actually mean?
I would
not interpret such a result as a decipherment, and I would not claim that a particular sequence means "root", "plant", or anything else.
The narrower question is simply:
Quote:Does the text contain statistical information about relationships between objects that can be identified independently from the illustrations?
Suppose that visually corresponding Herbal and Pharmaceutical objects consistently shared more rare sub-word components than unrelated combinations, and that the same effect could also be detected when comparing the f102 Pharmaceutical material with the group of Herbal pages represented there.
I can think of at least several possible explanations.
- Meaningful text.
If both entries concern the same plant, some lexical, morphological, or otherwise object-specific component might survive in both contexts even if the rest of the text is completely different.
- Artificial or self-referential pseudo-text.
Some form of artificially generated text would still be possible. However, the generation procedure would apparently have to be more complicated than simply producing locally similar words or drawing from a general "plant vocabulary".
It would somehow have to preserve object-specific relationships across different parts of the manuscript.
- A common source or exemplar.
Both the illustrations and some of the associated textual elements might have been copied from a common source.
In that case, the person producing the Voynich Manuscript would not necessarily even have to understand the repeated textual elements. It would be sufficient to reproduce information that was consistently associated with the same object in the exemplar.
In all of these cases, I think a positive result would still be interesting.
At minimum, it would suggest that the textual system contains information statistically associated with particular recurring visual objects, even if we remained completely unable to read a single word.
What would a negative result mean?
Probably much less.
Failure to find such an effect would obviously not demonstrate that Voynichese is meaningless.
A plant name might not occur in the Pharmaceutical section at all. Different abbreviations might be used. The Herbal and Pharmaceutical texts might discuss completely different properties of the same object. A Pharmaceutical label might refer to something other than the nearby plant fragment. While the absence of such a correspondence does reduce the likelihood of the text being meaningful, as we know, absence of evidence is not evidence of absence.
So I would interpret a negative result only as:
Quote:This particular method failed to detect an object-specific textual relationship.
My questions
Has anyone already carried out this particular kind of controlled comparison?
I have found discussions comparing identical labels or complete words, as well as broader studies of vocabulary similarity between the Herbal and Pharmaceutical sections. However, so far
I have not found an analysis that combines all of the following:
- Selects Herbal–Pharma correspondences in advance using only the illustrations, without looking at the text when deciding which objects match.
- Tests object-level relationships by comparing individual Pharmaceutical labels with rare or informative sub-word structures on the corresponding Herbal pages, rather than only searching for identical complete words.
- Separately tests the f102 material at the page, panel, or section level against the group of Herbal pages visually represented there.
- Compares the observed similarities against suitable alternative Herbal pages or groups, rather than simply identifying interesting matches after the fact.
- Takes into account the fact that several corresponding Pharmaceutical fragments may occur on the same page or panel and therefore should not necessarily be treated as statistically independent observations.
If something like this has already been done, I would be very grateful for a reference.
If it has not, I would be equally interested in hearing whether there is some methodological problem that makes either of these tests less informative than they appear to me.
In particular, I would appreciate criticism concerning the choice of sub-word features, suitable control groups, statistical independence, or anything else I may be overlooking.
As mentioned above, I am very much an amateur here, so corrections are very welcome.