![]() |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Deconstructing the VMS : A positionless ledger (part 1) - Printable Version +- The Voynich Ninja (https://www.voynich.ninja) +-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html) +--- Forum: Analysis of the text (https://www.voynich.ninja/forum-41.html) +--- Thread: Deconstructing the VMS : A positionless ledger (part 1) (/thread-6040.html) |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
RE: Deconstructing the VMS : A positionless ledger (part 1) - Jorge_Stolfi - 01-09-2026 (01-09-2026, 01:25 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.(27-08-2026, 11:43 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.I'm still trying to understand why so many possibilities. For example: where is suffix a after ch if hapax are excluded? The text on f8r.18 is "...daiin.chor.cha.r,cheal,cha,m". I can't understand what the Scribe did there, but there is a gallows from the next line intruding into the space between "cha" and "r". It s almost certain that the intended words were "...daiin.chor.char.cheal.cham" (or "...daiin.chor.char.cheal.chaiin" if you share my belief that m is an abbrev of iin). The word ".char." occurs 77 times in my transcription, and the word pair ".chor.char." occurs twice, not counting that one. In my transcription of You are not allowed to view links. Register or Login to view. there is no "cha" as an isolated token. But my decisions at what is and is not a space differ from Takahashi's. Can you give the line number and/or quote the context? All the best, --stolfi RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 01-09-2026 (01-09-2026, 06:36 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(01-09-2026, 01:25 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.(27-08-2026, 11:43 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.I'm still trying to understand why so many possibilities. For example: where is suffix a after ch if hapax are excluded? Takahashi <f8r.P3.18;H> okar.cphaiin.chaiin.eldaiin.chor.cha.rchealcham Zandbergen/Landini retain that cha but the transcription itself shows the ugly boundary. Takahashi <f42r.P1.5;H> qokar.chockhy.chotor.cha.kary Zandbergen/Landini has it as qokar.chockhy.chotor.chy.kary. That means both Takahashi occurrences supporting CH→A as a suffix are questionable for different reasons. The You are not allowed to view links. Register or Login to view. occurrence has the spacing problem you pointed out, and the You are not allowed to view links. Register or Login to view. occurrence is read as chy rather than cha by Z/L. But, I'm no paleographer, I'm simply a number cruncher. I report on the data I find and let others with more experience tell me what I should be crunching. The ledger will never be perfect because the transcriptions don't always agree, particularly on things like spaces and ambiguous glyphs. All anyone can really provide is a best reconstruction from the available transcription. So for cases like these, I would simply treat them as transcription or segmentation uncertainty in the data. Takahashi gives two occurrences of cha, so that's what the ledger records. Looking at the individual cases tells us that CH→A probably shouldn't be regarded as a particularly well-supported transition. That's the data I have so that's the result I report. But if you step back from individual questionable cases and look at the larger transition structure, that's where the ledger starts becoming interesting. A handful of uncertain readings don't create the overall pattern. The pattern comes from the thousands of repeated transitions that aren't in dispute. RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 01-09-2026 (01-09-2026, 06:36 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.In my transcription of You are not allowed to view links. Register or Login to view. there is no "cha" as an isolated token. Do you have a link to your most current transcription? The link on Rene's site reports a 404 linking to your site and the only one it lists as available is: #=IVTFF Eva- 1.5 # # # <f0.A> {} # Last edited on 1998-12-05 11:22:03 by stolfi If that is the correct one, here's the ledger for that file with gallows retained.
And here's two json files of the full corpus both stripped and non-stripped gallows. There's a lot more data in those files including counts and it should break the midfix and suffix into core/shared/hapax/gallows categories.
Stolfi_full_stripped.zip (Size: 314.25 KB / Downloads: 0)
Stolfi_full_non_stripped.zip (Size: 444.34 KB / Downloads: 0)
RE: Deconstructing the VMS : A positionless ledger (part 1) - Jorge_Stolfi - 01-09-2026 (01-09-2026, 07:12 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Do you have a link to your most current transcription? The one I am using is You are not allowed to view links. Register or Login to view. It is not quite in Rene's IVTFF format, so you will have to process it a bit before it is accepted by Rene's tools. Mostly, it has no page header lines (you would have to copy them from Rene's), and instead of Rene's position codes "l+", "l0", etc, I use "«", "=", or "»" to indicate the position of the line's start or end relative to the corresponding rail (ideal margin line). I also use "f85v2" instead of "fRos" for the page number of the Rosettes fold-out. Actually my own transcription (code ";U") is only 80% complete; the other 20% are copied from Rene's transcription (code ";Z"). The Starred Parags section is all ";U". All the best, --stolfi RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 01-09-2026 (01-09-2026, 09:02 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(01-09-2026, 07:12 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.Do you have a link to your most current transcription? So for the type work I do, I pretty much ignore most tags. I'm only concerned with the text and it's data. So, for your transcription I did the following: "." - solid word break "," - possible word break "-" - major text break For all of those, I converted them to spaces. If that's not the intended interpretation of the breaks, let me know and I can rework it. ? - splats. While I could make some educated guesses as to what splats might be, I strip any word that contains even 1 splat. I'm interested in letter transition accuracy which shouldn't change with splat removal. All of your other tags, stripped. You also use b, j, u and v. I included them. In this ledger build I used words length 3+ and considered rare words to be count 3-. Your b only occurs as a suffix therefore, it has no midfix or suffixes that appear after it. And your u has only one occurrence as tauiin where it's a midfix so it only has a rare i as a midfix that follows it. All other occurrences are as suffix. That's why those two letter rows will look a bit empty. Also for my work, I've created a 'mappings' JSON file that I work with. It basically breaks the Voynich down into it's unique words and then records a good bit of data about each word. I've included it for you to look over. With gallows stripped
With gallows retained
Files:
mappings_Stolfi_join25e1.zip (Size: 805.55 KB / Downloads: 0)
Stolfi_join25e1_stripped_hapax.zip (Size: 1.84 MB / Downloads: 0)
I wasn't able to attach the full gallows file as even zipped it exceeded the 2mb limit. I got it down to size with other formats but the forum rejected them. Sorry. *edit* I noticed something unusual in your transcription, the C row.
With gallows intact, standalone c mostly transitions into f/k/p/t. In 98.1% of those token occurrences, the gallows is immediately followed by h, so the pattern is: c + gallows + h. When the gallows is stripped, those forms collapse to atomic ch , which is why the c row changes so dramatically. There are 217 c-initial types in the transcription, 896 tokens, and every one has a gallows. After stripping, 871 of those 896 tokens become ch-initial. Outside gallows words, I find only one minimum-length-3 word in your transcription containing standalone c: ocsesy. So the main result is in your transcription, standalone c is overwhelmingly tied to gallows constructions, especially c + gallows + h. RE: Deconstructing the VMS : A positionless ledger (part 1) - Jorge_Stolfi - 02-09-2026 (01-09-2026, 10:33 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.So for the type work I do, I pretty much ignore most tags. I'm only concerned with the text and it's data. So, for your transcription I did the following: Seems reasonable. The "-" is usually a figure intruding into the text line. Line breaks in the text are line breaks in the file. However, treating "," as a space, like ".", may not be correct. (That is why we use "," instead of ".") Unfortunately spaces are often uncertain, and each transcriber may choose to ignore them or transcribe them with "," or ".", inevitably influenced by his intuition of what is a valid token. Thus a text that was supposed to be "ched.araiin" may be transcribed as "chedar.aiin" or "cheda,raiin" or etc etc. Thus any analysis that depends on word breaks will have some amount of "noise". In your case, strings that your program considers to be suffixes could be middle parts of the intended word. Quote:? - splats. I strip any word that contains even 1 splat. It is a sensibe decision. The "?" may be an unreadable glyph, but may also be a "weirdo" glyph that occurs only a few times,and thus probably a Scribal error. Quote:You also use b, j, u and v. I included them. I used them in my transcription because the goal was to record the text as accurately as possible within the constraints of EVA . However, in my analyses I consider "g" = "m", and "u" = "n". I would also say that "j" = "d" and "eeb" = "iin", "eb" = "in", "b"= "n". I also consider the following to be errors by the Author, the Scribe, or the Transcribers: text = re.sub(r"ir", "iin", text) text = re.sub(r"is", "iin", text) text = re.sub(r"m", "iin", text) text = re.sub(r"hh", "he", text) text = re.sub(r"ih", "ch", text) text = re.sub(r"i([ktpf])h", r"c\1h", text) Quote: I noticed something unusual in your transcription, the C row. ... So the main result is in your transcription, standalone c is overwhelmingly tied to gallows constructions, especially c + gallows + h. I did not follow the details of this comment, but yes: in my word model, there is no standalone c or C. There are only Ch, CTh, CKh, CPh, CFh. Thus when I see something that looks like a c outside of these cases, I assume that it is a distorted e and transcribe it as "e". There are rare cases of CHh, CTHh, CKHh, CPHh, CFHh; I transcribe them as such, but I think those are actually Che, CThe, CKhe, CPhe, CFhe . All the best, --stolfi RE: Deconstructing the VMS : A positionless ledger (part 1) - Dunsel - 02-09-2026 (02-09-2026, 12:00 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Thus any analysis that depends on word breaks will have some amount of "noise". In your case, strings that your program considers to be suffixes could be middle parts of the intended word. Exactly. It would be most tedious to go through each transcription and try to judge what the transcriber meant and you'd always be guaranteed to get something wrong. (02-09-2026, 12:00 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(01-09-2026, 10:33 PM)Dunsel Wrote: You are not allowed to view links. Register or Login to view.? - splats. I strip any word that contains even 1 splat.It is a sensibe decision. The "?" may be an unreadable glyph, but may also be a "weirdo" glyph that occurs only a few times,and thus probably a Scribal error. For the moment, I think it's reasonable. If it is a weirdo, it's not going to be in most 'normal' transcriptions. None that I know of transcribe the two "V" wierdos on f1r. And, as in the case of the ASCII characters in the ZL transcription, that would be for a later analysis. I'm just trying to get the big picture for the moment. Eventually, with the ledger, weighting, some OCR and spell checking algorithms, I think it would be possible to put some scoring system on repaired splats (those that have most of their letters). (02-09-2026, 12:00 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.I did not follow the details of this comment, but yes: in my word model, there is no standalone c or C. There are only Ch, CTh, CKh, CPh, CFh. Thus when I see something that looks like a c outside of these cases, I assume that it is a distorted e and transcribe it as "e". There are rare cases of CHh, CTHh, CKHh, CPHh, CFHh; I transcribe them as such, but I think those are actually Che, CThe, CKhe, CPhe, CFhe . And that explains it perfectly. From my side, it didn’t look like the other transcriptions I’ve loaded into the ledger, so I was trying to understand why the "c" row looked so different. Running your transcription through the ledger-building code recovered your "c" rule and put it directly into the ledger. Now that you’ve explained the transcription convention, the "c" row is exactly what I would expect to see from the rule you described above. So, I would say those are pretty accurate ledger renditions of your transcription. |