quimqu > 01-06-2026, 12:55 PM
| Folio | Section | Observed behaviour |
|---|---|---|
| f86v | Text-only | strongest transversal hub |
| fRos | Cosmological | cross-section connector |
| f111r | Marginal stars | dense lexical hub |
| f113r | Marginal stars | dense lexical hub |
| f115v | Marginal stars | highly connected |
| f108v | Marginal stars | bridge-like behavior |
| f72v | Zodiac | unexpectedly connected |
| f67r | Astronomical | distant lexical links |
| f89r | Pharmaceutical | cross-section overlap |
| f76v | Biological | emerges at relaxed thresholds |
quimqu > 02-06-2026, 11:01 AM
dashstofsk > 02-06-2026, 11:03 AM
(01-06-2026, 12:55 PM)quimqu Wrote: You are not allowed to view links. Register or Login to view.rare and semi-rare tokens
quimqu > 02-06-2026, 11:13 AM
(02-06-2026, 11:03 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.(01-06-2026, 12:55 PM)quimqu Wrote: You are not allowed to view links. Register or Login to view.rare and semi-rare tokens
I am interested in this. But how do you define rare and semi-rare? Is it words that occur only 2,3,4... times in the manuscript? What is the bar for a word to no longer be rare?
dashstofsk > 02-06-2026, 11:54 AM
quimqu > 02-06-2026, 01:07 PM
nablator > 02-06-2026, 01:18 PM
(01-06-2026, 12:55 PM)quimqu Wrote: You are not allowed to view links. Register or Login to view.So I started wondering if the opposite approach might actually be more informative.
Quote:In addition to creating the A matrix as described, which uses straight TF values, two weighting schemes are also employed to modify the values contained in A. The two schemes applied are Term Frequency-Inverse Document Frequency (TF-IDF) and Log-Entropy (LE).You are not allowed to view links. Register or Login to view.
dashstofsk > 02-06-2026, 01:24 PM
Dunsel > 02-06-2026, 02:02 PM
(01-06-2026, 12:55 PM)quimqu Wrote: You are not allowed to view links. Register or Login to view.I think this opens a new way of checking for relationships between folios and sections.
| Scribe | Scribe-local hapax |
|---|---|
| 1 | 1338 |
| 2 | 1261 |
| 3 | 525 |
| 4 | 298 |
| 5 | 210 |
| Pair | Matches | Jaccard Overlap |
|---|---|---|
| 1–2 | 105 | 4.21% |
| 1–3 | 55 | 3.04% |
| 1–4 | 37 | 2.31% |
| 1–5 | 22 | 1.44% |
| 2–3 | 80 | 4.69% |
| 2–4 | 29 | 1.89% |
| 2–5 | 31 | 2.15% |
| 3–4 | 22 | 2.74% |
| 3–5 | 29 | 4.11% |
| 4–5 | 18 | 3.67% |
| Rank | Pair | Jaccard Overlap |
|---|---|---|
| 1 | 2–3 | 4.69% |
| 2 | 1–2 | 4.21% |
| 3 | 3–5 | 4.11% |
| 4 | 4–5 | 3.67% |
| 5 | 1–3 | 3.04% |
| 6 | 3–4 | 2.74% |
| 7 | 1–4 | 2.31% |
| 8 | 2–5 | 2.15% |
| 9 | 2–4 | 1.89% |
| 10 | 1–5 | 1.44% |
| Token | Scribes |
|---|---|
| cheockhy | 1, 2, 3, 4 |
| odor | 2, 3, 4, 5 |
| ofaiin | 1, 2, 3, 5 |
| okody | 1, 2, 4, 5 |
| Pair | Matches | Jaccard Overlap |
|---|---|---|
| 1–2 | 35 | 3.80% |
| 1–3 | 21 | 3.05% |
| 1–4 | 19 | 3.11% |
| 1–5 | 16 | 2.74% |
| 2–3 | 32 | 4.93% |
| 2–4 | 22 | 3.79% |
| 2–5 | 23 | 4.18% |
| 3–4 | 8 | 2.31% |
| 3–5 | 15 | 4.82% |
| 4–5 | 10 | 4.22% |
| Rank | Pair | Shared Tokens | Jaccard Overlap |
|---|---|---|---|
| 1 | 3–5 | 40 | 15.87% |
| 2 | 2–3 | 81 | 12.27% |
| 3 | 1–2 | 107 | 11.33% |
| 4 | 4–5 | 15 | 10.07% |
| 5 | 1–3 | 61 | 8.96% |
| 6 | 3–4 | 22 | 7.80% |
| 7 | 1–4 | 38 | 6.60% |
| 8 | 2–4 | 32 | 5.51% |
| 9 | 1–5 | 21 | 3.61% |
| 10 | 2–5 | 18 | 3.09% |
| Rank | Pair | Shared Tokens | Jaccard Overlap |
|---|---|---|---|
| 1 | 2–3 | 380 | 16.51% |
| 2 | 3–5 | 145 | 16.00% |
| 3 | 1–2 | 520 | 15.42% |
| 4 | 1–3 | 292 | 11.95% |
| 5 | 2–5 | 199 | 9.91% |
| 6 | 4–5 | 61 | 9.84% |
| 7 | 3–4 | 101 | 9.59% |
| 8 | 1–4 | 192 | 8.83% |
| 9 | 2–4 | 168 | 7.84% |
| 10 | 1–5 | 139 | 6.54% |
quimqu > 02-06-2026, 02:13 PM
(02-06-2026, 01:24 PM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.So out of all the rare words ( of which I count 1963 ) 84 fall exclusively within the Herbal pages? Also 82 ( Herbal-Stars ) fall exclusively within both the Herbal and Stars pages? Similarly, 5 exclusively within the Text pages? Or have I got this wrong?
But also what is your objective? Is it to determine whether rare words are distributed randomly or are localised within sections or by illustration type?
It seems to me that localised is going to be adequately explained under both the meaningful and meaningless hypotheses.