The logic and theory behind it is simple.
The VMS has a highly structured, specialized architecture:
First, I use an in-depth frequency analysis of this language and the VMS to assign one—or, in this case, multiple—meanings to specific VMS morphemes within its structure and architecture, meanings that reasonably align with the frequencies of the other language. To this end, I apply a few rules that sort these morphemes according to VMS logic to ensure its repetitive nature: Furthermore, I need 7 rules to distribute spaces to specific end markers that best match 7 end markers, a few line-start glyphs and line-end glyphs, and I determine the word lengths according to another pattern or with a certain degree of randomness; then, in theory, I should actually be able to generate a VMS-like text using the real language.
Since morphemes in normal languages are also highly repetitive, I can replicate the core structure of VMS with some adaptation—since the structure of VMS will be much steeper than that of the language, I naturally need a distribution that assigns some morphemes from the other language to the few morphemes in VMS. So you’re going from a lot of information to very little information. You’re compressing. And compression toward less diversity is always feasible.
To generate the high Happax rate, I don’t encode individual words but rather these very morphemes, and as mentioned above, I split the words at specific “morpheme boundaries” that I define in advance according to the VMS logic.
Of course, this works better with a slightly more sophisticated cipher optimized for the VMS.
But what did I do instead? I adapted the language to the structure of the VMS. Of course, I could also do this with a Markov generator—and, in fact, in any way I choose. The more finely I do this, the more rules I use, the more stable it would become—but also the more nonsensical...
I managed to do this with very simple rules because I utilized the structure of the VMS on a different level, but that doesn’t prove anything.
