LUXVAR / scale & benchmark lineageCore-30 was the reference set.
The research stack grew far beyond it.
The public record now separates the canonical 30-word reference lexicon from larger generated corpora and later experimental benchmarks. Scale is reported without treating every generated item as an independently validated lexical unit.
30 reference words801 stable roots~1.3B generated forms10,800 V2 records
03 / 10× scale testA freshly generated 300-word variant.
Appendix E of the Morphological Hijacking manuscript introduces a fresh 20-root × 15-prefix variant: 300 words, ten times the Core-30 word count. Its frozen GPT2-small baseline was weaker than the original adversarial Core-30 condition (ARI −0.053, PHR 0.869), consistent with a milder prefix/root token-length imbalance.
With the full trained recipe, leave-one-root-out evaluation over five held-out roots and three seeds each produced pooled ARI = 0.992, SD = 0.007. The paper explicitly limits this result to the easier same-script, same-word-order condition; it is not presented as outperforming the harder cross-script 7B tests.