Skip to content
VARZINIndependent research
Menu
LUXVAR / scale & benchmark lineage

Core-30 was the reference set.
The research stack grew far beyond it.

The public record now separates the canonical 30-word reference lexicon from larger generated corpora and later experimental benchmarks. Scale is reported without treating every generated item as an independently validated lexical unit.

30 reference words801 stable roots~1.3B generated forms10,800 V2 records
Key distinction

“Bigger than Core-30” does not mean one canonical lexicon was simply enlarged. The papers introduce several distinct artifacts for different scientific questions.

01 / canonical reference

Core-30 remains the anchor set.

Core-30 is the 30-word reference lexicon used in early clustering, semantic-axis, and morphological-hijacking experiments. It remains useful because it is fixed and historically traceable, but it is no longer an adequate description of the full experimental scale of VARZIN/LUXVAR.

02 / broader LUXVAR layer

801 stable roots and a combinatorial corpus.

Designed roots

801 stable roots

The LUXVAR v2.2 manuscript analyzes a broader designed root inventory. It describes these roots across six semantic fields and other generator dimensions. That source-level description is distinct from Core-30’s five designed axes and is not silently collapsed into the same taxonomy.

Generator

~1.3 billion forms

The manuscript reports an exhaustive combinatorial generator producing about 1.3 billion forms. This is generated structure by design, not a corpus collected from human language use.

Balance check

n = 2,000,000 audit slice

At the reported audit scale, H/Hmax = 0.9995 and the stated first-order Cramér’s V associations remain below 0.02.

Taxonomy note: Core-30 uses five designed axes; the broader v2.2 corpus description uses six semantic fields. VARZIN preserves that discrepancy as documented rather than rewriting one into the other.

03 / 10× scale test

A freshly generated 300-word variant.

Appendix E of the Morphological Hijacking manuscript introduces a fresh 20-root × 15-prefix variant: 300 words, ten times the Core-30 word count. Its frozen GPT2-small baseline was weaker than the original adversarial Core-30 condition (ARI −0.053, PHR 0.869), consistent with a milder prefix/root token-length imbalance.

With the full trained recipe, leave-one-root-out evaluation over five held-out roots and three seeds each produced pooled ARI = 0.992, SD = 0.007. The paper explicitly limits this result to the easier same-script, same-word-order condition; it is not presented as outperforming the harder cross-script 7B tests.

04 / true group position

A distinct 360-word algebraic benchmark.

Appendix F builds a separate 5 roots × 6 prefixes × 12 positions = 360-word lexicon with an opaque, shuffled suffix mapping to prevent trivial numeral/ordinal leakage.

GPT2

True-label ARI 0.791 ± 0.105

Shuffled-label control: 0.025 ± 0.006.

Mistral

True-label ARI 0.929 ± 0.047

Shuffled-label control: 0.299 ± 0.044.

Qwen

True-label ARI 0.996 ± 0.004

Shuffled-label control: 0.232 ± 0.048.

Boundary: this is targeted group-position classification after explicit remediation. The same paper separately reports that this success does not establish homomorphic or systematic composition.

05 / VARZIN V2

The later frozen experiment uses 10,800 records.

Dataset

10,800 records

12 structural positions, four synthetic morphemes, 30 contexts, 24 training contexts, 6 held-out contexts, and 30 instances per position/context configuration.

Representation

Mean BA 0.9666

Across all 29 hidden-state indices, the preregistered surface-span representation was highly decodable; 25/29 indices reached BA = 1.000.

Composition

0 / 58 PASS

Across 29 indices × two composers, no combination met the complete preregistered positive criterion on unseen ordered pairs.

06 / evidence hierarchy

What “scale” means in the current VARZIN record.

Supported description

  • Core-30 is a fixed canonical reference subset.
  • The broader LUXVAR design contains 801 stable roots.
  • The generator can produce ~1.3B designed forms.
  • Later papers introduce independent 300-word, 360-word, and 10,800-record benchmarks.
  • Scaling experiments preserve their own controls, targets, and limitations.

Do not infer

  • That 1.3B generated forms are 1.3B empirically validated words.
  • That every later dataset is part of one canonical lexicon.
  • That group-position recovery proves composition.
  • That high probe accuracy proves semantic understanding.
  • That results from one benchmark automatically transfer to another.
07 / primary records

Trace each scale claim to its source.