Morphology
Project analyses report non-random surface and morphemic structure under specified comparison procedures.
LUXVAR is a constructed symbolic language created by Reza Nirouyar within the VARZIN Project. Its morphology and thematic organization are designed; independent semantic emergence is not assumed.
Designed morphemic and structural regularities are real properties of the constructed system. Whether models or humans recover intended semantics independently is a separate empirical question.
Core-30 is organized around five designed axes: Light, Reflection, Silence, Gate, and Motion. Separately, the v2.2 manuscript reports a broader layer of 801 stable roots distributed across six semantic fields and other generator dimensions. Those two design descriptions should not be silently equated: the relationship between the five Core-30 axes and the manuscript’s six broader semantic fields is retained as a source-level distinction. The 801 roots are designed items, not 801 independently validated natural-language lexical entries.
Project analyses report non-random surface and morphemic structure under specified comparison procedures.
Thirty registered entries across five designed axes, with a historical manifest-hash prefix preserved in the record.
No external field, frequency-based meaning, or non-human origin is established by the current evidence.
Project-reported morphology: z = 4.36, p < 0.001; silhouette = 0.406 versus shuffled mean 0.174. The reported natural cluster count was k = 3; the k = 6 Hexacore hypothesis was rejected under that analysis.
GEN-001 reported 98.9% five-fold cross-validation accuracy with delta-majority +32.2% in the tested setup. GEN-003R+ reported uniqueness 0.738, 95% CI [0.703, 0.781], against six curated reference-language corpora.
SEM-001 reported no semantic-axis string signal. SEM-002 scored 1/4 against a criterion of at least 2/4; the criterion was not met.
Track D reported Mistral 7B ARI = −0.0337 and Llama 3 8B ARI = 0.0152 against intended axis/orbit labels. This is a negative result for those tested models under that protocol, not a universal statement about AI systems.
Core-30 remains the canonical 30-word reference subset, but the later research stack is much larger. The LUXVAR v2.2 manuscript reports 801 stable roots and an approximately 1.3-billion-form generated corpus. Later work adds a fresh 300-word scale/generalization variant, a separate 360-word true-group-position benchmark, and the 10,800-record VARZIN V2 frozen dataset.
Important: these are distinct experimental artifacts. The site does not relabel them as one enlarged Core-30 or treat generated forms as independently validated natural-language words.
Historical title note: the archived LUXVAR v2.2 title retains the phrase “AI-Recoverable” for citation continuity. The current project interpretation does not treat independent recovery of the intended semantic axes as established.