Morphology
Do strings contain designed, non-random morphemic or phonotactic regularities under the specified controls?
VARZIN treats morphology, clustering, semantic-axis recovery, explicit training effects, and held-out composition as distinct questions. The purpose of the audit is to prevent success on one task from being promoted into evidence for another.
Some tested procedures recover surface or morphemic structure, while direct frozen-model tests did not recover the intended semantic axes. Later targeted training is a separate intervention, not evidence of spontaneous recovery.
Do strings contain designed, non-random morphemic or phonotactic regularities under the specified controls?
Do models or algorithms produce stable groupings, and what labels or structure do those groups actually track?
Do outputs align with the intended axes without being given those axes through supervision or projection?
Interpretive rule: agreement between models, good classification accuracy, or non-random morphology does not by itself demonstrate independent recovery of intended semantics.
The LUXVAR preprint reports Claude/Gemini cluster ARI = 0.812 as inter-model surface or morpheme agreement. The project correction explicitly warns that this should not be read as semantic-axis recovery, and the raw pairwise scoring was not independently re-verified at correction time.
The later frozen, unremediated test reports Mistral 7B ARI = −0.0337 and Llama 3 8B ARI = 0.0152 against intended axis/orbit labels. Under that protocol, the intended axes were not recovered.
SEM-001 reported no semantic-axis signal at string level. SEM-002 scored 1/4 against a criterion of at least 2/4, so the criterion was not met.
Later project documentation reports strong targeted recovery after explicit training or projection. Those experiments can be meaningful as representation-learning results, but they are not evidence that frozen models independently discovered the intended LUXVAR axes.
The documented experimental record distinguishes Appendix F group-position classification from Appendix G held-out composition. Appendix G reports seen performance of 0.996 but TRUE 0.106, SHUFFLED 0.281 and WRONG_OP 0.175 in the stated procedure; V2 Qwen Phase2 separately reports 0/58 passes. These are scoped negative results, not universal claims about LLM algebra.
The project-listed record Morphemes vs. Manifolds: Diagnosing Structural Blindness and Morphological Hijacking in Large Language Models via Finite Affine Orbits is the central publication for this line of work.
Representation recovery and systematic composition are treated as separate questions. The version-specific public record is 10.5281/zenodo.22679978.