Independent research on AI evaluation infrastructure and structural systemic risks.
The failure mode this archive describes is not model collapse, not benchmark contamination, not distribution drift. It is a functional shift of the evaluation categories themselves — boundaries persist while what they separate becomes progressively opaque.
Recent papers
Substrate Independence as a Structural Requirement
Three necessary conditions — lineage, anchor, rate — that any contamination-resistant evaluation design must satisfy, and what seven candidate families actually purchase against them.
The Reflexive Contamination of Categorical Benchmarks
The mechanism by which recursive use erodes a categorical benchmark’s discriminative function — requiring neither vendor intent, training-data leakage, nor exogenous drift.
The Audit Asymmetry Problem
Capacity conditions — bandwidth, semantic reach, self-reference depth — under which substrate-independent audit ceases to be informative.
A Category-Level Failure Mode Not Captured by Distribution Metrics
Externality conditions for substrate-independent audit, derived from the 2008 Gaussian copula collapse.