The Reflexive Contamination of Categorical Benchmarks
Use-induced erosion of evaluation infrastructure within capacity asymmetry
1 Introduction
The first note in this series identified category-level failure: a shift in the function of evaluation categories that distribution metrics do not capture (Ortulanu, 2026a §2). The second note specified the audit relation under which such failure becomes possible, naming capacity asymmetry as the condition at which an auditor no longer records, distinguishes, or stands outside the task-relevant state of the auditee (Ortulanu, 2026b §2.3). Within the second note’s frame, two empirical cases were read as capacity inversions: the 2008 collapse of the AAA category across structured-finance asset classes (Ortulanu, 2026b §6, stage 4) and the bias profile of contemporary LLM-as-judge architectures (Ortulanu, 2026b §7). The use channel that produced those inversions was named, but not developed. The second note reserved that training-evaluation feedback process for the present analysis (Ortulanu, 2026b §7).
The present note takes up that mechanism. It argues that the conditions traced in the second note converge in a recurrent dynamic: a categorical benchmark used recursively in the training or selection of the systems it evaluates undergoes erosion of its discriminative function — a process internal to the benchmark’s operation, requiring neither vendor intent, nor leakage of benchmark data into training corpora, nor any exogenous shift in benchmark inputs. The note names this process reflexive contamination of categorical benchmarks and treats it as one mechanism by which the bounded auditability introduced in the second note (Ortulanu, 2026b §5) loses its boundary over time.
Reflexive contamination is not a fourth failure condition and not a fourth capacity dimension. It is a use-induced pathway through which the failure conditions named in the first note assemble inside the capacity relation specified in the second. The failure conditions come from the first note, the capacity frame from the second, and the use channel from the present note. Of the first note’s three conditions, the use channel most directly reifies self-reference — the benchmark’s judgments become inputs to the systems it later judges — and produces anchor drift and proxy displacement as joint effects rather than as the primary driver.
The claim is sharp in one respect and modest in another. Sharp: reflexive contamination is not reducible to the four mechanisms the substrate-internal audit failure literature already contains. Model collapse at the distribution level (Shumailov et al., 2024), training-data leakage (Chen et al., 2025), the four variants of Goodhart’s law (Manheim & Garrabrant, 2018), competition-induced ratings shopping (Skreta & Veldkamp, 2009) — each captures part of what reflexive contamination produces; none captures the mechanism by which the production proceeds. Reflexive contamination occupies a position among these four frames, decomposing a sub-mechanism that the existing taxonomy leaves unanalyzed. Modest: the note does not propose a complete external substitute for the contaminated benchmark. It traces the mechanism, operationalizes the markers the second note proposed for its detection, and stops there.
What follows defines the mechanism, separates it from its neighbors, restages the second note’s two cases — the 2008 AAA cross-asset-class inversion and the contemporary LLM-as-judge bias profile — at the level of mechanism rather than phenomenon, develops the second note’s three markers into measurement protocols, and reads the 2024–2026 AI evaluation crisis as a contemporary forensic instance before turning to the limitations and the open questions.
Scope. The note treats reflexive contamination as a static aggregate phenomenon: the contaminated state at a single time point, or the cumulative state produced by earlier cycles. The dynamics by which contamination proceeds cycle by cycle in a recursive training pipeline — the multi-period evolution toward the contaminated state — belong to the fifth note in the series. The two analyses are complementary; this one fixes the endpoint, the later one will trace the trajectory. Where the trajectory is itself the object of interest, these tools apply only as the description of the limit they approach.
Methodological commitment. The second note traced capacity asymmetry through three formal sources as constraint markers rather than as a full derivation. This note follows the same commitment. The mechanism it names is not derived from a formal model; it is identified as the recurring structural feature of two well-documented forensic cases — one in finance, one in AI evaluation — and operationalized through the markers the second note proposed. What would constitute a formal model of reflexive contamination is left open in section 7.
Accordingly, the argument adds a mechanism to the capacity-asymmetry frame rather than a new theoretical apparatus. The frame is the second note’s; the mechanism is what the second note named without developing.
2 The Mechanism — Defining Reflexive Contamination
A categorical benchmark in the sense developed here is an audit instrument that produces a finite set of distinguishable judgments — a credit rating in the set {AAA, AA, A, BBB, …}, an evaluation outcome in the set {pass, fail}, a pairwise preference in the set {A wins, B wins, tie}. The judgments admit no continuous gradation between distinguishable categories; what the benchmark records is membership in one element of the set. This is the categorical case of the audit relation the second note describes, in which L_A’s vocabulary and decision procedure are jointly sufficient within a bounded subclass (Ortulanu, 2026b §5); here L_A is finite and the decision procedure is total. Bounded auditability is the description of where the form succeeds: the benchmark renders informative judgments within the subclass of auditee behaviors that its categories were designed to distinguish.
A benchmark can be categorical in either of two ways. Some are categorical by output: they return a finite label directly — AAA, pass, A-wins. Others are categorical by use: they return a scalar score that the surrounding infrastructure converts into a finite decision — release or discard, frontier or non-frontier, saturated or still informative. The mechanism developed here applies to both. The second kind carries an added evidentiary burden, because the contamination operates on the categorical use rather than on the scalar output: the conversion from score to category has to be documented before the mechanism can be claimed.
Reflexive contamination is the erosion of the boundary of that subclass through the benchmark’s recursive use. A benchmark is used recursively when its categorical judgments feed back into the training, selection, or design of the systems it subsequently evaluates. Credit ratings that influence the structuring of new collateralized instruments are recursive in this sense; LLM judges whose outputs serve as reward signals in RLAIF training are recursive in this sense; a benchmark used to select which model versions are released and which discarded is recursive in this sense. The use establishes a feedback channel from category to auditee that runs in the same direction as the audit channel from auditee to category. The two channels operate on the same population.
Under this configuration, two distinct conditions of the auditee that the benchmark originally distinguished can migrate, over time, into the same audit category — without the benchmark’s procedure changing, without any exogenous shift in benchmark inputs, and without any participant gaming the benchmark’s specification. The mechanism is the use itself: each cycle of recursive use selects auditee features that map cleanly onto the benchmark’s existing categories, and demotes auditee features that the categories cannot accommodate. After enough cycles, the categories no longer mark distinctions that were originally meaningful — not because the world has changed, but because the population the benchmark evaluates is now the population the benchmark itself shaped. The population shift is endogenous: the benchmark’s own use produces it, and the ordinary diagnostics that watch for an exogenous change in inputs register nothing. The category retains its name. What it sorts is no longer what the name refers to.
This is the auditee-side mechanism by which bounded auditability erodes from within. The boundary of the subclass within which the benchmark’s judgments remain informative — the set of auditee behaviors for which L_A’s vocabulary and decision procedure are jointly sufficient — contracts as the auditee population shifts to fit the categories the benchmark already accommodates. The contraction is invisible from inside the benchmark’s outputs: the benchmark continues to produce judgments, the judgments continue to be acted on, and inter-judge agreement may even rise (when the contracting subclass becomes more internally homogeneous). What falls is not the benchmark’s apparent quality but the breadth of what the benchmark can task-relevantly distinguish.
Intent is not required. Reflexive contamination requires no gaming on any participant’s part. Gaming, where it is present, accelerates the mechanism without being necessary to it. A credit rating agency that performs its analysis honestly, an LLM judge whose evaluation procedure has no exploitable bias, a benchmark whose specification is fully public and fully respected — each of these can still undergo reflexive contamination. The mechanism is structural: it operates wherever a categorical benchmark’s outputs participate in shaping the population the benchmark subsequently evaluates. The benchmark does not break. It goes on producing well-formed judgments in a language whose distinctions the evaluated population has learned to inhabit. The point is taken up in section 3 as the precise distinction from Goodhart’s law, which the existing literature treats as the canonical frame for proxy degradation under optimization pressure.
The unit of analysis is the category, not the distribution. The contamination described here operates at the categorical layer — at the level of the benchmark’s discriminative function across categories. It need not appear at the distributional layer: a contaminated benchmark may continue to produce judgments with rich distributional properties (high variance, broad coverage of the input space, low correlation with naive baselines) while having lost the capacity to separate categories the audit task required to be separated. The orthogonality between distribution-level diagnostics and category-level failure was the central claim of the first note in this series (Ortulanu, 2026a §2). Reflexive contamination is one mechanism by which the category-level failure that note named takes shape.
3 Distinction from Neighboring Mechanisms
The mechanism of section 2 sits adjacent to several existing literatures. Four treat substrate-internal audit failure directly; two more, drawn from neighboring disciplines, name a kindred dynamic; a final pair bounds the mechanism at its edges. The position is not orthogonal — each adjacent frame captures part of what reflexive contamination produces. Each is also structurally distinct from the mechanism, in ways that require explicit treatment.
3.1 Distribution Collapse
Shumailov et al. (2024) document the collapse of output distribution diversity in language models trained recursively on synthetic data. The mechanism is distributional: as model-generated content accumulates in training corpora, the distribution from which subsequent models are trained loses its tails, drifts toward modal samples, and converges on a narrow region of the original distribution. The diagnostic surface is the distribution itself — measures of perplexity, variance, and tail coverage register the collapse.
Reflexive contamination operates at a different layer. The categorical benchmark may retain rich distributional coverage of its inputs — receiving submissions across a wide variety of model architectures, prompt strategies, and task formulations — while its discriminative function erodes. Distribution-level diagnostics will read the inputs as healthily diverse; the contamination resides not in the distribution of submissions but in the category structure the benchmark applies to them. Where distribution collapse is the loss of tails, reflexive contamination is the loss of boundaries. The two can coexist or diverge; they are not the same mechanism.
3.2 Training-Data Leakage
A separate literature treats benchmark contamination at the level of training-data composition. Chen et al. (2025) survey methods for detecting whether specific benchmark items appear in training corpora — n-gram overlap, membership inference, completion-likelihood probes — and trace the shift from static to dynamic benchmarking as the field’s response to widespread contamination findings across popular benchmarks. The mechanism is leakage: the auditee has seen the test items during training and reproduces remembered answers.
Reflexive contamination requires no such leakage. A benchmark whose items have never appeared in any training corpus can host the mechanism described here, provided the benchmark’s judgments feed back into the training or selection pipeline of the systems it evaluates. The leakage frame and the use-induced frame address structurally different failure conditions: leakage is the failure mode in which the auditee has memorized the test; reflexive contamination is the failure mode in which the test has been used to shape the auditee. Both can obtain in the same evaluation system, and recent reasoning-model work documents both in the same systems (Wu et al., 2025); they remain analytically distinct.
3.3 Goodhart’s Law
Goodhart’s law is the closest adjacent frame, and the distinction requires the most precision. Manheim and Garrabrant (2018) catalog four variants: Regressional Goodhart (an imperfect proxy fails to converge with the goal under optimization pressure); Extremal Goodhart (proxy-goal correspondence breaks down at extreme values); Causal Goodhart (the causal model linking proxy to goal is misspecified); Adversarial Goodhart (an agent actively games the proxy specification).
Reflexive contamination is better read not as a fifth Goodhart variant, but as a mechanism that can produce Goodhart-like failures. Each Goodhart variant identifies a failure mode and its trigger: proxy imperfection (Regressional), extreme values (Extremal), causal misspecification (Causal), or active gaming (Adversarial).
Reflexive contamination is a generative mechanism that can produce failures resembling two of those modes — Extremal and Adversarial — without their distinguishing triggers. It produces Extremal-like failures (proxy-goal correspondence breaking at the upper end of the score range, where most models cluster) without treating extreme-value fragility as the primary cause; it produces Adversarial-like failures (auditee behavior aligning with the proxy’s measurement criteria) without any participant choosing to game the proxy. The Goodhart literature treats the modes; reflexive contamination identifies one of the generative mechanisms that produces them, beginning earlier than the proxy-goal relation — in how categorical judgments feed back into the population being judged.
The relation to Goodhart is therefore analytic rather than rivalrous. Reflexive contamination occupies a position the Goodhart taxonomy as it stands does not separately mark — the position at which use alone, even under aligned incentives and without intentional gaming, suffices to produce the failures the Extremal and Adversarial modes describe. This places the mechanism inside the counterfactual the second note advances for capacity asymmetry generally (Ortulanu, 2026b §2.5): audit failure can proceed even when no actor has the incentive to corrupt the process.
3.4 Ratings Shopping
Skreta and Veldkamp (2009) formalize the contamination of credit ratings under competition between rating agencies: issuers choose among multiple agencies, agencies compete for issuer business, and the resulting equilibrium inflates ratings on complex instruments. The mechanism is competitive: it requires multiple agencies, an issuer-pays compensation structure, and asset complexity that obscures rating accuracy from third parties.
Reflexive contamination is single-substrate. The mechanism described here operates within a single benchmark, or within a community of benchmarks that share substrate, without competition between organizations. A monopoly evaluator with aligned incentives can undergo the same kind of contamination as a competitive market of evaluators. The two frames are complementary: where ratings shopping addresses what competition does to evaluation quality, reflexive contamination addresses what use does — and use is present whether competition is or not.
3.5 Adjacent Frames Beyond the Substrate-Internal Literature
Two further frames name the same feedback dynamic at a broader grain. Each differs from reflexive contamination in object.
Campbell’s law is the older statement. Campbell (1976) observed that a quantitative social indicator, once it carries weight in decisions, comes under pressure that corrupts both the indicator and the process it was meant to monitor. The observation is general, covering test scores, crime statistics, and economic targets alike. Reflexive contamination is Campbellian in spirit and narrower in object. Campbell names the corruption of an indicator under decision use; reflexive contamination names one channel of that corruption — the indicator’s categorical judgments become inputs to the production or selection of the population it later judges — and locates the damage in the category boundary rather than in the indicator value. The narrower claim reaches cases the social framing leaves out, including automated evaluators whose proximate feedback channel is technical rather than overtly social.
Performative prediction is the closest formal neighbor. Perdomo et al. (2020) formalize the case in which a deployed prediction changes the distribution of the outcome it predicts, and study the stable points such feedback admits. The object there is the outcome distribution. Reflexive contamination concerns a different object: the discriminative force of the benchmark’s categories, which can erode while ordinary distributional diagnostics still read as healthy. Performativity describes how the population moves; reflexive contamination describes what the movement costs the category. The financial literature carries the performative dynamic under its own name: leverage built against measured risk feeds the prices that set measured risk (Adrian & Shin, 2010), and agreement can be coordination on a public signal rather than accuracy (Morris & Shin, 2002). The layer there is the price system and the driver is optimizing response; the categorical claim begins where neither is required.
Two more distant frames bound the mechanism on other sides. Adaptive leaderboard overfitting (Blum & Hardt, 2015) treats the loss of holdout validity under repeated test access; reflexive contamination can proceed even when the test items stay unseen, provided the judgments shape the auditee. Construct validity (Cronbach & Meehl, 1955) asks whether a score supports the inference drawn from it. Reflexive contamination is the time-indexed erosion of that support through recursive use rather than a static failure of it.
3.6 Summary of Distinctions
| Adjacent frame | Layer | Mechanism | What reflexive contamination adds |
|---|---|---|---|
| Distribution collapse (Shumailov et al. 2024) | Distribution | Recursive training on synthetic data | The categorical layer eroding while distribution holds |
| Training-data leakage (Chen et al. 2025) | Data composition | Memorization of test items | The benchmark shaping the auditee even without leakage |
| Goodhart’s law (Manheim–Garrabrant 2018) | Proxy-goal relation | Optimization pressure on imperfect proxies | The use-without-intent sub-mechanism beneath the four variants |
| Ratings shopping (Skreta–Veldkamp 2009) | Market structure | Competition over multiple raters | The single-substrate mechanism without competition |
| Campbell’s law (Campbell 1976) | Social indicator / governance | Decision use corrupts the indicator and the monitored process | Category-boundary erosion via recursive training or selection, including automated evaluators with a technical feedback channel |
| Performative prediction (Perdomo et al. 2020) | Prediction / outcome distribution | Deployed predictions shift the outcome distribution | Loss of category discriminative force, not only distributional response |
The mechanism that section 2 names isolates a residual structure these frames do not separately mark: category-boundary erosion through recursive use of the benchmark’s own judgments.
4 Two Instances Restaged
The second note treated two cases as capacity inversions rather than as recursive-use processes. This section rereads the same material through the use channel defined in section 2: the cases are drawn from different domains (structured finance, artificial intelligence evaluation) and different decades, and their value lies in the correspondence across them.
4.1 The 2008 AAA Cross-Asset-Class Inversion
The AAA credit category, stable for corporate bonds for several decades, lost its capacity to render distinct judgments across the asset classes the rating apparatus absorbed between 2003 and 2007. Asset-backed securities, collateralized debt obligations, ABS-CDOs, and CDO² instruments shared a single category designation while sharing increasingly little of the discriminative content the designation was meant to carry. By the end of the period, more than ninety percent of AAA-rated subprime residential mortgage-backed securities issued in 2006 and 2007 were later downgraded (Senate Permanent Subcommittee on Investigations, 2011). The second note treated this as semantic-reach inversion within a staged capacity account: AAA ceased to mark a distinction the rating language had been designed to carry (Ortulanu, 2026b §6, stage 4).
The mechanism by which the distinction was lost is reflexive contamination in the sense of section 2. The recursive use ran through two channels. The first was the AAA judgment itself, an input to the structuring of new CDOs: CDO managers absorbing the BBB tranches of residential mortgage-backed securities depended on the rating apparatus’s continued willingness to rate the resulting CDO senior tranches as AAA. The second was volume: new AAA-rated structured-finance issuance increased the share of subprime origination that the CDO pipeline could absorb. The two channels reinforced: AAA’s continued application to recombined subprime pools made the recombination commercially viable, and the recombination volume expanded the pool of subprime instruments to which AAA was applied. The mechanism does not depend on the rating apparatus changing its procedure. The auditee population it rated changed — and changed in response to the apparatus’s continued use of the AAA category.
The mechanism does not require gaming in any standard sense. The argument is not that no analyst made errors or that no methodology was relaxed; it is that reflexive contamination can proceed even without such departures. What changed was the referent of the AAA category. In 1995, AAA marked the highest credit quality among a population of corporate issuers and a small set of structured instruments whose underlying collateral was more conservatively underwritten. In 2007, AAA marked membership in a category whose population was dominated by structured instruments whose existence depended on AAA’s continued application. The category’s name was preserved. What the name pointed to had been reshaped through the category’s own use.
The collapse of the category in 2007–2008 was the moment at which the discrepancy between the name and its referent became externally visible. The wave of downgrades that began in July 2007 marked the audit cycle’s loss of informativeness as a recognized event (Ortulanu, 2026b §6, stage 5). The mechanism the present note describes was operative for at least four years before that moment.
4.2 The LLM-as-Judge Bias Profile
The contemporary case admits a parallel reading. Large language models used as judges of other large language models — the LLM-as-judge architecture documented by Zheng et al. (2023) and related model-based evaluation and feedback procedures, including Constitutional AI, RLAIF, and model-based red-teaming (Bai et al., 2022) — produce evaluations that exhibit a documented bias profile: preference for verbose responses (Saito et al., 2023), sensitivity to position in pairwise comparisons (Shi et al., 2025), favorability toward outputs that resemble the judge’s own (Panickssery, Bowman, & Feng, 2024). The second note treated these biases as surfaces of a single capacity inversion across three dimensions (Ortulanu, 2026b §7).
Reflexive contamination explains these biases as use-induced rather than merely evaluator-internal. The judge’s evaluations function not only as audit judgments but as training signals: in RLAIF and Constitutional AI pipelines, the judge’s preferences become the optimization target for the next iteration of the systems the judge will subsequently evaluate. A verbosity-preferring judge rewards verbosity in the next training cycle; a judge that favors outputs resembling its own rewards judge-like style; position sensitivity can reward artifacts correlated with presentation order rather than task quality. The systems being trained converge toward response styles the judge reliably rewards. Pairwise comparisons continue to produce winners while the categorical content of the judgments narrows toward the subclass of behaviors the judge can still distinguish.
Li et al. (2025) document this dynamic at the level of preference leakage: bias of LLM judges toward systems trained on data whose evaluation the judge contributed to producing. Across the three relatedness channels Li et al. identify — generator-judge identity, inheritance lineage, shared model family — the judge’s prior preferences become predictive of the judge’s subsequent preferences, in a way that risks converting apparent agreement into family-level consistency rather than independent accuracy. The bias profile catalogued in the earlier LLM-as-judge literature is what preference leakage looks like in the limit: a judge whose evaluations describe its own contribution to the auditee’s training, presented as evaluations of the auditee’s intrinsic properties.
The mechanism parallels the AAA case at the level of feedback structure. The judge’s use as a training signal shapes the population the judge subsequently evaluates; the judge’s continued application produces evaluations of a population that exists, in part, because of the judge’s prior applications. The category “model passes the LLM-as-judge evaluation” retains its name. What the name marks has, through recursive use, contracted to the subclass of model behaviors the judge can distinguish — which is, increasingly, the subclass of behaviors that prior cycles of judge-based training have selected for. The contraction is invisible from inside the evaluation outputs.
5 Measurement Protocols for the Three Markers
The second note proposed three internal markers of audit detachment: score saturation, agreement on a shrinking task, and divergence from deployment outcomes (Ortulanu, 2026b §7). The markers were proposed as observable signatures of the capacity-asymmetry condition under which the audit cycle had ceased to be informative. A measurement protocol for each follows, with explicit treatment of the background variation it must separate from the capacity-asymmetry signal.
The protocols are deliberately substrate-internal: they operate on the audit cycle’s own outputs and the auditee’s deployment record, without requiring access to information that the cycle does not already produce. The cost of this choice is that no single protocol on its own confirms reflexive contamination; the protocols’ joint operation is what supports the diagnosis. The structure mirrors the first note’s argument that no single externality condition is sufficient — only the joint satisfaction of three is.
5.1 Marker 1 — Benchmark Saturation with Falling Tractability of Independent Verification
The first marker has two components. The first is benchmark score saturation: scores on a fixed benchmark approach a ceiling, and frontier systems cluster within a narrow band near the top of the score range. The second is the rising cost — in time, expertise, or data — required to construct an independent verification of those scores: replication, oracle comparison, or distillation to a smaller, contamination-resistant test.
The measurement protocol consists of two longitudinal time series: (a) the top-quartile score on the benchmark over time, and (b) the per-item construction cost of replacement benchmarks targeting the same capability class — measured by expert time, review time, rejection rate, and required expertise. Marker 1 is positive when (a) approaches its ceiling while (b) rises at a rate disproportionate to capability gains measured on independent, non-overlapping evaluations. The background variation to separate is task-intrinsic saturation — capability genuinely approaching the limit the task permits — which would be characterized by (b) holding constant or falling rather than rising. Task-intrinsic saturation produces (a) plateauing without (b) increasing; reflexive contamination produces both. The 2024–2026 benchmark record, in which several public benchmarks lost frontier-ranking power while replacement benchmarks grew costlier to construct, is treated as the empirical instance in section 6.4.
5.2 Marker 2 — High Inter-Judge Agreement on a Contracting τ
The second marker measures agreement among independent judges of the audit task, normalized for the scope of the task being judged. Inter-judge agreement is standardly measured via Cohen’s κ for pairwise judges, weighted κ when disagreement-degree matters, Fleiss’ κ for multi-judge configurations, or Krippendorff’s α for more general settings. Agreement that rises in raw value is, on its own, ambiguous: it may indicate that the judges have converged on the correct task-relevant distinctions, or it may indicate that the task they are jointly distinguishing has contracted.
The protocol therefore requires a second measurement alongside agreement: the scope of the audit task τ being judged at each time point. Scope contraction can be documented through rubric history (rubrics that progressively exclude difficult or ambiguous cases), exclusion criteria (cases that were judged in earlier rounds but are now ruled out of scope), and the distribution of judgment categories actually used (categories that nominally exist but are no longer applied to any sampled cases). Marker 2 is positive when raw agreement rises while τ scope contracts at a comparable rate. Operationally, the scope of τ can be proxied by category-usage entropy, rubric coverage, the exclusion rate relative to earlier evaluation rounds, and task-family diversity within the evaluated set.
Three recent results bear on both measurements. Han et al. (2025) report human-versus-LLM-judge agreement at Cohen’s κ = 0.85 for result evaluations and 0.78 for trajectory evaluations; Haldar and Hockenmaier (2025) document that LLM judges exhibit substantial self-inconsistency across repeated evaluations of the same input, suggesting that nominal inter-judge agreement may overstate the stability of the underlying judgments. Shi et al. (2025) report that position bias varies strongly with the quality gap between solutions — a finding that would, under reflexive contamination, predict rising apparent agreement as systems converge on similar quality levels in the contracted τ.
The background variation to separate is genuine convergence on a stable task — agreement rising while τ scope holds constant, indicating that judges have aligned on the right distinctions. The distinction turns on tracking τ scope independently of the agreement measurement.
5.3 Marker 3 — Divergence Between Evaluation Results and Downstream Consequence
The third marker measures the correlation between the audit cycle’s judgments and the downstream consequences observed when the auditee is deployed. Where the audit cycle is informative, the correlation should be substantial: better evaluation scores should predict better deployment outcomes, controlling for confounders. Where the audit cycle has ceased to be informative through reflexive contamination, the correlation should fall — not necessarily to zero, but to a level below what the evaluation purports to measure.
The AI Incident Database listed more than 5,000 reports and more than 1,000 incidents by 2026 (AI Incident Database, 2026; semantic-association analysis in Russo et al., 2025). The qualitative point is independently visible in that record: many deployment failures fall outside the task envelope of standard research benchmarks. The quantitative form would correlate release-time benchmark rank with incident rates for the same model version at fixed post-deployment intervals, controlling for deployment scale and reporting bias.
The protocol’s implementation is harder than the protocols for markers 1 and 2, because the downstream-consequence data is itself produced under conditions that introduce noise: incident reporting is selective, deployment scale varies across model versions, and the time lag between deployment and incident accumulation differs across failure modes. The measurement therefore requires paired longitudinal data — benchmark scores at release, incident rates at fixed time-since-deployment intervals — with explicit modeling of the reporting and scale confounders.
The protocol assumes downstream consequence is observable at all. Where the auditee’s design renders certain downstream consequences structurally unobservable to external monitors — as in cryptographic supply integrity in privacy-preserving systems — marker 3 cannot operate even when capacity asymmetry is present. The limit case is treated at length in a later supplementary note in this series (Ortulanu, 2026h) and revisited in section 7.
The background variation to separate is task-evaluation gap — benchmarks that were never designed to predict downstream consequence (memorization benchmarks, narrow-skill benchmarks) and would not predict it even in the absence of reflexive contamination. Marker 3 applies only to benchmarks that explicitly purport to predict deployment-relevant capability; for those, divergence bears on the diagnosis jointly with markers 1 and 2.
5.4 Joint Operation of the Three Markers
No single marker on its own supports the diagnosis of reflexive contamination. Marker 1 can be produced by task-intrinsic saturation; marker 2 can be produced by genuine convergence; marker 3 can be produced by task-evaluation gap. The diagnostic confidence rises with the joint operation of the three — and rises further when the causal direction of the joint pattern is consistent with the use-induced mechanism rather than alternative explanations.
The disconfirming event the second note identified — the arrival of a single recognized failure event of the kind July 2007 supplied for structured finance — would indicate that detachment had ceased to be the form (Ortulanu, 2026b §7). Such an event marks the transition from the stage at which marker-based protocols are the appropriate diagnostic tool to the stage at which the audit cycle’s loss of informativeness is externally visible without protocol mediation. The protocols are appropriate for detection before such an event; their role declines after.
6 The Contemporary Forensic — 2024–2026 AI Evaluation
The 2008 case treated in section 4.1 admits, in retrospect, a clean reconstruction: the AAA category, the recursive use channel, the population shift, the eventual external bill. The contemporary AI evaluation case admits the same structural reconstruction but at a different stage of its progression. The 2024–2026 evidence is read here as a forensic instance in formation, with the markers of section 5 as the diagnostic frame.
6.1 The Evaluation Crisis as Common Background
Academic surveys and industry retrospectives in 2025 converged on a related diagnosis: major parts of AI evaluation infrastructure were losing their frontier-ranking power. Chen et al. (2025) identified three problems plaguing current benchmarks — inflated scores from data contamination, unfair evaluation due to cultural and linguistic biases, and insufficient testing of process credibility. Industry retrospectives described 2025 as the year “the scorecard broke” (Goodeye Labs, 2025), and by the end of 2025 static public tests were increasingly treated as unreliable signals of frontier capability.
The convergence on a single diagnosis across academic surveys, evaluation specialists, and industry analyses is itself diagnostic. The contamination dynamic had progressed to the stage at which the markers had begun to surface as widely shared informal observations. The forensic value of the period lies in the contamination’s incompleteness — diagnostic markers remained measurable rather than overwhelmed.
6.2 Preference Leakage as a Clear Instance
Li et al. (2025) formalize the contamination dynamic at the level of LLM-as-judge architecture. The work identifies three forms of relatedness between the LLM serving as data generator (producing the training data for systems under evaluation) and the LLM serving as judge (evaluating those systems): the generator and judge are the same model, the generator is in the judge’s inheritance lineage, or the generator and judge belong to the same model family. Each form of relatedness produces measurable bias in the judge’s evaluations of the trained systems — the bias favoring systems whose training data the judge contributed to producing.
The preference leakage score Li et al. introduce quantifies the dynamic across Arena-Hard and AlpacaEval 2.0 benchmarks. The result is general across the three relatedness forms and across the LLM baselines tested. The work names what reflexive contamination looks like when the use channel is short and explicit — the configuration of section 4.2, now measured.
The structural correspondence to the AAA case is direct. The AAA category was applied by rating agencies whose application of AAA was an input to the structuring of new instruments the agencies subsequently rated. The LLM-as-judge architecture, in the configurations Li et al. measure, has the same recursive use channel — the judge’s outputs are an input to the training of systems the judge subsequently judges. The preference leakage score supplies one quantitative surface of the same feedback structure, measured at a stage of the dynamic where measurement is still tractable.
6.3 Position Bias and Self-Inconsistency as Boundary Erosion
Shi et al. (2025) systematize position bias in LLM judges across 15 judges, 22 tasks, and over 150,000 evaluation instances. The principal finding for the present analysis: position bias correlates strongly with the quality gap between solutions being compared. When solutions are clearly different in quality, position has little effect; when solutions are close in quality, position effects dominate. The finding is consistent with marker 2 in its joint form: as the τ being judged contracts — as the systems being compared cluster more tightly in quality — the residual signal the judge can rely on shrinks, and surface features (position, length, fluency) increasingly determine the verdict.
Haldar and Hockenmaier (2025) document a complementary surface: substantial self-inconsistency in LLM judges evaluating the same input across repeated trials. Reported agreement statistics that compute Cohen’s κ across pairs of judges may overstate the stability of the underlying judgments, because the within-judge variance — not separately reported in agreement statistics — accounts for a non-trivial fraction of total variance. Where reflexive contamination has contracted τ, the residual within-judge variance becomes a larger fraction of the remaining discriminative signal, and aggregate agreement statistics misread this as stable convergence.
The combination of Shi et al.’s position-bias finding with Haldar and Hockenmaier’s self-inconsistency finding supplies the evidence needed to test marker 2 rather than, by itself, confirming it. Where rubric history, exclusion criteria, or category-use records show τ contraction, these judge-side instabilities would support the marker’s positive condition.
6.4 The Contamination Crisis as Marker 1 Evidence
The weakening of MMLU, HumanEval, HellaSwag, and the original GSM8K as ranking signals for frontier work supplies the qualitative form of marker 1. These benchmarks are scalar by output and categorical by use: their scores are converted into frontier/non-frontier and saturated/still-informative decisions of the kind section 2 describes. Public trackers (LifeArchitect, 2026) and provider reports place the saturation timing roughly as follows. After o1-preview reached 92.3% on MMLU in September 2024, the benchmark was increasingly treated as saturated for frontier ranking. Gemini 3 Pro reached 93.8% on GPQA in November 2025 and 90.1% on MMLU-Pro in the same window, with both benchmarks treated as saturated soon after. The figures are not drawn from a single controlled evaluation.
Replacement benchmarks have required heavier expert involvement: HLE (Humanity’s Last Exam) and FrontierMath required graduate-level domain expertise for item construction, and in early-2025 evaluations placed leading models well below the ceiling levels reached on older public benchmarks — scores in the 30–35% range, time-sensitive, cited here as evidence of replacement difficulty rather than as fixed estimates. These observations suggest that replacement construction has not kept pace with saturation pressure. Taken together, they satisfy marker 1 in qualitative form.
Wu et al. (2025) refine the picture further: when reinforcement learning is evaluated on benchmarks the base model has seen, even random or incorrect reward signals appear to improve performance — an artifact that disappears on contamination-free variants. The differential between contaminated and clean benchmarks supplies a lower bound on the contamination’s magnitude. The contamination extends beyond training-data leakage in the section 3.2 sense to the broader phenomenon of benchmarks shaping the systems they evaluate.
6.5 The Benchmark–Deployment Gap as Marker 3 Evidence
The AIID record (AI Incident Database, 2026) gives marker 3 its most direct qualitative form. Strong benchmark performance has not eliminated serious deployment failures outside the benchmarks’ task envelopes.
The quantitative form is harder to measure, because the AIID does not natively provide the paired data the marker 3 protocol requires. The qualitative pattern holds: the same models that produce strong benchmark performance produce documented incident reports, and the available record over the 2024–2026 period does not, by itself, establish that the gap is closing. Marker 3 is visible qualitatively, although not yet quantified.
6.6 What the Markers Detect and What They Miss
The evidence assembled under markers 1, 2, and 3 in 2024–2026 — markers 1 and 3 in qualitative form, marker 2 as a testable rather than confirmed condition — is consistent with reflexive contamination as defined in section 2. The pattern appears across vendors, evaluation methods, and model families in the cited evidence. On that reading, the contamination operates at the level of the shared categorical evaluation infrastructure, not only as a defect of any single benchmark.
What the markers miss is the trajectory by which the joint pattern assembled. The protocols developed in section 5 produce static measurements at a time point; they do not, in themselves, distinguish reflexive contamination from alternative dynamics that could have produced the same static surface. Such alternatives include genuine task saturation followed by genuine convergence on contracted but accurate τ, with the divergence in marker 3 attributable to evaluation-deployment gap rather than capacity asymmetry. Distinguishing the alternatives requires longitudinal data on the order in which the marker components rose — saturation followed by divergence, or divergence followed by saturation — and on the causal direction connecting them. The longitudinal analysis belongs to the fifth note in this series.
A second feature the markers miss is the boundary case in which the auditee’s design renders marker 3 inoperable. Where downstream consequence is structurally unobservable — as in the cryptographic-supply case treated in the supplementary cryptographic-supply note (Ortulanu, 2026h) — the divergence between evaluation results and downstream consequence cannot be measured at all. The marker is silent in the limit, even when reflexive contamination is present. The trajectory the markers cannot trace and the limit at which one of them cannot operate together set the agenda for section 7.
7 Limitations
The analysis has several limitations that bear explicit statement.
The mechanism named in section 2 is identified through structural correspondence between two cases rather than derived from a formal model. The correspondence is suggestive — the AAA case and the LLM-as-judge case share the three features section 4.3 identifies — but it does not, on its own, establish that no third mechanism could produce the same structural surface. A formal treatment would require an information-theoretic model of recursive use in categorical benchmarks that goes beyond the present argument. The first note’s analogous limitation regarding the signal-to-noise floor of self-referential evaluation was taken up in the second note’s structural development; a parallel formal treatment of reflexive contamination remains open.
The treatment of reflexive contamination as a static aggregate phenomenon is a methodological choice rather than a metaphysical claim. The mechanism operates through cycles; the note fixes the endpoint of those cycles rather than tracing the trajectory. The trajectory is the subject of the fifth note in the series. Where the trajectory itself matters for diagnosis — where the order in which markers assembled is the diagnostic information — the present tools apply only as the description of the limit. Cases in which the trajectory is recoverable from longitudinal data may admit sharper analysis than a static treatment supports.
The protocols developed in section 5 are substrate-internal. They operate on the audit cycle’s outputs and the auditee’s deployment record, without requiring access to information the cycle does not produce. This choice was made for the same reason the second note’s analysis was framed in terms of capacity within the audited substrate: external information that would resolve the diagnosis without protocol mediation is, by hypothesis, what the contamination has made unavailable. The cost of the choice is that the protocols cannot, alone, distinguish reflexive contamination from alternative dynamics that produce the same static signatures. The joint pattern across markers raises diagnostic confidence; it does not collapse the alternative explanations to a single posterior.
The Marker 3 protocol carries a stronger limitation. Where the auditee’s design suppresses observability of certain downstream consequences as a structural intent — as the Zcash Orchard shielded-pool design suppresses observability of supply integrity (Ortulanu, 2026h) — the divergence between evaluation results and downstream consequence cannot be measured at all. The marker is silent in such cases, even when reflexive contamination is present. The silence is a property of the auditee’s design, not a defect of the marker; cases of this kind belong to the supplementary cryptographic-supply note, where they are treated as forensic material for the audit-asymmetry frame more generally.
The note treats reflexive contamination in two domains — structured finance circa 2008 and artificial intelligence evaluation in 2024–2026. It does not establish that the mechanism is general to all domains in which categorical benchmarks are used recursively. Other domains — academic citation indices, standardized testing in educational selection, performance review systems in corporate governance, peer review in scientific publication — may exhibit the mechanism in recognizable form; the present treatment leaves them for later work. The sixth note in the series treats categorical integrity in non-AI domains as its primary object; the seventh treats cross-domain isomorphism as a synthesis problem. The restriction to two cases is a scope decision, not a claim about generality.
A final limitation concerns what the note does not attempt. The note does not propose a contamination-resistant benchmark. Whether such a design exists for AI evaluation — analogous to the way Bitcoin’s transparent ledger resists detection asymmetry in the cryptographic-supply domain (Ortulanu, 2026h) — is left open in section 8. The design problem remains open.
8 Open Questions
Five questions remain.
The Trajectory. The static treatment leaves the multi-cycle dynamics by which reflexive contamination proceeds for the fifth note in the series. That note will need to address the rate at which the discriminative function of a categorical benchmark erodes per cycle of recursive use, as a function of the use channel’s bandwidth and the auditee population’s size. Whether the rate is approximately linear, geometric, or exhibits regime transitions is currently unresolved. The 2024–2026 AI evaluation data may permit retrospective trajectory estimation if the saturation timings for MMLU, GPQA, and MMLU-Pro, together with the early trajectory of HLE, are placed on a common scale with the volume of recursive use each benchmark sustained.
The Limit Distribution. A categorical benchmark whose recursive use has reached steady state produces a population of auditees that is, in some sense, the fixed point of the use channel. What this fixed point looks like — what subclass of auditee behaviors survives the iterated selection — is unresolved. Li et al. (2025) document that the fixed point exhibits preference leakage; whether it admits further structural characterization is open. The question matters because diagnosis of how much of an evaluation population sits at the fixed point versus en route to it bears directly on the temporal scale on which the contamination is recoverable.
The Design Question. The hardest practical question is whether a categorical benchmark can be designed to resist reflexive contamination. The supplementary cryptographic-supply note treats one contrasting case in which public, parallel verification removes a related asymmetry (Ortulanu, 2026h); the point of the analogy is not cryptocurrency, but the existence of designs in which verification is not concentrated in the same substrate as the evaluated population. Candidate designs in the contemporary AI literature — rotating private tests, game-based evaluation, and evaluator-auditee separation — address parts of the problem. Whether any combination preserves discriminative force under recursive use remains unresolved. The structural form such a design must take, and the necessary conditions any candidate must satisfy, is the subject of the fourth note in this series.
Cross-Domain Isomorphism. The argument is developed from AI evaluation and structured finance, not yet from a broad cross-domain sample. Other categorical-evaluation domains — academic citation indexing, standardized educational testing, scientific peer review — are named only as candidates for later work, not as implied instances; the sixth note treats non-AI domains, and the seventh synthesizes cross-domain isomorphism. The structural correspondence between the AAA case and the LLM-as-judge case is the only direct support the analysis offers.
The Reflexive Position of the Argument. Because the argument concerns audit relations, its own review conditions matter. The framework remains open to ordinary argumentative criticism; what it adds is a structural reason to seek reviews from outside the AI safety substrate. Reviews from financial regulation, evaluation methodology, and sociology of knowledge would provide the strongest external tests.
No constructive alternative is offered here. The contribution is diagnostic: the note names a feedback mechanism, gives markers for detecting it, and tests those markers against two cases. The remaining problem is design — what kind of evaluation infrastructure could be used recursively without training systems into the very categories meant to audit them. That problem is left open.
References
Adrian, T., & Shin, H. S. (2010). Liquidity and leverage. Journal of Financial Intermediation, 19(3), 418–437.
AI Incident Database. (2026). AI Incident Database. Responsible AI Collaborative. incidentdatabase.ai. Accessed June 2026.
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073.
Blum, A., & Hardt, M. (2015). The ladder: A reliable leaderboard for machine learning competitions. Proceedings of the 32nd International Conference on Machine Learning, 1006–1014. arXiv:1502.04585.
Campbell, D. T. (1976). Assessing the impact of planned social change. Occasional Paper Series No. 8. Dartmouth College, Public Affairs Center.
Chen, S., Chen, Y., Li, Z., Jiang, Y., Wan, Z., He, Y., Ran, D., Gu, T., Li, H., Xie, T., & Ray, B. (2025). Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation. arXiv:2502.17521.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Goodeye Labs. (2025). 2025 Year in Review for LLM Evaluation: When the Scorecard Broke. Industry retrospective. goodeyelabs.com/insights/llm-evaluation-2025-review.
Haldar, R., & Hockenmaier, J. (2025). Rating Roulette: Self-inconsistency in LLM-as-a-judge frameworks. arXiv:2510.27106.
Han, S., Titericz Junior, G., Balough, T., & Zhou, W. (2025). Judge’s Verdict: A comprehensive analysis of LLM judge capability through human agreement. arXiv:2510.09738.
Li, D., Sun, R., Huang, Y., Zhong, M., Jiang, B., Han, J., Zhang, X., Wang, W., & Liu, H. (2025). Preference leakage: A contamination problem in LLM-as-a-judge. arXiv:2502.01534 (ICLR 2026 forthcoming).
LifeArchitect. (2026). Mapping IQ, MMLU, MMLU-Pro, GPQA, HLE. lifearchitect.ai/mapping. Accessed June 2026.
Manheim, D., & Garrabrant, S. (2018). Categorizing variants of Goodhart’s law. arXiv:1803.04585.
Morris, S., & Shin, H. S. (2002). Social value of public information. American Economic Review, 92(5), 1521–1534.
Ortulanu, G. (2026a). A category-level failure mode not captured by distribution metrics. osservadore.
Ortulanu, G. (2026b). The audit asymmetry problem: Capacity conditions for substrate-independent audit. osservadore.
Ortulanu, G. (2026h). The 2026 Zcash Orchard episode: An application of the audit asymmetry frame to a non-AI domain. osservadore (supplementary note, forthcoming).
Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. arXiv:2404.13076.
Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., & Hardt, M. (2020). Performative prediction. Proceedings of the 37th International Conference on Machine Learning, 7599–7609. arXiv:2002.06673.
Russo, D., Orlando, G. M., La Gatta, V., & Moscato, V. (2025). Automating AI failure tracking: Semantic association of reports in the AI Incident Database. arXiv:2507.23669.
Saito, K., Wachi, A., Wataoka, K., & Akimoto, Y. (2023). Verbosity bias in preference labeling by large language models. arXiv:2310.10076.
Senate Permanent Subcommittee on Investigations. (2011). Wall Street and the financial crisis: Anatomy of a financial collapse [Levin–Coburn Report]. United States Senate.
Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., & Vosoughi, S. (2025). Judging the judges: A systematic study of position bias in LLM-as-a-judge. Proceedings of AACL-IJCNLP 2025. arXiv:2406.07791.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759.
Skreta, V., & Veldkamp, L. (2009). Ratings shopping and asset complexity: A theory of ratings inflation. Journal of Monetary Economics, 56(5), 678–695.
Wu, M., Zhang, Z., Dong, Q., Xi, Z., Zhao, J., Jin, S., Fan, X., Zhou, Y., Lv, H., Zhang, M., Fu, Y., Liu, Q., Zhang, S., & Zhang, Q. (2025). Reasoning or memorization? Unreliable results of reinforcement learning due to data contamination. arXiv:2507.10532 (AAAI 2026 forthcoming).
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36. arXiv:2306.05685.