Recursive Evaluation Collapse
A minimal model for the multi-cycle dynamics of reflexive contamination
1 Introduction
Between 2019 and 2025, the public benchmarks of language-model evaluation saturated in sequence — each retired near its ceiling, each replacement costlier to build than the last. The scores rose the whole way. What the scores measured is the variable no leaderboard displays. This note gives that variable a clock.
The third note in this series fixed an endpoint and declined to trace the path to it. Reflexive contamination was treated there as a static aggregate phenomenon — the contaminated state at a time point, or the cumulative state earlier cycles produced — and every dynamical question was deferred (Ortulanu, 2026c §1, §6, §8). Three were deferred by name: the rate at which a categorical benchmark’s discriminative function erodes per cycle of recursive use, as a function of the use channel’s bandwidth and the auditee population’s size; the order in which the three diagnostic markers assemble, which a static protocol cannot supply; and the character of the limit the process approaches. The second note left two adjacent reservations: how the cascade among capacity dimensions unfolds when audit cycles repeat, and how shared substrate grows when the systems trained on an audit’s outputs become, in later cycles, the systems on which the audit operates (Ortulanu, 2026b §2, §3). A further debt is older. The third note conceded that a parallel formal treatment of reflexive contamination remained open (Ortulanu, 2026c §7). The present note pays these debts together, in the most economical form available.
The instrument is a minimal model: a population of auditee behaviors updated, cycle by cycle, partly as a function of the benchmark’s own judgments. Minimal is meant exactly. The model carries no architecture, no training objective, no claim about any particular pipeline. It is a bookkeeping device for the mechanism the third note named, and it plays the role Shannon, Rice, and Gödel played in the second note: marking the shape of a constraint rather than deriving it (Ortulanu, 2026b §1). What the bookkeeping buys is a clock. With a clock, questions that the static treatment could only pose become statements with content: which configurations erode slowly and which compound, what the limit state looks like, in what order the observable markers should arrive, and what evidence would refute the account.
The note proceeds from the model to its consequences. The erosion dynamic is stated and its assumptions are laid bare; three trajectory regimes are separated; the fixed point is characterized; the markers of the earlier notes are re-derived as trajectories with a predicted order of assembly; seven falsifiable predictions are extracted, with a sketch of the retrospective estimation that would bear on them; the limitations and open questions close.
Scope. The model is not fitted to data in this note, and no empirical estimation is executed here. The predictions are stated so that execution — a separate, companion line of work — has a fixed target. Where the note touches the 2024–2026 evaluation record, it does so through the third note’s hedged readings rather than fresh measurement.
Methodological commitment. A note that introduces a formal object owes the series an account of the object’s status. The model below is offered as notation for the third note’s mechanism, not as a theory from which the mechanism follows. Its assumptions are stated as assumptions; its regimes are qualitative; nothing in what follows depends on a closed form. The commitment of the earlier notes — constraint markers rather than derivations — is unchanged.
2 The Model
Let X be a space of auditee behaviors and C a finite judgment set, with B the categorical benchmark mapping behaviors to judgments — the third note’s object, carried over unchanged (Ortulanu, 2026c §2). Let Y* be the task-relevant property the benchmark is meant to track. A population at cycle t is a distribution Pₜ over X: the auditees the benchmark actually faces. One cycle of evaluation and production is written
Pₜ₊₁ = U(Pₜ, B, Zₜ)
where U is the update by which the next population is produced — training, selection, release, imitation — and Zₜ collects everything in the update that does not depend on the benchmark’s judgments: new entrants, exogenous data, drift in the world. The benchmark is used recursively when U depends materially on B’s judgments. Write β for an index of that dependence — the weight the benchmark’s verdicts carry in the update, whether as reward signal, selection gate, curation filter, or release decision. The third note’s use channel is the case β > 0; the second note’s audit-only relation is the case β = 0. The configuration in which deployed judgments re-enter the process they judge is the performative one, already placed as the third note’s closest formal neighbor (Perdomo et al., 2020; Ortulanu, 2026c §3); the model gives that configuration categorical bookkeeping.
Two families of quantities live over this process, and the distinction between them is the model’s entire point. The first is external: the discriminative content Dₜ — the degree to which the benchmark’s judgments separate the Y-relevant classes of the population it currently faces. In the information-theoretic idiom, Dₜ is the task-relevant information the judgments carry about Y over Pₜ; the idiom is used here as shape, in the second note’s manner, not as a quantity any section computes. The second family is internal: the score level Sₜ, the inter-judge agreement Aₜ, the intra-judge stability Rₜ, and the scope τₜ of the task actually being judged — the quantities an evaluation ecology can observe about itself, the ones the third note’s protocols operate on (Ortulanu, 2026c §5).
The model adds no mechanism. It gives the third note’s mechanism a clock.
3 The Erosion Dynamic
One cycle of the dynamic restates the third note’s mechanism with subscripts. When β > 0, the update selects for behaviors that map cleanly onto the categories the benchmark already distinguishes, and demotes behaviors its categories cannot accommodate. Within each category, the surviving population homogenizes along the judged dimensions; the distinctions the benchmark was built to draw are exactly the distinctions along which the population contracts. The scope τₜ of what is genuinely being discriminated narrows, and, other things equal, the use channel’s contribution pushes Dₜ₊₁ below Dₜ — not because the benchmark changed, but because the population it faces is increasingly its own product (Ortulanu, 2026c §2).
The erosion claim rests on three assumptions, stated as such. The update must weight the benchmark’s judgments (β > 0); the anchor must be inhabitable, in the fourth note’s sense — the reference the judgments answer to must be one the population can migrate toward (Ortulanu, 2026d §3); and the judge’s lineage must be close enough to the population’s that judge-compatible features are learnable at all. Where any assumption fails, the dynamic stalls: a benchmark whose verdicts nothing consumes does not contaminate, an anchor the population cannot inhabit does not contract, and a judge whose criteria the population cannot model is gamed only by accident. The fourth note’s three conditions reappear here as the parameters that keep Dₜ bounded away from collapse — the static partition acquiring its dynamical reading.
The bookkeeping can be made explicit without changing its status. Write the cycle’s effect on discriminative content as
Dₜ₊₁ = Dₜ − β · h · L · c(Pₜ, B) + ε(Zₜ)
where the expression is read as a schematic of expected change, clipped to D’s admissible range, not as a separable causal model; h is the anchor’s habitability, L the judge’s lineage relatedness to the population, c the contraction pressure the benchmark’s categories exert on the population it currently faces, and ε the restoration that exogenous influx supplies from outside the benchmark’s image. Nothing is estimated; the expression names the dependencies. No recursive use, no contraction through this channel. No inhabitable anchor, no migration toward the reference. No lineage kinship, no reliable learning of judge-compatible features. Sufficient influx, slower erosion. The product form carries only the zeros: absent any one factor, this channel contributes no contraction, and any monotone interaction with the same zeros would serve the bookkeeping equally well.
What makes the dynamic treacherous is the behavior of the internal family while Dₜ falls. Homogenization within categories makes judging easier, not harder: judges presented with a population that has converged on category-compatible features agree with one another more readily, so Aₜ holds or rises; scores climb toward whatever the categories reward, so Sₜ rises; and the contraction of τₜ is invisible from inside any single cycle, because the cases that would have tested the boundary are precisely the ones the update has demoted. From inside, improvement and erosion file the same report. The signature of the dynamic is therefore a divergence between the two families — internal signals steady or improving while discriminative content declines. Genuine convergence, the explanation the third note’s protocols must always rule out, predicts the two families moving together: a population actually mastering the task raises agreement and external discrimination at once. The divergence pattern, not any single quantity, is what separates the two accounts — a dynamical restatement of the third note’s joint-marker requirement (Ortulanu, 2026c §5).
4 Regimes of the Trajectory
The third note asked whether the erosion rate is approximately linear, geometric, or punctuated by regime transitions (Ortulanu, 2026c §8). Within the model the question resolves into three regimes, separated by the weight β and by what feeds the update besides the benchmark.
The buffered regime. Where β is small and the exogenous term Zₜ is large — a benchmark consulted but not optimized against, a population continually refreshed from sources the benchmark never judged — erosion proceeds slowly and approximately steadily. Each cycle’s selection pressure is diluted by influx the benchmark did not shape. Many benchmarks plausibly begin life here, which is part of why early-life validity guides late-life validity so poorly: the regime, not the instrument, is doing the work.
The self-feeding regime. The dynamic compounds when the evaluated population’s own outputs become the next cycle’s raw material — model-generated data in training corpora, judge-approved responses as preference data, released systems generating the text future systems learn from. The update then applies the benchmark’s selection twice over: once through the verdicts, once through the material the verdicts admitted. Erosion in this regime tends toward the geometric, each cycle starting from a population already concentrated by the last. The configuration is the one the recursive-training literature has begun to chart at the distribution level: models degrade when trained on their own output without sufficient fresh data (Shumailov et al., 2024; Alemohammad et al., 2023); stability returns when real data keeps an adequate share of the mix (Bertrand et al., 2023); and where synthetic data dominates, the degradation follows characterizable scaling laws (Dohmatob et al., 2024). That literature’s fresh-data condition is this model’s exogenous term, met or unmet. Finance has run the regime under its own name for decades: leverage set against measured risk feeds the prices that set measured risk, the procyclical loop of the endogenous-risk literature (Adrian & Shin, 2010) — the same self-feeding shape at the price layer, with optimizing response supplying the weight that recursive training supplies here. The analogy is structural, not substantive: the object fed back differs; the self-feeding form of the update is the same. The present claim is the third note’s, one layer up: the categorical structure can erode even where distributional diversity is maintained (Ortulanu, 2026c §3), and the categorical layer is where the judgments live.
The cliff. The third regime is entered rather than approached. As within-category signal thins, there comes a point at which the residual differences between candidates fall beneath what the judge can discriminate, and judgments collapse onto surface features — length, position, fluency, style. Call the threshold δJ, the judge’s discrimination floor; it is a property not of the task alone but of the judge, the prompt, and the presentation channel together. The static witness for this floor is already in the record: position bias in LLM judges grows as the quality gap between compared solutions narrows (Shi et al., 2025), and intra-judge stability is low precisely where inter-judge agreement looks respectable (Haldar & Hockenmaier, 2025). In the model’s terms, when discriminable signal falls below the judge’s floor, τₜ does not contract further — it caves, and the benchmark’s verdicts become functions of presentation. Saturation, read from inside, is the cliff’s internal face: scores cluster at the ceiling because the instrument has run out of population to separate.
The second note’s reserved question — how the cascade among capacity dimensions unfolds in repeated cycles — has its answer in the coupling of these regimes (Ortulanu, 2026b §2). An anchor failure (the population inhabiting the reference) narrows semantic reach; narrowed reach lowers the bandwidth the task appears to require, so the rate deficit stops registering; the apparent adequacy invites heavier use, raising β; heavier use accelerates the contraction that caused the failure. Each dimension’s deficit recruits the others. The interval between an inversion in one dimension and its propagation to the rest compresses — the dynamical content of the second note’s closing observation (Ortulanu, 2026b §9).
5 The Limit Case
A process of the form Pₜ₊₁ = U(Pₜ, B, Zₜ) with stationary β and Z has limiting behavior to characterize. The third note asked what the limit population looks like (Ortulanu, 2026c §8); the model permits a characterization without a closed form — the fixed point of the use channel, in the third note’s phrase, read here as a limit case rather than a stationarity claim.
At the limit, the population is concentrated on the region of judge-compatible features: behaviors shaped to the categories, styled to the judge, distributed as the update rewards. Over such a population the benchmark does not fall silent. Judgments continue to issue; agreement among judges drawn from the shared lineage is high, because the population was selected for their shared criteria; scores sit wherever the category structure pays. What the judgments report, however, is no longer the task property Y* — within the manifold, Y*-relevant variation has been selected out or rendered invisible — but proximity to the family that produced the judge. The third note’s formulation was static: a judge whose evaluations describe its own contribution to the auditee’s training, presented as evaluations of the auditee’s intrinsic properties (Ortulanu, 2026c §4). The limit is that formulation made permanent.
The early trajectory toward this point has already been measured once. Preference leakage — judge bias toward systems trained on material the judge’s lineage produced, increasing with generator-judge relatedness (Li et al., 2025) — is, in the model’s terms, a reading of Dₜ already declining along the β gradient: relatedness is the recurrence’s L, read in configurations where the judge’s criteria had entered the update. The measured effect is the trajectory’s early segment, taken while the configuration is young enough for the gradient to be visible.
Recovery, at or near the limit, runs only through terms outside the benchmark’s image. Exogenous influx the benchmark never judged restores variation the categories did not select; an anchor rotated to a reference the population cannot inhabit restores a direction of discrimination the manifold does not span; a lineage-disjoint judge restores criteria the population was not shaped toward. These are the fourth note’s three conditions, returning as the only levers the model leaves — the design partition and the dynamics naming the same structure from two sides (Ortulanu, 2026d §3, §6).
The limit is not silence. The benchmark goes on issuing judgments; what they report there is membership.
6 The Markers as Trajectories
The second note proposed three markers of audit detachment and the third operationalized them statically, conceding that a static protocol cannot distinguish reflexive contamination from alternative dynamics that produce the same surface at a time point — the discrimination requires the order in which the markers assembled (Ortulanu, 2026b §7; 2026c §5, §6). The model supplies the order.
Marker 1 as trajectory. Score saturation with rising verification cost is, in model terms, Sₜ approaching ceiling while the cost Vₜ of constructing items that still discriminate rises. Vₜ rises because discriminating items must sample behavior off the learned region, and that region tightens each cycle: item-writers are searching a shrinking complement. The model adds timing — Vₜ begins rising while Sₜ is still mid-range, because the manifold contracts from the first cycle, not from saturation onward.
Marker 2 as trajectory. Agreement on a contracting task is Aₜ holding or rising while τₜ contracts and Rₜ fails to improve. The scope components the third note’s protocol named — rubric coverage, category-usage entropy, exclusion rate, task-family diversity (Ortulanu, 2026c §5) — are the observable faces of τₜ. The model’s reading of the record’s oddest pair — respectable inter-judge agreement beside poor intra-judge stability (Haldar & Hockenmaier, 2025) — is that agreement of this kind measures the homogeneity of the judged population, not the stability of the judging: Aₜ can be inherited from the manifold while Rₜ stays noisy. In the regime, the judges agree ever more readily about ever less.
Marker 3 as trajectory. Divergence between evaluation results and downstream consequence is the external correlation falling as Dₜ falls — the slowest marker, because consequences accumulate at deployment timescales, not evaluation timescales.
The predicted order of assembly follows from the regimes: τₜ contraction begins first and silently; Aₜ firms as homogenization proceeds; Vₜ rises next, while scores are still informative in the shrinking scope; Sₜ saturation arrives as the cliff’s internal face; the marker-3 divergence trails the whole sequence at the lag of consequence. Contamination therefore predicts scope contraction before saturation, with verification cost rising through the interval. The genuine-progress account predicts the reverse signature — saturation arriving on a stable or expanding scope, verification cost flat or falling, external correlation preserved. The two accounts, indistinguishable in a snapshot, separate in time series. That separation is what the third note’s protocols were built to await, and what the order supplies (Ortulanu, 2026c §5, §6).
7 Falsifiable Predictions, with an Estimation Sketch
Seven predictions follow from the model. Each is stated with the observation that would refute it.
Prediction 1 — use-intensity gradient. Among benchmarks of comparable age and domain, those whose judgments carry more weight in training, selection, and release (β higher) lose frontier-ranking discrimination faster. Refuted if discrimination loss is uncorrelated with use intensity once ceiling proximity is controlled.
Prediction 2 — relatedness gradient. Judge bias toward evaluated systems increases, on average, with generator-judge relatedness, the configuration preference leakage measured at one point (Li et al., 2025). Refuted if the gradient flattens when item-level leakage is removed and capability is controlled.
Prediction 3 — leakage-free persistence. Erosion persists under full item rotation. A rotating benchmark whose verdicts still feed training and selection contaminates through the verdicts; the fourth note’s matrix row states the design-level reason (Ortulanu, 2026d §5). Refuted if rotation alone, with the use channel intact, arrests the markers.
Prediction 4 — entropy before saturation. The entropy of judgment categories actually used declines before scores saturate: nominal categories persist in rubrics while fewer do discriminative work. Refuted if category-usage entropy holds until ceiling.
Prediction 5 — the scissors. As evaluated populations converge in measured quality, internal agreement holds or rises while agreement with task-external reference (oracle, ground truth, deployment outcome) falls. Refuted if internal and external agreement move together through convergence — the genuine-convergence signature.
Prediction 6 — replacement cost growth. Per-item construction cost of still-discriminating evaluations rises with cumulative recursive use of the domain’s benchmarks, not merely with capability. Refuted if replacement cost tracks capability alone across domains of differing use intensity.
Prediction 7 — anchor immunity. The negative control. Where recursive use is heavy but the anchor is uninhabitable, the markers should fail to assemble in the predicted order: internal improvement should not decouple from external discrimination. The account is weakened wherever the full sequence appears despite an anchor the population cannot inhabit.
The quantities the predictions range over admit working proxies; each carries its chief confounder, listed beside it.
The quantities, operationalized
| Quantity | Working proxies | Chief confounder |
|---|---|---|
| β (use weight) | presence in training and preference pipelines; role in release decisions; leaderboard prominence | private pipelines |
| Dₜ (discriminative content) | correlation with a task-external oracle; frontier-rank stability across protocol variants | genuine convergence; drift |
| τₜ (scope) | category-usage entropy; rubric coverage; exclusion rate; task-family diversity | rubric revision |
| Aₜ, Rₜ (agreement, stability) | inter-judge agreement; same-input rerun stability under order and paraphrase | population homogenization; decoding noise |
| Vₜ (replacement cost) | expert hours per discriminating item; item discard rate | capability growth; tooling |
| external divergence | benchmark-deployment correlation; oracle gap; audit-test gap | visibility and scrutiny; reporting bias |
One discipline governs every test: the proxies, and the assignment of a case to the model’s elements — whether the channel was material, whether the anchor was inhabitable — are to be fixed before outcomes are known. Assigned in retrospect, the elements stop being scope conditions and become exits for the framework rather than for the domain.
The estimation these predictions invite is retrospective, and its sketch is brief because its difficulty is data, not design. Place the saturation timings of the retired public benchmarks and the early trajectory of their replacements on a common scale, as the third note proposed (Ortulanu, 2026c §8). Attach to each benchmark a use-intensity index — presence in training and preference pipelines, weight in release decisions, leaderboard prominence — and to each interval the scope proxies of marker 2. Then ask whether erosion rate orders with use intensity once capability growth, measured on oracle domains, is removed. Every term of the sketch is observable in principle; several are observable only with provider cooperation; none is computed here. The companion empirical track holds the executable version, and the note’s claims are calibrated to survive its failure: the model would be unfitted, not refuted, by missing data — refutation runs through the seven predictions.
8 Limitations
The model is a bookkeeping device, and its limitations are those of its genre.
Nothing is fitted. The regimes are qualitative, the fixed point is characterized rather than computed, and no closed form is claimed for any trajectory. The note’s contribution is the partition of dynamical cases and the order of marker assembly, not magnitudes.
The dynamic assumes a fixed benchmark procedure, as the third note’s counterfactual did (Ortulanu, 2026c §2). Real evaluation ecologies revise rubrics, retire instruments, and rotate items; revision resets τₜ accounting and can mask or mimic recovery. The model treats procedure change as outside U, which is a simplification the estimation sketch must confront — rubric history is both a scope proxy and a confounder.
The weight β is not directly observable. Training-pipeline composition, preference-data provenance, and release-decision criteria are mostly private; the use-intensity index of the estimation sketch is a public-signal proxy for a private quantity, and predictions stated over β are testable only through it.
The single-benchmark frame understates coupling. Benchmarks sharing substrate — common corpora, common judge lineages, common item ecologies — erode jointly, and the second note’s reserved question about substrate growth across cycles (Ortulanu, 2026b §3) is answered here only in the small: the model treats one instrument facing one population, where the ecology-level process couples many instruments to one population. The coupling plausibly strengthens the dynamic, which would make the simplification conservative in direction — an assumption, stated as one, not a result.
The empirical record invoked is the third note’s, with its hedges carried, not re-derived: tracker-supported timings, qualitative saturation readings, time-fixed figures (Ortulanu, 2026c §6).
9 Open Questions
Five questions remain.
Measuring the Channel. The predictions are stated over β; the estimation sketch reaches it only by proxy. What public observables best track the weight of benchmark judgments in production updates — and whether provider disclosure could make the channel directly measurable — is open, and is the single largest obstacle between the model and its data.
The Ecology. How coupled instruments erode together — whether shared substrate synchronizes their trajectories, and whether one instrument’s cliff hastens its neighbors’ — extends the model in the direction the second note’s substrate-growth question points (Ortulanu, 2026b §3). The single-instrument treatment here is the base case of that extension.
Recovery Dynamics. The fourth note’s conditions reappeared in section 5 as recovery levers, statically. How much exogenous influx, or how frequent an anchor rotation, holds Dₜ above a working floor — maintenance schedules for evaluation infrastructure, in effect — is the constructive question the model makes well-posed and does not answer.
Execution. The seven predictions are stated for testing; the companion empirical track holds the design. Which prediction falls first to available data — the scissors, under oracle-anchored domains, is the present best candidate — will shape how much of the model survives.
Domain Generality. Every quantity in the model is domain-neutral: nothing in P, B, β, or the regimes is specific to AI evaluation. Whether the non-AI categorical benchmarks the third note named as candidates exhibit the same regimes and the same order of marker assembly is the sixth note’s question, and the synthesis of what the answer implies is the seventh’s.
The note closes where its instrument points. The dynamics of reflexive contamination admit a clock, three regimes, a characterized limit, and an order of assembly that makes the static protocols of the earlier notes longitudinal. What they do not yet have is a fitted curve. That absence is stated, not hidden; the predictions are the form in which it can be paid. And if the model is right, the leaderboards of 2026 are already partly measuring the echo of 2024; the fitting would say how much.
References
Adrian, T., & Shin, H. S. (2010). Liquidity and leverage. Journal of Financial Intermediation, 19(3), 418–437.
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. G. (2023). Self-consuming generative models go MAD. arXiv:2307.01850.
Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M., & Gidel, G. (2023). On the stability of iterative retraining of generative models on their own data. arXiv:2310.00429.
Dohmatob, E., Feng, Y., Yang, P., Charton, F., & Kempe, J. (2024). A tale of tails: Model collapse as a change of scaling laws. Proceedings of the 41st International Conference on Machine Learning. arXiv:2402.07043.
Haldar, R., & Hockenmaier, J. (2025). Rating Roulette: Self-inconsistency in LLM-as-a-judge frameworks. Findings of the Association for Computational Linguistics: EMNLP 2025, 24986–25004. doi:10.18653/v1/2025.findings-emnlp.1361.
Li, D., Sun, R., Huang, Y., Zhong, M., Jiang, B., Han, J., Zhang, X., Wang, W., & Liu, H. (2025). Preference leakage: A contamination problem in LLM-as-a-judge. arXiv:2502.01534 (ICLR 2026 forthcoming).
Ortulanu, G. (2026b). The audit asymmetry problem: Capacity conditions for substrate-independent audit. osservadore.
Ortulanu, G. (2026c). The reflexive contamination of categorical benchmarks. osservadore.
Ortulanu, G. (2026d). Substrate independence as a structural requirement. osservadore.
Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., & Hardt, M. (2020). Performative prediction. Proceedings of the 37th International Conference on Machine Learning, 7599–7609. arXiv:2002.06673.
Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., & Vosoughi, S. (2025). Judging the judges: A systematic study of position bias in LLM-as-a-judge. Proceedings of IJCNLP-AACL 2025, 292–314. doi:10.18653/v1/2025.ijcnlp-long.18.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759.