Substrate Independence as a Structural Requirement
A capacity profile for contamination-resistant evaluation
1 Introduction
Through 2024 and 2025, two national institutes built to evaluate frontier models ran on new mandates, public methods, and benchmarks descended from the ecology they were built to audit. The offices were independent. The instruments were not. That distinction — between who holds the judgment and what the judgment is made from — is this note’s subject.
Three reservations made earlier in this series come due together. The first note in the series proposed three externality conditions an audit instrument must satisfy to block category-level failure — provenance independence, external anchoring, and speed parity — and read four current safety approaches as trade-offs among them (Ortulanu, 2026a §5, §6). The second note relocated those conditions inside a capacity relation between auditor and auditee, and reserved two derivations for later work: the derivation of substrate independence as the necessary form of the audit instrument when self-reference depth requirements cannot be met by governance alone (Ortulanu, 2026b §8), and the conditions under which sampling and ensemble methods can substitute for bandwidth parity (Ortulanu, 2026b §9). The third note named the mechanism that makes the question urgent — reflexive contamination, the erosion of a categorical benchmark’s discriminative function through the recursive use of its judgments in training and selecting the systems it evaluates — and closed by asking what structural form a contamination-resistant design must take, and what necessary conditions any candidate must satisfy (Ortulanu, 2026c §8). The present note takes up the three reservations as one problem.
The argument distinguishes two senses in which an audit instrument can be independent of what it audits. Independence by governance concerns who forms the judgment: ownership, staffing, funding, mandate, and the incentive relations that determine what happens to the judge when the judgment displeases. Independence by substrate concerns what the judgment is formed from: the lineage of the representational primitives, training history, and decision procedure out of which the criterion is built. The two senses come apart precisely where the third note’s mechanism operates — under recursive use, a governance-independent judge built from the auditee’s lineage converges on family-level consistency rather than independent judgment. From the second note’s capacity dimensions, held against the use channel, the note derives three design-level necessary conditions — a lineage condition, an anchor condition, and a rate condition — and reads seven candidate families across eight matrix rows. Rotation is split by verdict type because fresh items and an external correctness oracle purchase different things.
The result is a design grammar rather than a favored architecture. It shows what each candidate purchases, what it leaves exposed, and which combinations can dominate their components.
The two senses of independence separate first. The three conditions then form a capacity profile, governance is located inside that profile, and seven candidate families are compared across eight rows before the combination problem closes the analysis.
Scope. The profile records what a design makes available at a time point. The fifth note supplies the trajectory generated when a dimension remains exposed.
Methodological commitment. The conditions inherit the second note’s capacity dimensions and are held against the third note’s use channel. Their formalization makes comparisons explicit without turning the dimensions into a scalar score.
2 Two Senses of Independence
An audit instrument can be independent of its auditee in two ways that current discussion runs together. The first is positional. An evaluator is governance-independent when the relations that hold it in place — ownership, staffing, funding, statutory mandate, the consequences it faces when its judgments displease — are separated from the auditee’s. This is the independence that regulation can confer, that institutional design can engineer, and that the frontier-regulation literature treats as the central design variable (Anderljung et al., 2023). The second is constitutive. An evaluator is substrate-independent when the material its judgments are formed from — the representational primitives, the training history, the decision procedure, in the second note’s usage the substrate (Ortulanu, 2026b §3) — descends from sources disjoint from those that produced the auditee.
The first note marked the gap between the two in passing: legal independence is not distributional independence, and two organizations that are legally distinct can draw from the same pre-training substrate and converge on overlapping judgments (Ortulanu, 2026a §5). An independent office can still operate a dependent instrument. What the present note adds is that the gap becomes load-bearing under recursive use. A governance-independent judge whose lineage is shared with the auditee population does not merely risk correlated error. Under the use channel, its judgments participate in shaping the population toward the judge’s own categories, and the agreement that results measures consistency within a shared family rather than independent accuracy. The preference-leakage results catalogued in the third note measure a close instance of this configuration: bias of a judge toward systems trained on material the judge’s own lineage produced, across generator-judge identity, inheritance, and shared model family (Li et al., 2026; Ortulanu, 2026c §6).
Four adjacent repairs should be separated from the question this note asks. Regulatory design fixes who may audit and on what mandate; it operates on position. Cryptographic attestation fixes the integrity of the record — what was computed, by what artifact; it operates on provenance of the computation, not of the judgment (Ortulanu, 2026a §6). Benchmark rotation fixes the freshness of the items — what is asked; it operates on exposure. Rubric extension fixes what the description language can express; it operates on reach (Ortulanu, 2026b §5). Each buys something real. None of them, as such, touches the question of what the judgment is made from and where its reference stands — the two components on which, the next section argues, resistance to reflexive contamination structurally depends. Section 5 therefore separates rotation from verdict type: rotation changes exposure, while an executable oracle can independently anchor correctness without anchoring task coverage.
Governance decides who holds the instrument. Substrate is what the instrument is made of.
3 The Three Conditions
The second note specified three dimensions of an audit relation’s capacity — bandwidth, semantic reach, self-reference depth — each held relative to the requirement the audit task induces from the auditee (Ortulanu, 2026b §2). The third note specified the channel through which recursive use erodes a categorical benchmark’s discriminative function even when capacity is initially adequate (Ortulanu, 2026c §2). Holding each capacity dimension against that channel yields a condition on design. The derivation is deliberately spare: a design resists use-induced erosion only if the use channel cannot act on the sources, the reference, or the observation tempo that sustain the capacity dimensions.
The lineage condition. The process that forms the criterion — the judge’s training, the rubric’s construction, the test’s generation — must draw, at the resolution the audit task requires, on sources sufficiently disjoint from the auditee population’s lineage. The condition descends from self-reference depth: where criterion and auditee share substrate, the criterion is co-produced with the objects it judges, and the audit returns to self-evaluation regardless of who administers it (Ortulanu, 2026b §3). Disjointness runs from artifact identity through fine-tuning inheritance and family lineage to shared pretraining ecology, shared human-domain substrate, and physically separate measurement channels. The condition asks for task-relevant distance, not maximal distance. Its sufficient resolution remains an empirical question (Ortulanu, 2026a §9).
The anchor condition. The reference must lie beyond the reach of the use channel, and the evaluator must retain task-relevant access to it. Both clauses bind. A population cannot inhabit a machine-checked proof or a physical outcome, but an auditor unable to observe the proof or outcome has no anchor in practice. Conversely, a perfectly visible preference model remains inhabitable by construction because the population trains directly toward it (Bai et al., 2022). The condition therefore joins non-migratability to verification access. A learned preference is visible and inhabitable; a game outcome is visible while the game class remains inhabitable; a physical measurement can be both visible and beyond the population’s power to move. Semantic reach requires a reference outside the population and a channel by which the judge can still reach it.
The rate condition. The observational apparatus must operate at the tempo of the auditee — the bandwidth requirement, unchanged from the second note, with its distinguishing property intact: of the three, it is the condition procedural repair cannot supply by itself, because the task-relevant capture rate is set by the apparatus, not by the rules under which the apparatus operates (Ortulanu, 2026b §4, §9). The condition concerns the rate of task-relevant judgment against the auditee’s rate of change; the latency of record verification is a different quantity, and the distinction keeps both attestation rows from purchasing the rate coordinate.
The three conditions follow the three capacity deficits. Lineage addresses self-reference depth; anchor plus verification access addresses semantic reach; rate addresses bandwidth. Removing one reopens a distinct failure: a fast external judge can still share the auditee’s blind spots, a lineage-disjoint judge can still lack a task reference, and an anchored independent judge can arrive after the population has changed. Within the inherited apparatus the partition is operationally irreducible, though a formal proof of minimality remains open.
3.1 The capacity profile
Let a design d have the profile
K(d) = (ℓ(d), a(d), r(d))
where ℓ records task-relevant lineage distance, a records the joint availability of a non-migratable anchor and verification access, and r records observation tempo relative to auditee change. The coordinates can begin as ordered grades rather than cardinal measures. Lineage inheritance, anchor manipulability, evidence accessibility, and the ratio of audit latency to deployment latency supply working observables.
Profiles admit a partial order. Design d₁ dominates d₂ when every coordinate of K(d₁) is at least as strong and one is stronger. Designs that trade lineage for rate remain incomparable. The matrix in section 5 is therefore a Pareto map, not a leaderboard. A scalar score would hide the bottleneck the conditions were derived to expose. Only three grades enter that order: Unmet < Partial < Met. Conditional marks a coordinate whose grade cannot be assigned until the paired mechanism or operating regime is fixed. Orthogonal and Enabler are role labels rather than grades: an orthogonal component buys no coordinate, while an enabler can verify a capacity another component constitutes. Profiled alone, both remain Unmet on that coordinate.
Composition is the constructive edge. For a composite d₁ + d₂, the profile is not automatically the coordinate-wise maximum: a physical anchor adds nothing if the judge cannot access its evidence, and lineage attestation adds little if the verified judge cannot keep pace. Components compose only where the evidence channel connecting them preserves the purchased capacity. The design rule follows: close enough to read what the population produces, constituted far enough away to judge it, and fast enough to retain the difference.
4 Where Governance Stops
The second note reserved the relation between governance and substrate independence (Ortulanu, 2026b §8). The relation is neither equivalence nor irrelevance. Governance can change a substrate when it changes the production of judgment; organizational separation alone cannot.
Governance assigns custody of the criterion, sets the incentives of its custodian, and determines the consequences of its application. It reaches the substrate when those powers secure independent evidence, procure disjoint tools, preserve data provenance, or grant access the auditee cannot condition. It stops at the institution’s boundary when the judge, items, and rubric still descend from the auditee’s lineage. The self-reference depth requirement constrains the production relation: whether the standpoint the task requires is made from material outside the process being judged (Ortulanu, 2026b §3). A separate office can operate a dependent instrument; an internal apparatus can contain an external measurement channel. Independence belongs to how the judgment is produced.
Sampling, statistical inference, and hierarchical audit can substitute for rate parity when failure modes are sparse, stable, and specified in advance (Ortulanu, 2026b §9). They lose that substitution when the failure class is what the audit must learn, because the sampler’s frame then carries the blind spot at issue. Frontier audit contains both cases. Known misuse patterns admit sampling; emergent task-relevant behavior does not inherit the same guarantee. Ensembles drawn from one lineage multiply looks while preserving correlated blindness (Ortulanu, 2026a §5, §7).
The institutional best case makes the partition concrete. The UK AI Security Institute and the US Center for AI Standards and Innovation hold formal governance independence — separated staffing, public mandate, formalized evaluation practice (UK AI Safety Institute, 2024; NIST, 2024; U.S. Department of Commerce, 2025). Read against the three conditions: the lineage condition is met in part (public-sector staffing and methods, while the evaluation tooling and test corpora still descend from the shared research substrate), the anchor condition is met in part (mandate-fixed criteria, while the benchmarks that operationalize them are drawn from the ecology the auditees train on). In the public 2024–2025 snapshot, the record did not establish rate parity: access remained partly cooperative and disclosed evaluations trailed frontier deployment. Activity expanded in 2026 — CAISI published evaluation guidance and conducted model assessments, while UK AISI reported a wider evaluation program (NIST, 2026; AI Security Institute, 2026) — without yet supplying a public comparison between audit latency and deployment latency. The rate coordinate is therefore unverified, not institutionally impossible. The snapshot illustrates the profile and its time index at once.
Reassigning the judge changes custody. Independence changes only when the production of judgment changes with it.
5 The Candidate Matrix
Seven candidate families recur in current discussion of contamination-resistant evaluation: the three the third note named — rotating private tests, game-based evaluation, evaluator-auditee separation (Ortulanu, 2026c §8) — and the four the first note assembled from adjacent domains — zero-knowledge attestation, hardware attestation, physical-feedback loops, and certified human-origin data (Ortulanu, 2026a §6, §9). Rotation requires two rows because an executable verdict and a preference verdict have different anchors. Read against the three conditions, each candidate purchases some and sacrifices others. The reading, not the ranking, is the point: the matrix locates each design’s contribution and its residue.
The matrix in summary
Ordered grades: Met / Partial / Unmet. Conditional = implementation-dependent. Orthogonal and Enabler describe roles, not grades; alone they purchase no coordinate. Judgments are task-dimension specific
| Candidate | Lineage | Anchor | Rate |
|---|---|---|---|
| Rotating tests, oracle-scored (Jain et al. 2024) | Partial — fresh items and executable judge; broader task ecology shared | Partial — correctness anchored; task coverage inhabitable | Conditional — refresh cadence is not deployment-tempo observation |
| Rotating tests, preference-scored | Partial — items fresh, judge lineage unchanged | Unmet — learned preference remains inhabitable | Conditional — item refresh, not deployment-tempo observation |
| Game-based evaluation | Partial — rules and state judge are external; scenario construction may share lineage | Partial — outcome is a world-fact, game class inhabitable | Conditional — where play runs at auditee tempo |
| Evaluator-auditee separation | Partial — organizational resolution only | Unmet | Conditional — rate depends on the separated apparatus, not the separation |
| Zero-knowledge attestation (Kang et al. 2022) | Orthogonal — record, not judgment | Orthogonal | Orthogonal |
| Hardware attestation | Enabler — verifies lineage claims | Orthogonal | Orthogonal — record verification, not behavioral observation |
| Physical-feedback loops | Met | Conditional — Met with authenticated access; Partial where the interface is gameable | Conditional — Met in high-frequency domains; Unmet elsewhere |
| Provenance-certified data channel (C2PA 2026) | Partial — source claims, not authorship | Unmet — provenance does not establish human authorship | Conditional — where authenticated capture and verification keep pace with data production |
Rotation is the most informative family, because it separates three properties current practice treats as one: item freshness, verdict type, and task coverage. A rotating private benchmark — live problems time-stamped past the training cutoff, in the documented instance (Jain et al., 2024) — defeats item-level leakage: the auditee has not seen the test. Where solutions are scored by execution against hidden test cases, the verdict also supplies an external oracle for correctness. Recursive use can then improve the property the oracle actually checks, the fifth note’s alignment counterexample rather than automatic erosion (Ortulanu, 2026e §3). The remaining exposure moves outward. Repeated optimization can narrow the population toward contest-compatible problem types, interfaces, and solution styles while correctness within that scope improves; the oracle anchors the verdict, not the breadth τ of the task presented to it. Preference-scored rotation lacks even that protection: fresh items leave the learned judgment and its use channel intact. Rotation is therefore leakage-resistant by construction, contamination-resistant only on the dimensions its verdict anchors.
Game-based evaluation comes closest, among benchmark-shaped designs, to the anchor condition. A win is adjudicated by rules and game state, not by a learned preference — a reference that does not adjust toward the population’s style. Public game-based leaderboards now run head-to-head ladders for frontier models (Kaggle, 2026). The opponent pool determines the population against which comparative strength is estimated; it is not the lineage of the verdict. Lineage exposure instead enters through scenario construction, interface, and any learned adjudication layered over the rules. The population can still migrate into the game class itself, overfitting the anchor’s domain even where it cannot inhabit the anchor’s verdict. Evaluator-auditee separation, the third named candidate, is the governance instrument of section 4 in design clothing: it buys lineage disjointness at organizational resolution only, and the judge architecture it separates (Zheng et al., 2023) remains constituted from the shared substrate. Practice instances run at scale: third-party evaluators maintain leaderboards built on expert-driven evaluations with prompt-exposure controls (Scale AI, 2026); the separation they buy is organizational, and the judging substrate is unchanged. Nor does separation supply tempo: where a separated evaluator keeps deployment pace, the pace comes from its apparatus, not from the separation.
The cryptographic candidates divide along the record-judgment line the first note drew (Ortulanu, 2026a §6). Zero-knowledge attestation of model inference (Kang et al., 2022) makes the record incorruptible without touching what renders the judgment; it is orthogonal to all three conditions rather than a partial purchase of any. The proofs are perfect answers to a question contamination does not ask. Hardware attestation stands one step closer: it can verify claims about computational provenance — which artifact, which weights, which pipeline — and so serves as an enabler for the lineage condition, a way to check disjointness assertions; it does not itself constitute disjoint lineage (Ortulanu, 2026a §9). Even when attestation runs at low latency, that speed verifies the record; it does not observe task-relevant behavior at auditee tempo, so the rate coordinate remains orthogonal.
The two remaining candidates are the ones the first note sketched as partial external channels (Ortulanu, 2026a §9), and they are the only rows whose proposed reference is external in kind — the only candidates that can satisfy the anchor’s non-migratability clause in principle. Physical-feedback loops — model outputs validated against measured outcomes in the world — satisfy lineage because the measurement apparatus does not descend from the model ecology. They satisfy the anchor only where authenticated outcomes remain available beyond a gameable measurement interface, and fail on rate everywhere the measured process is slower than deployment, which is most domains. The first note’s qualification stands: the channel works where high-frequency physical measurement is available at low latency — robotics telemetry, online systems whose outcomes arrive with the action. The first note’s certified human-origin channel (Ortulanu, 2026a §9) meets the lineage condition for the training-data channel in principle; its nearest deployed instrument — content-provenance credentials of the C2PA kind (C2PA, 2026) — certifies claims about the source and history of material, not human authorship itself. It can enable an external data channel only when paired with separate authorship attestation. Its rate contribution is conditional on authenticated capture and verification keeping pace with data production. Its contribution to constructing an anchored system is therefore indirect at best; provenance alone does not purchase the anchor coordinate, and the adoption and privacy costs noted in the first note stand.
One episode outside the AI domain bears on the lineage condition, read at the right weight. The Zcash Orchard proof-soundness defect of 2026 sat unobserved for four years inside the system’s resident review ecology and surfaced under an AI-assisted audit workflow brought in from outside that cycle — not a fully lineage-disjoint probe, since the auditor knew the system and the assisting model shares the broader substrate, but a review constituted differently enough to carry different blind spots; a supplementary cryptographic-supply note in preparation takes up the episode at length (Ortulanu, 2026h). The illustration claims no more than this: review assembled outside the resident lineage saw what the resident lineage had learned not to.
No row meets all three conditions unconditionally at general frontier scale. At that scale, every candidate purchases at most two, and the profiles cluster — benchmark-shaped designs offer tempo compatibility in principle but sacrifice anchor; physical channels buy anchor and usually miss deployment tempo. The matrix at design level reproduces the trade-off profile the first note found at the institutional level (Ortulanu, 2026a §6), which is what the derivation would predict: the candidates are partial solutions to the same partition.
6 The Satisfiability Question
The profile turns satisfiability into a comparison among composites.
Every coordinate is purchased, enabled, or approximated by at least one row. No single candidate occupies the upper corner. Restricted domains supply local existence cases: where physical outcomes arrive at deployment tempo and remain available to the evaluator, a feedback channel can meet lineage, anchor, and rate together (Ortulanu, 2026a §9). Breadth is the open variable. General frontier audit ranges over behaviors for which no single physical reference is available, and some preference-shaped domains may contain no uninhabitable anchor at all.
Combination is where the question lives. The conditions constrain different components — sources, reference, tempo — and do not exclude one another by content. They do create tension: lineage-disjoint apparatus tends to run slower than the substrate it audits, and non-migratable anchors often sit in slower processes (Ortulanu, 2026a §5). A composite can join physical anchors where timescales allow, lineage-checked judges where they do not, and attestation across the evidence channel. Its effective profile is set by the weakest necessary connection, not the strongest component.
If the upper corner is empty at frontier scale, the profile becomes a price statement. Recursive use would then require a declared choice among lineage distance, anchor quality, and tempo. The matrix makes that choice visible before deployment rather than after the external bill.
The conditions say what resistance costs. The profile says which designs pay which part.
7 Limitations
The analysis has limitations that bear explicit statement.
The lineage condition is stated without the resolution at which it binds. The first note flagged the missing measure — at what distributional distance audit independence is restored — as an open empirical question, and nothing here closes it (Ortulanu, 2026a §9). The condition’s force is comparative: candidates can be ordered by how much lineage disjointness they purchase, even while the sufficient threshold is unknown.
Formal minimality remains unproved. Operational irreducibility follows within the inherited capacity apparatus because each coordinate blocks a distinct deficit. A different decomposition of the same constraint surface may still exist (Ortulanu, 2026b §2).
The candidate matrix is a reading of the design landscape through August 2026, and it is time-indexed. New designs, or new combinations, can change the rows without changing the derivation; the conditions are the stable part of the analysis, the matrix its dated application.
Every judgment is indexed to a task dimension. An executable oracle can anchor correctness while leaving task-family coverage exposed; a physical outcome can be external while its measurement interface remains gameable. The matrix records the effective evidence channel, not the prestige of the component carrying it.
The treatment is static. A design that meets two conditions and fails one does not fail totally; it erodes, and the rate at which it erodes — whether partial satisfaction buys years or weeks — is the dynamics question this note’s scope excludes. The fifth note in the series takes up the trajectory.
The evidence base is structural rather than empirical. The matrix assesses what each candidate’s architecture can and cannot purchase; it does not measure how purchased conditions perform in deployment. The companion empirical track proposed alongside the third note’s markers would bear on this directly, and remains separate work.
8 Open Questions
Five questions remain.
The Combination Problem. Whether some composite of the matrix’s rows meets all three conditions at frontier scale is the practical heart of the matter, and it is unresolved. The composability argument of section 6 shows no contradiction among the conditions; it does not exhibit a combination. The candidates’ clustering — tempo compatibility in benchmark-shaped designs with inhabitable anchors, uninhabitable anchors in physical channels with latency — suggests the search space is narrow.
The Resolution Question. The lineage condition binds at an unspecified resolution. Training-data overlap, procedural inheritance, and downstream judgment correlation could turn ℓ(d) from an ordered grade into a checkable measure (Ortulanu, 2026a §9).
Partial Satisfaction. Most feasible designs will meet the conditions partially, and the informative question becomes temporal: how fast does the discriminative function erode as a function of which condition fails and by how much. That trajectory — including whether failure of the anchor condition dominates, as the third note’s mechanism suggests it should — is the subject of the fifth note in this series.
Domain Translation. The conditions are derived inside AI evaluation, but nothing in the derivation is specific to it: lineage, anchor, and rate are statements about any categorical benchmark used recursively. Whether the non-AI domains the third note named as candidates exhibit the same partition — and whether their occasional resistances are purchases of the same three conditions — is the sixth note’s question.
The Reception Question. The second note asked why the dominant framing of frontier-safety work positions capacity asymmetry outside the space of recognized problems (Ortulanu, 2026b §9). The conditions derived here sharpen that question: substrate-level independence is not technical work in the sense the field currently rewards, and the first note’s analysis of why such work goes unrecognized applies to the present derivation with full force (Ortulanu, 2026a §7). The frame-level treatment belongs to the seventh note.
The contribution is the profile: two senses of independence separated, three coordinates derived, seven candidate families located, and composition stated as a partial-order problem. A missing coordinate exposes a channel; whether that exposure produces erosion depends on the proxy and trajectory the fifth note supplies. The design grammar remains compact: close enough to read, far enough to judge, fast enough to retain the difference.
References
AI Security Institute. (2026). Research agenda. aisi.gov.uk. Accessed August 2026.
Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O’Keefe, C., Whittlestone, J., et al. (2023). Frontier AI regulation: Managing emerging risks to public safety. arXiv:2307.03718.
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073.
C2PA. (2026). C2PA specifications, version 2.4. Coalition for Content Provenance and Authenticity. spec.c2pa.org. Accessed August 2026.
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., & Stoica, I. (2024). LiveCodeBench: Holistic and contamination free evaluation of large language models for code. arXiv:2403.07974.
Kaggle. (2026). Game Arena. kaggle.com/game-arena. Accessed August 2026.
Kang, D., Hashimoto, T., Stoica, I., & Sun, Y. (2022). Scaling up trustless DNN inference with zero-knowledge proofs. arXiv:2210.08674.
Li, D., Sun, R., Huang, Y., Zhong, M., Jiang, B., Han, J., Zhang, X., Wang, W., & Liu, H. (2026). Preference leakage: A contamination problem in LLM-as-a-judge. Proceedings of the International Conference on Learning Representations. arXiv:2502.01534.
NIST. (2024). AI Risk Management Framework: Generative AI Profile (NIST AI 600-1). National Institute of Standards and Technology.
NIST. (2026). Center for AI Standards and Innovation. National Institute of Standards and Technology. nist.gov/caisi. Accessed August 2026.
Ortulanu, G. (2026a). A category-level failure mode not captured by distribution metrics. osservadore. doi:10.5281/zenodo.19627563.
Ortulanu, G. (2026b). The audit asymmetry problem: Capacity conditions for substrate-independent audit. osservadore. doi:10.5281/zenodo.20092930.
Ortulanu, G. (2026c). The reflexive contamination of categorical benchmarks. osservadore. doi:10.5281/zenodo.21252512.
Ortulanu, G. (2026e). Recursive evaluation collapse. osservadore (forthcoming).
Ortulanu, G. (2026h). The 2026 Zcash Orchard episode: An application of the audit asymmetry frame to a non-AI domain. osservadore (supplementary note, forthcoming).
Scale AI. (2026). AI model leaderboards. labs.scale.com/leaderboard. Accessed August 2026.
UK AI Safety Institute. (2024). Inspect: An open-source framework for large language model evaluations.
U.S. Department of Commerce. (2025). Statement from U.S. Secretary of Commerce Howard Lutnick on transforming the U.S. AI Safety Institute into the Center for AI Standards and Innovation. Press release.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685.