We are looking for collaborators and critics. If you think you can contribute, send a short email to hezhonghao2030@gmail.com — a few lines on what you work on and what you would want to push on is enough. We will add you to the Slack channel from there.
Six profiles we are specifically hoping to hear from are at the bottom. If you disagree with the central thesis, that counts — the failure conditions below are the ones we consider binding.
The short version
Recursive self-improvement (RSI) — an AI system that iteratively improves itself — is the dominant story for how AI becomes transformative, and the implicit end-game for several frontier labs. It rests on a hidden premise: that “better” can be verified cheaply and automatically.
Every domain where RSI-style bootstrapping has worked supplies an automatable verifier — game self-play, formal proof, code against tests. Where the target is instead human taste, judgment or values, no such verifier exists, and self-improvement saturates.
Past that point we propose a different regime: recursive co-improvement (RCI) — a loop in which the AI sharpens the human’s capacity to articulate and verify judgment, and the human’s sharpened judgment in turn refines the AI’s objective, each round enabling a harder next one.
These are two parallel regimes, not one ladder. RSI is the right account below the boundary. We are not arguing against it; we are arguing about where it stops.
The central thesis
Verification decides whether an improvement loop produces real gains or merely measured ones — and past the verification boundary the only verifier available is a human, so the human is the thing that has to get better.
Three claims, in the order they have to survive:
- Improvement loops are limited by their verifier, not their generator. Ungrounded self-critique degrades: ten rounds produce a 55% decline in informational change, and a single grounding step restores forward movement. The gap between measured and real gain grows with recursion depth — reward hacking rises from 26.4% of optimisations at ten steps to 57.8% at a hundred. Verifier-guided training provably converges on the verifier’s knowledge centre, plateauing or reversing unless the verifier is reliable.
- Past the boundary, no mechanical verifier exists to be improved. The generation–verification gap does not close on its own there, because nothing automatable is available to close it.
- So the human verifier becomes the object of improvement — not as a data source and not as an approver, but as a capacity the loop can move rather than a fixed input it consumes.
A claim we deliberately do not make. It is tempting to say progress in RSI is attributable to improvements in verifiers. We think that is too strong. It is a causal claim over observational data with at least four uncontrolled alternatives: base-model capability, inference compute, elicitation of latent capability rather than acquisition of new capability, and scaffolding.
Several results cut the other way. Training on reasoning chains that all reach the wrong answer has beaten training on human annotation. Inference compute alone has matched expert physician annotation in a domain with no ground truth. Developer-written context files have lowered task success. The survey mapping this literature calls the pattern “a qualitative pattern… not a measured law.”
Our argument does not need the causal claim. It rests on availability: where no mechanical verifier exists, the human is the only verifier there is.
The five levels
The axis is how automatable is verification of task performance. Time-to-feedback, legibility, and whether a market prices the output all correlate with the axis but none of them defines it.

The five levels. RSI is the right account to the left of the boundary; RCI is the regime to the right of it.
| Level | What verification looks like | Example domains | Where it breaks |
|---|---|---|---|
| L1 | A sound, cheap, generator-independent checker settles correctness outright | Machine-checked proof, chess and Go, hidden-test code judging | Nothing breaks. The human is not in the loop and does not need to be |
| L2 | Automatable only through a cheap proxy standing in for the real objective | Unit tests, benchmark labels, verifiable-reward training | Proxy and objective diverge: optimisation pressure selects against the very quality the check was meant to certify |
| L3 | Grounded in fact and possible in principle, but too slow, costly or confounded to keep pace with generation | Weather, clinical trials, materials discovery, chip fabrication, vehicle autonomy | Verification cannot keep up with generation, so human judgment has to bridge the latency |
| L4 | No mind-independent verifier exists — a community’s collective verdict constitutes the standard rather than tracking something behind it | Peer review, appellate review, consensus clinical guidelines, theory choice | The verdict is slow and fallible and nothing behind it does the correcting. Peer review’s inter-rater reliability sits around κ ≈ 0.17 |
| L5 | Even the human signal is latent. It has to be elicited and refined rather than read off, because the standard is still under construction | Morality and constitutional questions, frontier aesthetics, ideal reflective preferences | There is no determinate standard yet to converge on. Only consistency conditions constrain the answer |
What would show us wrong
Stated up front, because a position with no failure condition is not a research programme.
- If a self-improvement loop with no human in it produces sustained real gains on a task with no checkable ground truth, the central claim fails.
- If human judgment, measured properly, turns out not to improve under the interventions we build, the programme fails regardless of what happens to RSI.
- If a human’s independence degrades as their accuracy improves — if co-improvement produces a better-performing but more correlated verifier — then the loop is consuming the asset it runs on, and we should say so rather than redefine the metric.
Background reading: our read of the ICLR 2026 RSI workshop and scalable oversight and the verification boundary.
Who we are hoping to hear from
We want collaborators and we want critics, and we are more interested in the second. Six profiles, with what we would actually want from each.
You work on recursive self-improvement
What we want: to be corrected. Our argument is about where RSI stops, so it is only as good as our picture of the frontier — and that picture comes from a survey and one workshop’s proceedings, not from building these systems. If we have mischaracterised what current self-improvement loops do, where they saturate, or why, we would rather hear it from you than from a reviewer.
You have thought about gradual disempowerment
Our failure mode above L3 — the human verifier degrading under load, sliding from reviewer to approver — is a mechanism-level version of what Kulveit et al. (2025) describe at the level of societal systems. What we want: help connecting the two scales. We can say what happens to one reviewer facing more candidates than they can check. We are much weaker on what happens to an institution.
You think about post-AGI social, economic and political development
The boundary implies a claim about which human activities stay load-bearing and which get hollowed out, and therefore about where economic and political agency sits after broad automation. What we want: pressure on whether that survives contact with how institutions really work.
The body of work we have most in mind here is Dr Iason Gabriel’s — AGI and Society Lead at Google DeepMind — rather than any single paper. The connection is closer than the topic labels suggest. Our boundary is defined against an ideal standard: the judgment a competent evaluator would endorse on reflection. That is the same object Artificial Intelligence, Values and Alignment takes as its central problem, namely fair principles that receive reflective endorsement despite wide variation in what people actually believe. What we want: to find out whether our L5 uses that concept in a way his account supports, or in a way it rules out.
You work on safety and alignment, especially scalable oversight or socio-technical alignment
This is the most direct engagement and where we most expect to be wrong. The scalable-oversight programme — debate, iterated amplification, recursive reward modelling, prover–verifier games, weak-to-strong generalisation, measuring progress on scalable oversight — argues that past the point of direct human checking the answer is a better protocol rather than a better human. We think those two literatures have not been reading each other, and we have written up why we think the protocols stop where they do. If that reading is unfair, we want to know.
You work on epistemic risks
The failure we care about most is not an AI overriding human judgment. It is human judgment quietly ceasing to be exercised while its institutional form stays in place.
The field-level map we work from is AI Epistemic Risks: Emerging Mechanisms and Evidence (Yang, Casper, Stray et al., 2026), a thirty-author survey that sorts the area into three mechanisms: persuasion and manipulation, cognitive offloading, and feedback loops. Our own argument lives inside two of them. A verifier who slides from reviewer to approver is a case of cognitive offloading, and the L4 failure — a collective verdict that constitutes the standard while its inputs increasingly come from the systems it is meant to check — is a case of the third. The individual-to-collective version of that third mechanism is developed in The Lock-in Hypothesis (Qiu, He, Chugh and Kleiman-Weiner, ICML 2025). One of us is a co-author on both.
What we want: work on measuring epistemic harm that does not reduce to task accuracy. The survey is good on mechanisms and honest about how thin the measurement is.
You are a philosopher of science, epistemology or ethics
L4 and L5 are doing philosophical work and we would rather do it well. L4 asserts that a collective verdict constitutes a standard rather than tracking one, which is close to Longino’s account of objectivity as a property of communities, and to Hart’s rule of recognition. L5 asserts that some standards are not merely unknown but not yet determinate. What we want: to find out whether that distinction is defensible, and whether we have reinvented something that already has a name.
To get involved: shoot a short email to hezhonghao2030@gmail.com saying what you work on and what you would want to push on. If you think you can contribute, we will add you to the Slack channel from there. Disagreement is welcome and is the most useful thing you can send.