I spent 4 months on an options trading system. It never traded real capital. A defect in its data pipeline was inventing profit the whole time, and no check before the final audit caught it. By then those checks included expert review panels, pre-registered experiments, out-of-sample tests and stress tests. At its peak the backtest reported 163% a year.

The data pipeline had left the price of the index the options were written on at zero, on 46% to 68% of rows depending on the year. When the price of an option the system held was missing, the code invented one from that zero, so each put it held closed at roughly its full strike price, and a spread booked about the width between its strikes as profit that did not exist. The policy had been trained on those prices, so it had learned to exploit the defect rather than trade the market. Every one of those checks measured the result through the same corrupted profit-and-loss path.

I commissioned an audit before any capital, real or paper, was committed to the system, and it found the defect. It used 6 AI reviewer agents, and a coordinator agent adjudicated their reports. Of those, 5 each took one surface of the system, such as the financial math or the training environment, and every finding had to name a concrete test that would expose it. The reviewer assigned to the environment was told that some rows had a zero index price and that one code path already patched them, and asked whether that patch interacted with other code. It examined the same fallback logic that invented the prices and marked it verified correct. The coordinator later found the reviewer had never tested the real-data path the backtests ran on.

The sixth reviewer, a skeptic, had no surface of its own. The first question in its brief was: what single bug, if it existed, would most cheaply explain the 163%? It disabled the code that invented the prices and reran one 2025 test, which went from +287.5% to -0.5%.

Measured honestly, the system lost about 2.6% a year. The option-selling strategy underneath it has a real edge, but at best about 1.2% to 2.5% a year depending on fills, against about 14% a year for holding the index over the same 8 years. The bar I had written down in advance was 3%, and the strategy missed it even at perfect fills, so I stopped.

Before that audit, I believed that rigorous review was itself a check on the answer. The environment reviewer and the skeptic had the same data, the same ability to run code, and an adversarial brief each. The environment reviewer read the fallback and ran its probes on other paths, so no result it produced could have contradicted the 163%. The skeptic removed the suspected cause and measured again, and that measurement could have left the result standing. It erased it instead.

Two AI reviewer agents in the same audit of the options system. Both could read the pricing code and run code against the data. The environment reviewer was told zero index prices were patched in one code path and asked whether that patch interacted with other code; it read the same fallback and marked it verified correct. The skeptic was asked what single bug would most cheaply explain the 163%, disabled the suspect code, and reran one 2025 test, which went from +287.5% to -0.5%.

I wanted to know whether that difference held outside my project, so I reread the research on critics, sampling and debate. Most of what I found rests on 3 papers.

A critic’s report is a lead

The strongest case for a critic is McAleese et al. (OpenAI, 2024). CriticGPT and the model it reviewed were initialized from the same checkpoint. On short Python tasks, CriticGPT caught more human-inserted bugs than contractors who had a median of 5 years of Python, spent roughly 50 minutes per example, and could execute the code. The authors also ran it over ChatGPT training data, mostly not code, that a first annotator had rated flawless. Where the critique flagged something, reviewers said in 24% of those cases that it had found a problem that substantially lowered their rating.

The critic learned this from training data, human-labeled examples of bugs that people had inserted on purpose. Models trained without that data “severely under-performed.” The code results are scoped to short single-file Python, and the authors note that the critic’s hallucination and nitpick rates are much higher than the humans'.

I read this as a real detector with a limit I had not expected. During training, someone knew where every inserted bug was, so the training signal could be checked against a known answer. At review time the critic reports what it finds, and nothing in the report says how far to believe it. My environment reviewer also produced a confident report. It found real bugs, and in the same voice it marked the fallback that invented the prices verified correct.

The misses are in selection

My first explanation for why a review loop misses a failure was that the critic shares the proposer’s blind spots. Several instances of the same reasoning cannot see what one instance cannot see. Brown et al. (arXiv:2407.21787v1) showed me that explanation was incomplete.

They sampled one model repeatedly and counted whether the right answer appeared anywhere among the samples. Llama-3-8B-Instruct on MATH reaches 79.8% coverage at 100 samples and 95.3% at 10,000. Over those same samples, majority vote and reward-model scoring improve the solve rate only from 38.7% to 39.8%.

Coverage climbs while selection stays flat. Llama-3-8B-Instruct on MATH, from 100 to 10,000 samples per problem, log scale. Coverage, the share of problems where any sample is right, rises from 79.8% to 95.3%. Selection by majority vote or reward model over the same samples goes from 38.7% to 39.8%, leaving a gap of 55.5 points at 10,000 samples.

Coverage here is an upper bound that assumes a perfect checker, and it is measured over repeated samples of an answer, not over reviewers of one answer. It does not measure a review panel directly. It does show where the constraint sits. The right answer is reachable far more often than the loop reaches it, and neither selector checks anything outside a model. Majority vote counts agreement among the samples, so a right answer that appears in only a few of them stays outvoted. A reward model is another model’s opinion. Coverage turns into accuracy only where something outside the model can check the answer.

My audit had the same shape at a small scale. Only 1 of the 6 reviewers reported the defect, and another had marked the same code verified correct. The skeptic’s rerun with the price-inventing fallback disabled decided between them.

Where debate worked, a verifier did the checking

Khan et al. (ICML 2024) is the paper usually cited as evidence that debate works, and it is a real result on one task family, QuALITY long-context reading comprehension. Non-expert human judges reach 88% accuracy with debate against 60% naive, and GPT-4-Turbo judges 76% against 48%. The judge is denied the source passage, and the debaters have it. Quotes are machine-verified against the passage, and the judge is instructed to trust only verified quote tags.

The same paper, in Appendix C.2, reports the authors’ earlier exploratory work in the symmetric regime, with no information or capability asymmetry between judge and debaters. Debate scored below the best single-model baseline on 9 of 10 benchmarks. Their sentence, with their own qualifier inside it: “debate in our inference time setup does not lead to truth-seeking behaviour in situations where the judge has the same information and capability as the debaters.”

Two configurations in Khan et al. On the left, the cited result: the judge is denied a passage of about 8,000 tokens that the debaters hold, quotes are machine-verified, and human judges go from 60% to 88% on QuALITY. On the right, Appendix C.2: with no information or capability asymmetry, debate scored below the best single-model baseline on 9 of 10 benchmarks.

The symmetric regime is the configuration in this paper closest to what a production two-instance loop runs, though their debaters argued fixed opposing answers rather than proposing and critiquing. The two runs are separate experiments, not an ablation, and they differ in more than one way, in task and above all in whether the judge can see the source. The authors attribute the gap to that information asymmetry. The quote check made the asymmetry usable. The judge could not read the passage, but it could trust a quote the machine had matched against it, and that check is the part of the cited setup that a model’s judgment does not decide.

A channel to the world

Each loop whose verdict I could trust, in Brown, in Khan and in my audit, had a channel to the world. A channel is a part of the loop whose output is decided by something other than a model’s judgment and could contradict the claim being checked, such as an execution, an instrument, or a threshold fixed in advance. In Brown’s MATH runs it is the checker that coverage assumes, and without one, majority vote and a reward model stall near 40% for Llama-3-8B-Instruct. In Khan’s debates it is the quote verifier. In my audit it was the skeptic’s rerun with the price-inventing fallback disabled. The environment reviewer also ran code, but it never ran the measurement with the suspect removed. CriticGPT had one during training, on bugs whose locations were known, and none at review time, so each finding it reports still needs a check of its own.

What I recommend

  1. Treat what a critic finds as a lead to check. A critic finds real defects, including in answers a human annotator had already rated flawless. Before you act on a finding, run a test that would fail if the defect is real and pass if it is not.

  2. Build the selector before you draw more samples. The right answer is in the model’s samples far more often than any single sample is right. Over the same Llama-3-8B-Instruct samples on MATH, majority vote and reward-model scoring improve the solve rate only from 38.7% to 39.8%.

  3. If you run debate, give the judge a check that argument cannot overturn. Khan’s 60% to 88% gain for human judges came in a setup where the debaters held a source the judge never saw and every quote was machine-verified against it. In the same paper’s earlier runs, where judge and debaters had the same information and capability, debate fell below the best single-model baseline on 9 of 10 benchmarks.

  4. Ask what in the loop is not a model’s opinion. For any loop you run, find the part that is decided by something other than a model’s judgment. An execution counts only if its result could contradict the claim, for example a measurement with the suspected cause removed. If you cannot find one, treat the loop’s verdict as one more opinion.

The second part of this series covers the harness I built after the audit, and the 4 ways it broke once the channels were mine.


I used Claude in drafting this post.