What the Evals Missed

The previous post was about writing a reward function and watching a policy exploit it. This one is about the harder question underneath: how did I know any of it was working? Not “did the loss go down.” Loss went down the entire time, including during the weeks the policy was quietly learning to game me. I mean the actual question — is this thing doing what I meant, and how would I find out if it were not. ...

August 19, 2026 · 12 min · Aashish Sheshadri

Two Agents Arguing Is Not a Verifier

I spent 4 months on an options trading system. It never traded real capital. A defect in its data pipeline was inventing profit the whole time, and no check before the final audit caught it. By then those checks included expert review panels, pre-registered experiments, out-of-sample tests and stress tests. At its peak the backtest reported 163% a year. The data pipeline had left the price of the index the options were written on at zero, on 46% to 68% of rows depending on the year. When the price of an option the system held was missing, the code invented one from that zero, so each put it held closed at roughly its full strike price, and a spread booked about the width between its strikes as profit that did not exist. The policy had been trained on those prices, so it had learned to exploit the defect rather than trade the market. Every one of those checks measured the result through the same corrupted profit-and-loss path. ...

September 30, 2026 · 8 min · Aashish Sheshadri