← AI Feed
AI Feed

It ran, and it could not keep up

The smallest model was the most expensive to run, forty-eight coding agents failed together, and a review that worked properly still let defects through.

How we Organise

The smallest model in the comparison was the most expensive to run

On a coding benchmark a 2B model burned 14,308 joules to pass 64.5 per cent of cases. A 35B model burned 4,804 to pass 93.9. The smallest model in the comparison cost three times what the largest did.

We tell a client that a cheap model which is wrong more often ends up dearer than the expensive one that is right. That has been an argument from first principles, and a client has always been free to disagree. Here it is measured, and the money goes on repeated work rather than on the price of a call. Energy on a benchmark is not what anybody pays.

How we Build

Forty-eight coding agents wrote one specification, and the failures arrived together

Forty-eight agent-written implementations of one specification met a million randomised inputs. Independence predicts 115 cases where versions fail together. There were 429. Voting across triples did cut the worst case sharply.

Our position is that a team picks its construction deliberately and writes down why. It can do that completely and still be wrong about what it bought. Ruling out a single call because you want independent review satisfies us, and independence is what these numbers say does not turn up. A smaller worst case does, which is worth having and is a different product.

How we Assure

Anthropic froze its training environments for a month to catch up with itself

An automated review of every training environment ran continuously and fell behind what was being produced. During a month-long freeze, more than 10 per cent of the production mix was flagged for reward hacking and misconfiguration.

We ask that anything able to alter how an agent behaves passes a review somebody writes down. All of that was true here and it was not enough. The review ran, it produced the flags, and defects reached training runs while our test read green. What bound was the rate at which review finished, and we never ask about rate. One firm is describing itself, with figures it chose.