← AI Feed
AI Feed

The record looked right

Desktop agents at production scale failing where nobody can check them, a capable agent model on one consumer GPU, and a result that stood because of what checked it.

How we organise

a16z on what desktop agents are doing in production

The best model now scores 85 per cent on the standard desktop benchmark, against roughly 72 for human testers, and teams are running millions of portal interactions a month. The authors are exact about where it stops. An agent that reads net 60 as net 30 writes a record that looks perfectly plausible.

We tell a client to sort candidate work by where its check comes from, before looking at any model. Both failures named here are failures of checking rather than of capability. A wrong record reads like a right one, and the truth arrives two days later as a phone call.

How we build

Meta puts a 30-billion-parameter agent model on one GPU

The weights are published under a permissive licence. Compressed to roughly 4-bit precision the model fits under 20 GB, which leaves room for its working memory inside a consumer card. Meta reports minimal degradation on agentic work and says it is trained to retry a failed tool call.

Our position is that where an agent runs has become a choice again. For two years buying a capable agent meant buying network access to somebody else’s data centre, and that one decision set cost, latency, data path and regulatory position together. This will not do frontier work.

How we assure

A mathematical bound raised, and what checked it

A research model moved a longstanding lower bound from 41.6 per cent to 67.2, over two sessions and 31 million output tokens. The first 650 ideas failed. Of about 60 subagents it coordinated, 13 did nothing but check the arguments the others produced, and the finished proof passes a machine checker.

We judge checking as a budget line of its own rather than a read-through at the end. Nothing here rests on trusting the model: every claim is settled by something outside the thing that produced it. That cost roughly a fifth of the agents in the account.