The default changed, not the policy
A spend cap enforced rather than reported, a review process carrying four jobs at twice the volume, and a sandbox turned on for everybody instead of documented.
How we organise
The gateway sits in front of models, servers and assistants. It reports cost by team and application. A budget can be a hard cap.
We tell a client that spend control has been a reporting problem. A cap makes it a setting. Somebody must then name the number at which an agent stops.
Lines of code per human-landed diff at Meta rose 106 per cent. Generation got faster and review did not. Review already carried four purposes at once.
Our reading is that this is a capacity problem in a quality costume. Style belongs in a linter. What remains is a smaller job a person can do well.
How we build
Agents run in the background, and everything they do goes into a replay-exact log. One case ran a thousand tool calls over a day.
We judge an unsupervised run by whether anyone can reconstruct it. The logging is the part worth copying. A chat transcript is not a record.
A five-member spreadsheet team
One general-purpose agent got 89.69 per cent of spreadsheet cells right. It got 34 per cent of the answers right. Same runs.
We hold that per-step accuracy compounds. Nine in ten per step is wrong most of the time. Splitting the roles inserts a check.
How we assure
Zed sandboxes its agent by default
The agent’s terminal is now sandboxed for every user, enforced by the operating system rather than the application. Git writes and network requests are blocked.
We read a default as worth more than an option. Most users change nothing. Those two denials are what turn a compromised agent into a supply-chain incident.
Mistral says its model matches or outperforms open models seven times its size. Aggregators reported a win over a 20-billion-parameter model. Mistral names none.
We ask that a forwarded claim be read back against the vendor’s wording. The hedge is doing the work. It disappears in the retelling.