← AI Feed
AI Feed

The model did not change, the system did

A benchmark score tripled without a new model, a harness that has become a training target, and a benchmark built from work a firm had already shipped.

How we organise

Building abundant intelligence

One model’s price dropped 80 per cent. Better context management took another model’s score from 13.3 per cent to 38.3, on six times fewer tokens. The model itself did not change.

We tell a client to price a successful outcome, retries and review included. The token price is the invoice line. The vendor with most to gain from it has conceded that the two diverge.

How we build

A model trained against its rivals’ harnesses

Qwen3.8-Max is live and will ship open weights. A sixteen-day autonomous run produced 265 commits. It was trained across named third-party harnesses, including its competitors’.

Our position is that the harness has become a training target. Two identical models in two scaffolds are not the same system, and the vendor knows which one it trained against.

An MIT-licensed model at frontier scores

V4 Flash moves from preview to release under an MIT licence, its own decoding module drafting seven tokens ahead. It posts a score that sat with the closed frontier this spring.

We judge an architecture decision by the date it carries. A review that concluded agentic work needs a closed API carries one, and results like this expire it.

How we assure

A benchmark built from merged pull requests

Eighty tasks, each from a pull request the firm’s agent shipped after review. The prompts come from what the engineer asked. Any task every model solves is deleted.

We read the evaluation that matters as the one built from work you have already shipped. That handles contamination and saturation at once. It could not be bought.

Ten results, each with a machine-checkable proof

Ten long-standing problems resolved or advanced by an internal model. Every argument was formalised so that a machine could check it, not just a reader. Claiming human authorship would misrepresent both contributions.

We ask a client for its attribution rule. Two disciplines are on show: a certificate on every claim, and an account of who produced what. Most firms have neither.