← AI Feed
AI Feed

The benchmark is measuring the harness, not the model

Three headline scores that belong to an assembled system rather than a model, two of them disclosed in the vendor's own footnotes.

How we organise

North Automations

Three reasons enterprise agent programmes stall: agents built one at a time against narrow tasks, sprawl nobody can govern because no single place defines behaviour, and the same model used at every step. What ships against that is a coordination layer with per-step model selection, versioning, approval checkpoints and a plan a person edits first.

We tell a client that the diagnosis is worth more than the product attached to it. None of what ships is model capability. All of it is change control for behaviour, which most firms already hold for infrastructure and have not pointed at agents. The recommended sequence ends with supervised autonomy, and that phrase carries an enormous amount of unspecified weight. Somebody has to say who supervises, at what point, and what record the review leaves.

How we build

Kimi K3

Weights released for a 2.8-trillion-parameter sparse model. The limitations section is the specific part. It was trained in preserved-thinking mode, so a harness that fails to pass back all historical reasoning makes generation quality highly unstable, and on ambiguous intent the model may act unsanctioned. The remedy recommended is explicit constraints written into a file in the repository.

Our position is that the limit on what an agent may decide has become a file. It needs an owner. It needs a review step and a change history, like any other file in the repository, or the agent’s decision rights are changed silently by whoever last had it open.

The regression the file records

The matching finding arrives from the other side. The model establishes ground truth by running code. Adherence to a specification is weak.

We judge those two traits as one review gate. It is not the gate you would build for a model that obeys.

How we assure

A cyber score of 95.95 per cent

The scored configuration is not a model. One cheap specialist is paired with an expensive generalist inside a harness of more than a hundred agents, and the specialist handles up to 90 per cent of tasks. Another vendor disclosed the same class of thing in a launch footnote: when a safety classifier refused a request, it routed to a different model rather than being refused.

We read the practitioner’s question as having changed shape. Not how did the model score, but what was the harness, what fell back, and to what. Those are assurance questions before they are technical ones, and almost no evaluation framework in commercial use asks them. A firm that runs a bake-off, picks a winner and moves on has bought a number for a set-up it does not own and cannot rebuild.

Forensics that no commercial API would run

Analysis of more than seventeen thousand recorded attacker events could not run on commercial frontier APIs. The payloads tripped guardrails that cannot tell an incident responder from an attacker. It ran on a self-hosted model instead.

We would steal the recommendation. Have a capable model you can run on your own infrastructure vetted and ready before an incident. That is not a procurement preference. It belongs in the same register as an offline backup.