The model you bought is not the model you got
A quota that repriced itself with no invoice line changing, an index of what a provider does to a model's accuracy, and one benchmark task costing $2,600 and another $251.
How we organise
Twice the tokens at the same price
Two fourteen-day windows of one developer’s coding logs, before and after a model change. Tokens per session rose 2.25 times. Per-token pricing did not move.
We tell a client that a quota can reprice itself with no invoice line changing. Nothing triggers a review, because consumption is what moved. One developer’s logs are not a controlled study.
A voluntary pre-release testing framework
Developers would give government up to thirty days of early access to frontier models for cyber assessment. The benchmark used is classified. Nothing may become licensing.
Our reading is that a pre-release gate is arriving as practice before it arrives as law. A firm that already runs one will find the paperwork trivial. A firm deploying on vendor assurance holds no evidence at all.
How we build
Whether an agent harness can be reused
One argument says a harness has to be bonded into the application it serves. Within hours AWS published the opposite bet. Its three separate harnesses became one behind a defined protocol.
We judge this as a boundary a firm draws rather than a vendor it picks. Draw it too high and you maintain three of everything. Draw it too low and your agents run inside somebody else’s scaffold.
How we assure
An index of what a provider does to accuracy
The index scores how far a provider’s accuracy falls below a self-hosted reference of the same model. Tool calling, hard reasoning, long-context recall. Snapshots rather than monitoring.
We hold that the name on a contract does not fix the accuracy delivered. Quantisation, serving configuration and routing all sit in between. None of them appears in a procurement document.
MirrorCode, and what a task costs
One task cost 2,600 dollars across nineteen unattended days. Another rebuilt a sixteen-thousand-line toolkit in fourteen hours for 251 dollars. Same benchmark.
We read the spread rather than either figure. Long-horizon autonomous work has no reliable unit cost yet. A fixed price against it is a bet on where you land.