Fifty-eight per cent were never once right
Relayed output that builds no judgement, a maintenance cost that became a cron entry, and 160 accounting tasks of which 58 per cent were never once solved.
How we organise
Pasting model output instead of answering costs the reader more than prompting would. Nobody in the chain checked it. Your own words are the proof.
We tell a client this reads as manners and is competence. A relay builds no judgement. The cost lands on the reader.
Build the tool locally, rebase onto upstream nightly, replace the running version. Maintenance was the expensive half. It is now a cron entry.
Our reading is that build-or-buy for internal tooling has moved, and few firms have repriced it. A tool you cannot read is one you cannot bend.
How we build
Six months of voice infrastructure
Almost none of the work was model work. The gains came from the media layer and the orchestration. Under load a CPU-side process saturated first.
We judge a capacity plan by what it was written against. Most are written against the GPU count. That is the part nobody had to discover.
On raw throughput one accelerator loses by about 40 per cent. On tokens per pound it wins. Coverage reported the second and called it faster.
We hold that throughput and cost per token are different claims. A latency-bound workload needs the faster node. A budget-bound one needs the cheaper.
How we assure
An accounting benchmark, run eight times
160 month-end-close tasks, every model run eight times. 58 per cent were never solved on any run. The steadiest model managed 2.6 per cent.
We read consistency rather than the leaderboard. A single-run score is capability, and nobody buys capability. A firm buys the same answer twice.
Speed up inference under a quality gate. The winning verified entry reached 510 tokens a second. A faster run at 536 failed the gate.
We ask who ran the gate. Fastest and fastest-verified are separate claims. Where the team making the change owns the measurement, that is a speed number.