← AI Feed
AI Feed

The model stopped being the variable

A multi-agent system delivered as one model, inference split across two kinds of hardware, and a voice channel wired straight to an execution path.

How we organise

A multi-agent system delivered as one model

Fugu is not a base model. It coordinates a pool of models behind a single compatible API, and posts 82.1 on one terminal benchmark against a 74.6 baseline. Specific models can be opted out of the pool from the console.

We tell a client that the architecture question has moved from which model to how you conduct several of them. Once orchestration is the unit, cost, quality and vendor optionality stop being a capability bet and become an organising decision somebody has to own. Note which control shipped first. Opting a named model out of the pool is what a regulated buyer asks for before anything else.

A founder reduces the gap to compute

The transcript is unconfirmed, so everything in it is attributed rather than established. The claim is that talent, model capability and applications all reduce to differences in compute. Hardware is said to pay for itself in about ten months.

Our reading is that a problem with a price tag is a convenient problem for somebody who has just closed a multi-billion-dollar round. The structural claim underneath probably deserves more attention than the valuation. If inference cost falls faster than training cost, the defensible position moves to whoever can serve a model at the lowest marginal cost, and giving weights away stops looking like altruism.

How we build

Inference split across two kinds of hardware

One rack takes the prompt side and a wafer-scale engine takes token generation, because the two phases have opposite hardware appetites. The headline is up to five times the tokens per second per watt. The footnote compares against one of the two vendors alone, rather than against a fleet of anybody else’s accelerators.

Our position is that serving architecture has become a cost lever rather than a procurement detail. None of it is free. Two hardware profiles mean one scheduler that understands both, a handover, and a class of failure that exists in neither half alone.

Scaling agentic reinforcement learning

The agent under training shares a sandbox with the machinery that grades it, so grading material is withheld until scoring time. The company says plainly that this is mitigation and not a guarantee.

We hold that a team standing up coding agents owns two systems: the build, and the thing that judges the build. The honesty of the second is what stops the first quietly gaming its own tests.

How we assure

Voice reaches the desktop

The offer is to direct multiple agents by voice. It reaches local files, plugins and connected tools.

We judge speech the least auditable input an enterprise has ever wired directly to an execution path. There is no typed record of what was asked and no diff to read before it runs, and the latency budget discourages a confirmation step. It arrived as a routine product update on enterprise tiers.

Two health releases, one checkable

One is a study of 13,917 participants with a blinded protocol and a published list of what it could not control. The other connects health records inside a consumer product.

We read both as diligent, so the difference is exposure rather than care. One published enough for somebody else to disagree with it. The other asks to be taken at its word, at a distribution several orders of magnitude larger. In a regulated domain, checkability is the whole of the assurance argument.

AI agents at work

Ninety-five per cent of executives are confident employees use AI responsibly, where 52 per cent of those employees admit using tools without approval. Only 34 per cent of firms apply the same security controls to agents as to people.

We ask a client to sit with the distance between 95 and 52. That gap does not close with better communication. Executives are describing the system they designed and employees the system they use, and both are answering honestly about different things. It is the ordinary condition of a control written down and never wired to anything that could fail. Most firms already own identity, least privilege, joiners and leavers, audit and revocation, and have simply not pointed any of it at the non-human workforce doing the work.

Reinterpreting the measures

Token usage is not a productivity measure. Pull-request throughput tells only part of the story.

We treat measurement as the assurance layer, and it is the one most firms have not rebuilt. Capability travels through products, arriving on enterprise tiers as an update nobody approved. Control travels through programmes, which have to be funded and staffed before anything moves.