← AI Feed
AI Feed

It held for one round

Eight months of prompts inside one firm, a code review that runs past its opening exchange, and a model that calls a question unanswerable and then answers it.

How we Organise

Sophistication in genAI use, read off eight months of one firm’s prompts

Across 713,564 prompts from nearly 4,000 employees over eight months, how well people used the tool neither improved over time nor improved lastingly after formal training.

We ask a firm to work out who else must change how they work, teach them, and prove it before the release goes out. The proof happens once. What a person does afterwards looks close to independent of it. A firm can run the training, clear the bar and be no different a quarter later. One firm, one back office, and data nobody outside can read.

How we Build

A code review benchmark that does not stop after the opening exchange

Automated code review is nearly always measured as a single decision. Measured across the rounds a real review takes, capability falls away as they accumulate. The defects most often missed are the subtle ones.

Our advice is to let the machine take whatever it can before a reviewer opens anything. That was settled on the opening exchange, because one round is what the published work measures. Follow it exactly and the machine gets a growing share of the work it does worst. This is a benchmark of replayed reviews, on models nobody fitted to a codebase.

How we Assure

A model that calls a question unanswerable and then answers it

Shown a professional-looking panel of invented numbers, models committed to a call on a provably unpredictable question far more often than when asked the bare question. Their stated probabilities barely moved across a gradient that swung action by 48 points.

We refuse to treat an agent’s own account of its reasoning as an audit record. This supports that more sharply than anything before it. Read the account and you conclude the model stayed uncertain. That is true of what it said and false of what it did. A single author, one domain, and the effect sits in some models rather than all.