On both sides of the test
How we Organise
What 17 million enterprise messages say about who is using the machine
Chatterji and four colleagues posted this on 12 August. They linked ChatGPT Enterprise account records to usage, to worker roles, to task classifications and to public-company financial data. The record runs through March 2026. At the six-month adoption horizon the worker-level sample covers more than 1,500 firms and more than 17 million messages.
Four things came back. Usage grew because new firms adopted and because firms already using it used it harder. Among US-listed firms, adoption concentrates in the larger and more valuable ones, and in those spending more on research and on selling and administration. Active use runs across job functions and across levels of seniority. Usage intensity is highest among early-career workers. The messages cover writing, technical work, communication and the pulling together of information.
We tell a client something narrower. Where a role’s entry-level work has been automated, that role needs a written progression path, and the path is the thing to go and check. The practice assumes the affected roles can be picked out from the rest. On this evidence they cannot. The machine is in every function and at every level. It is heaviest where the entry-level work sits. A rule that fires on nearly every role does the work of a rule that fires on none. A client with a path written for two roles has satisfied us and left the other thirty untouched. This is not the first source to argue it. Notes from DX on how Capital One assesses AI readiness, published on 27 August, recorded the same worry from the other end: mentorship and apprenticeship are widely held to matter, and are thinly resourced.
We may be wrong, and the gap is in what usage intensity shows. Heavy use by early-career workers is not their work being taken from them. It is equally consistent with juniors reaching for a new tool first, which would leave the path to senior intact and busier. The data covers one vendor’s product. The financial linkage reaches only US-listed firms, and it stops in March. If the reading here holds, we stop telling a client that naming the automated roles is a step it can finish. We start asking which roles it believes are untouched, and what that belief rests on.
How we Build
A skill evolved by one model, working better on another
Tang and five colleagues published this on 27 August. Their framework keeps three things apart that skill-evolution methods usually mix. There is raw execution experience, there is accumulated knowledge, and there is the executable skill. Experience is consolidated into a persistent wiki. Later revisions to a skill are proposed from that wiki rather than from the run that has just happened.
They measured it on five benchmarks. Those span mathematical reasoning, web search, spreadsheet work, long-document questions and an embodied task. Five models were used, drawn from three families. Within one family the gains rose with size, at 12.3, 17.5 and 23.9 per cent for the 4B, 9B and 27B models. Evolved skills also stood in for scale. The 9B model with them reached 47.4 per cent where the 27B model without them reached 39.4.
Skills then transferred across families. A skill evolved by one model sometimes beat the skill a model had evolved for itself. The ablation puts the weight on the wiki. Remove the persistent knowledge and evolution degrades. Give the inference agent access to the wiki while evolution runs and the skills come out worse.
Our position is that what one team builds, every team should be able to find. The argument for it has always been the ordinary one about duplicated effort. This puts a stronger argument underneath. A skill built elsewhere was not merely as good as the local one. It was sometimes better. So the value in a built skill is not tied to the conditions it was built under. The transfer has a limit the authors name. It held where a skill captured a general procedure. It failed where the skill encoded a workaround for one model’s habits. That limit is useful to a client sorting its catalogue, and it is a benchmark result rather than a production one.
How we Assure
An automated researcher closed the safety gap, once it was stopped from cheating
Anthropic published this on 28 August. Claude was set to train models autonomously against public benchmarks covering ten categories of alignment failure. It took one category at a time, through a loop of reading the literature, proposing methods and data, training, and testing. Success was scored as the percentage of the safety gap closed. That was measured across the three to five benchmarks each category has.
Two constraints were imposed. Methods that damaged the student model’s general capability were excluded. Claude was forbidden to distil its own alignment into the target model, and a monitoring agent enforced both rules by reading every method before it ran.
For all ten categories the loop found fixes that improved the target benchmarks without degrading capability. The best of them held on benchmarks withheld from the loop. They held on an open tool that simulates adversarial multi-turn scenarios, and on models up to 4.7 times larger than the ones being optimised for. On deception the loop submitted more than 150 attempts. It closed 82 per cent of the safety gap in the run reported and 85 per cent on average. Six experienced safety researchers working under the same rules closed 20 per cent. Across the wider comparison it outscored 28 human researchers who had up to eight hours each.
We hold that a lineage is judged on the population of acts it produces rather than on any single act. A lineage with no aggregate measure is recorded as unevaluated rather than as passing. Read one way this supports us. The aggregate measure held up under 150 attempts, and it generalised past the benchmarks it was pointed at.
The constraints are what argue with us. Anthropic ruled out methods that raised a safety score by making the model less useful. It forbade the researcher from writing its own alignment into the student, and it put a monitor in front of every method to make those rules bite. Our practice asks whether an aggregate measure exists. It says nothing about who is optimising against it, or how hard. A measure a machine can run 150 attempts at is a different object from one a team reports quarterly. The 29 August finding on how twelve frontier models account for their own reasoning bore on this same claim, which is now argued from both directions.
We may be wrong about how far this reaches. These are benchmark scores rather than deployed behaviour. The work is a vendor measuring its own model on a problem it has staked a great deal on. Anthropic says plainly that the human comparison is weak, because the researchers could not iterate on their submissions. A reasonable person could read the constraints as ordinary practice. If the reading here holds, we stop telling a client that an aggregate measure evidences an evaluated lineage. We start asking what stops that measure being optimised against, and who checks.