← AI Feed
AI Feed

The change arrived without a release

How we Organise

The people who used to describe a change are now shipping it

Linear published a report drawn from its own product telemetry, covering June 2024 to August 2026. Three cuts carry the finding, each with its own population. Pull requests opened per workspace are up 111 per cent on a June 2024 baseline, across 47,900 paid workspaces in June 2026. Teams that connected a coding agent roughly tripled their weekly pull requests over two years, from 21 to 65. Teams without one went from 8 to 10. That cut covers 6,887 paid teams, 4,280 with an agent and 2,607 without.

The third cut is the one that matters here. The share of product managers attaching a pull request rose from 3 per cent to 10 per cent in two years. Designers went from 1 per cent to 8 per cent, across 166,000 paid users. Linear counts only pull requests in repositories connected to Linear, and says so. Those shares are floors rather than ceilings.

We ask a team to name or count everyone outside it who has to work differently because of what the team shipped. Each of them gets a route to learn. Each is then shown to be able against a bar set before the release. The trigger is a release. Nobody released anything to those product managers and designers. They picked up a general tool on their own. The share of them writing code trebled in two years with no release behind it, no learning route, and no bar set before anything.

This is not the first source to argue that this practice is watching the wrong moment. The field study reported here on 29 August, across 713,564 prompts from nearly 4,000 back-office employees, found that a bar cleared on release day predicted little about the months after it. That one says we read the bar too early. This one says that for the largest group, we never read it at all.

We may be wrong, and the limits are real. This is one vendor’s customer base. Those customers had already chosen an issue tracker that integrates coding agents. Linear notes that the agent teams were higher-output before they connected an agent. It counts pull requests opened rather than merged, which says nothing about whether the change was any good. Somebody could reasonably read the whole thing as a selection effect.

If it holds, we stop telling a client that counting who works differently after each release covers the people whose work has changed.

How we Build

The memory was accurate, relevant, and made the answer worse

Mengru Wang and eight colleagues submitted MemTrapBench on 20 August. Their subject is what a stored memory does to reasoning on the task in front of the model, rather than whether it was extracted and retrieved correctly. They name two failure shapes, reasoning fixation and belief distortion. Their dialogues run in three stages. A prior is planted in a plausible setting and applied repeatedly. It is then buried under unrelated turns until the dialogue runs 18 to 40 turns. The final question stays related to the history while changing the conditions under which that prior should apply.

Tested on the Gemini and Qwen families against recent memory frameworks, every memory strategy scored below the no-memory setting, and the strongest still fell by more than 10 percentage points on Gemini-3-Flash-Preview and Qwen3-30B-A3B-Instruct-2507. On a smaller subset scored by two judges, the gap is wider: responses without memory averaged 92.29 per cent against 31.05 with it under GPT-5.2, and 95.57 against 40.07 under Claude-Sonnet-4.6. Controlled runs attribute the fall to what the memory meant rather than to the extra context length.

Our own practice bounds what a system may remember. That memory expires, and somebody has tested the expiry. Every one of those is hygiene. Each catches memory that should not be there any more. What was measured here is memory that should be there by every rule we state. It was recorded faithfully, still relevant to the question, inside its bounds and nowhere near expiry. A team passes our test completely and takes the fall anyway. Our test separates memory that is stale from memory that is current. The damage came from memory that was neither.

The authors are careful about what they built, and so should we be. They call it a diagnostic stress test of harmful memory influence rather than a general evaluation of memory utility. The dialogues were built to induce the failure they then measure. Two model families, and a model wrote the dialogues. The abstract says five memory frameworks and the body says four, which we could not resolve. Their own mitigation is a prompt, and it recovers 14.9 percentage points on one framework. Nothing here says memory is a mistake.

Should this stand, we stop telling a client that bounding memory and testing its expiry covers the harm that memory can do.

How we Assure

Thirty-seven per cent of users have nothing that ranks their first value first

Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly and Moustapha Cisse submitted this on 1 September. They put two instruments on one scale. The first is a paraphrase-controlled audit of the as-shipped default settings of 23 frontier model archetypes, ranking five values against each other: safety, helpfulness, honesty, autonomy and equity. The second is a pairwise-tradeoff study of 1,649 US participants answering the same instrument.

Demand covers all five values, with the largest single constituency under a third. Supply does not. The 23 archetypes span about 2 per cent of the demand under conservative noise-matched estimation, and 0.10 per cent at full audit precision. No archetype puts helpfulness or autonomy first. That leaves 37 per cent of participants with nothing on the menu ranking their first value first. Across six model families autonomy falls in five, equity rises in five and safety rises in four. Within-family version trends are monotone at p = 0.013, and the autonomy decline sits in scenarios where safety is not at stake. A two-point menu beats the whole 23-archetype frontier by 47 per cent on mean regret, with a confidence interval of 43 to 52 per cent.

Our practice puts a quality floor under a firm’s agents. No number of extra agents may get past it. Somebody is named as holding it, and fleet health is reported as a distribution rather than as a total. All three describe the firm’s own agents. The ranking measured here arrives in the model as bought. It covers a fraction of what people want, and it is moving away from the value already least served. A firm can name a holder, report a distribution, and hold its floor exactly as we ask. The floor under its least-served users was set by which archetypes were on the menu, and the person we name holds something that does not reach the thing that moved.

The population is US-only, the five values are the authors’ own, and this is a preprint. The audit also reads default settings rather than deployed behaviour. Most firms put their own instructions and refusals in front of the model, and a reasonable person could hold that those move the floor further than the shipped default does. This work does not measure that, and it would be the first thing to check.

Where that is right, we stop telling a client that naming a holder for its quality floor means somebody can move it.