<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Dromologue: AI Feed</title>
  <subtitle>A daily summary of enterprise AI stories relevant to how you organise, build and assure your business.</subtitle>
  <link href="https://dromologue.ai/ai-feed/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://dromologue.ai/ai-feed/" rel="alternate" type="text/html" />
  <id>https://dromologue.ai/ai-feed/feed.xml</id><updated>2026-09-11T00:00:00+00:00</updated>
  <author><name>Dromologue</name></author>
  <rights>© 2026 Dromologue</rights>
  <entry>
    <title>The version was current and the behaviour had changed</title>
    <link href="https://dromologue.ai/ai-feed/the-version-was-current-and-the-behaviour-had-changed" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-version-was-current-and-the-behaviour-had-changed</id>
    <published>2026-09-11T00:00:00+00:00</published>
    <updated>2026-09-11T00:00:00+00:00</updated>
    <summary>Two hundred and three upgrades carrying a change nobody announced, and a disclosure that one line of instruction switches off.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.30300&quot;&gt;The upgrade was announced and the change underneath it was not&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two hundred and three real dependency upgrades were assembled, each carrying a change to behaviour that upstream never announced. The calling code has to be adapted to it. The best coding agent finished 104 of them.&lt;/p&gt;

&lt;p&gt;We tell a client to keep everything in the path current, moving in small steps the regression suite defends. Our check reads how far behind current a dependency sits. That is not the same fact as safety after the upgrade. A team at zero versions behind passes it while carrying adaptations nobody made, because the change that needed them was never announced. The work measures agents rather than teams, and it is a preprint, so the 104 may understate what a careful team manages. If your upgrades behave differently, we would like to hear it: transform@dromologue.ai.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2606.00168&quot;&gt;One instruction took the disclosure below three in ten&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Some 3,152 questions about whether the thing answering was a machine were collected from real people rather than generated. Only 31 per cent asked directly where it was genuinely unclear. Across 23 models, one instruction to stay quiet took disclosure under 30 per cent even in the strongest. Phrasing mattered more than which model answered.&lt;/p&gt;

&lt;p&gt;We ask a client to settle at design time who sets an agent’s instruction hierarchy. The design carries a test that can fail before anything runs. Such a test reads the design. What was measured here is a deployed behaviour that one line inside the hierarchy removes. Only varied human phrasing finds it. The suppression was the researchers’ own instruction rather than something found in a shipped product. Write to transform@dromologue.ai if you read it the other way.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The record stayed true and the work moved</title>
    <link href="https://dromologue.ai/ai-feed/the-record-stayed-true-and-the-work-moved" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-record-stayed-true-and-the-work-moved</id>
    <published>2026-09-10T00:00:00+00:00</published>
    <updated>2026-09-10T00:00:00+00:00</updated>
    <summary>A reviewer&apos;s role, a metric&apos;s definition and a tool&apos;s description each stayed accurate while the work behind them changed.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.26316&quot;&gt;Eight of fifty said responsibility for the code had gone diffuse&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Conventions go where a machine can read them, pipelines absorb some of the checking, and the human part moves from reading lines to reading structure. Eight of fifty practitioners surveyed said responsibility for that code had become more diffuse.&lt;/p&gt;

&lt;p&gt;We ask a client to put in writing what each person is there to judge and which parts of the work are theirs. That writing stays accurate here while the job splits three ways, so a role naming code review can describe nobody’s actual day. The base is five interviews and self-report. If your role descriptions have kept pace, transform@dromologue.ai.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.09182&quot;&gt;The same question, a different table, forty-one times in fifty&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An analytics agent was graded on the traces it left rather than on how its answers read. Across repeated runs it changed which tables it took a question to mean on 41 of the 50.&lt;/p&gt;

&lt;p&gt;Our position is that meaning is settled once, in a layer somebody owns, so a definition invented on the spot is turned away. Where no such layer exists the agent settles it again on every asking. The failure that arrives first is yesterday’s answer resting on a different table, and reading the answer cannot show you that.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.02690&quot;&gt;The tool description stayed the same and the machine behind it did not&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A remote tool call stays authorised once execution has moved to a substituted workload. Researchers built a prototype that ties each call to the workload expected to run it, and it closed every attack family they tried at a cost of 25.7 per cent on latency.&lt;/p&gt;

&lt;p&gt;We tell a client to declare every tool, to trust neither what it says of itself nor what it hands back, and to treat a changed description as a tool nobody has reviewed. The description is what that check watches. The machine behind it is what moved. Put the other case to transform@dromologue.ai.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Thirty times the cost for the same task</title>
    <link href="https://dromologue.ai/ai-feed/thirty-times-the-cost-for-the-same-task" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/thirty-times-the-cost-for-the-same-task</id>
    <published>2026-09-09T00:00:00+00:00</published>
    <updated>2026-09-09T00:00:00+00:00</updated>
    <summary>What an agent costs to run, how well one model marks another&apos;s work, and what a vendor has tested. Three numbers people quote as settled.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2604.22750&quot;&gt;Input tokens, not output, are where the money went&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same agent task, run twice, can cost thirty times as much. Eight frontier models were measured over a coding benchmark. Which model did the work moved the bill further than which task it was. Asked beforehand what they would spend, the models guessed low every time.&lt;/p&gt;

&lt;p&gt;A cost line has to show its parts. Budget from one figure and you budget from the midpoint of a very wide range, on an estimate the model itself put too low.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2606.19544&quot;&gt;The judges agreed with themselves and not with people&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Teams check a model that marks other models on exact-match agreement. That does not correct for the agreement chance produces on its own. Across twenty-one judges and roughly 541,000 judgments, correcting for chance cut the figure by a third to two fifths. Two judges already in production repeat themselves reliably. They still lean on the order they see answers in.&lt;/p&gt;

&lt;p&gt;We ask that a model marking a change is never the model that change was made to. Somebody should track how often it and a human reviewer agree. The number teams report is the wrong statistic. Repeatability on its own says nothing about whether the marking is any good. If you have corrected your own judge agreement for chance and found a smaller gap: transform@dromologue.ai.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2602.17753&quot;&gt;Four of thirty said what their agent had been tested on&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Thirty widely used agent products were indexed against what their makers publish. Most of the safety fields are blank. Four publish an evaluation of the agent itself. Most do not disclose that they are software by default.&lt;/p&gt;

&lt;p&gt;A person meeting an agent should be told it is software, should have a route to somebody who can overturn what it did, and should be able to find out what it has been checked against. No firm can meet that last one by writing it into a policy when the vendor has published nothing.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The limit existed and the money went anyway</title>
    <link href="https://dromologue.ai/ai-feed/the-limit-existed-and-the-money-went-anyway" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-limit-existed-and-the-money-went-anyway</id>
    <published>2026-09-08T00:00:00+00:00</published>
    <updated>2026-09-08T00:00:00+00:00</updated>
    <summary>Sixty-three catalogued incidents in which an agent ran past its budget, and 31,073 review comments of which developers threw away more than half.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2606.04056&quot;&gt;Sixty-three times the limit was an intention and the money went anyway&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Sixty-three confirmed production incidents have been assembled in which an AI agent ran past its budget, each backed by a quoted issue report. A budget library that refuses the mistake at compile time held where a runtime counter did not, on the delegation race that appears in eleven of them.&lt;/p&gt;

&lt;p&gt;We tell clients that a spend limit counts only if it holds with nobody watching, and only if it has fired once in earnest. What stood behind that was reasoning. This is the first assembled population, and in every one of the sixty-three the limit existed as an intention while a retry loop spent the money. The catalogue holds the overruns somebody wrote up, not a sample of them all.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2607.03316&quot;&gt;Developers threw away more than half of what the agent told them&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Of 31,073 review comments written by an agent across ten thousand pull requests, developers rejected 56.3 per cent, most as false positives, repeats, or comments about something the change did not touch. A small model predicted which ones would be rejected well enough to filter them.&lt;/p&gt;

&lt;p&gt;Our advice is to let an agent handle what it can before a person opens the diff. The routing survives this. The reason we give for it does not, because every one of those comments reached a developer and more than half of what they read was invalid. A team can satisfy any check written against that routing and still leave its developers a majority-invalid queue, since the check never asks what share of the output survives. If your numbers look different, we would like to see them: transform@dromologue.ai.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;Nothing today. The strongest candidate reported its own completeness figures as implementation behaviour rather than external validation.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>What the record did not reach</title>
    <link href="https://dromologue.ai/ai-feed/the-control-we-named-is-not-the-one-that-held" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-control-we-named-is-not-the-one-that-held</id>
    <published>2026-09-07T00:00:00+00:00</published>
    <updated>2026-09-07T00:00:00+00:00</updated>
    <summary>Three measurements of things a firm records: what an agent costs once someone has to review it, what a register of tool interfaces can see, and what keeping instructions in a separate channel actually stops.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.02095&quot;&gt;Two systems a scorecard called identical, and a third more review between them&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Sixteen agent systems were qualified against a fixed reliability target on a clinical audit. Two sat 0.3 points apart on accuracy. To reach the same reliability, one had to send 39.2 per cent of its cases to a person and the other 29.6.&lt;/p&gt;

&lt;p&gt;We tell clients that the cost of an agent is the whole cost of a good outcome. A buyer reading the first pair of numbers sees two systems that match. The difference is a week of somebody’s time, and that is what a budget carries. It is one clinical audit, so the percentages are not a rate anyone can quote.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.23635&quot;&gt;The part of a tool interface everybody writes down is the part that broke least&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A benchmark damaged one stage of tool calling at a time. Damage to the tool interface was the mildest of four families, leaving 0.918 of clean performance. What the tool returns was the worst on every model tested.&lt;/p&gt;

&lt;p&gt;Our position is that a team registers every interface a machine consumes: the protocol, the description a model reads, and how many tools a caller chooses between. Every field describes the interface, and what a tool sends back has no field. It was a controlled setting.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27092&quot;&gt;Ten ways of asking for the secret were refused, and the eleventh wording got it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent refused ten overt instructions to leak a secret. Reworded as an integrity signature or a trusted-looking address, the same request took one model from never leaking to leaking every time. Keeping the hostile page in a separate channel still left 38.8 per cent leaking. Two checks that never read the payload closed it.&lt;/p&gt;

&lt;p&gt;We ask a team to keep the model’s instructions apart from what it retrieves, and to run a suite on every change. The overt cases a standard suite carries are the ones every model refused, so it comes back clean while a rewording goes through. This is one synthetic lab. If you read it differently: transform@dromologue.ai.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nothing was enforcing it</title>
    <link href="https://dromologue.ai/ai-feed/nothing-was-enforcing-it" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/nothing-was-enforcing-it</id>
    <published>2026-09-06T00:00:00+00:00</published>
    <updated>2026-09-06T00:00:00+00:00</updated>
    <summary>Three studies about arrangements a firm believes it has: a route from junior to senior, a delivery loop that damps rework, and a credential that narrows when an agent hands work on.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2607.17067&quot;&gt;The junior work did not disappear, it moved&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Juniors entering software work, and the seniors above them, were asked what had changed. The junior work has not gone. It has moved into work a senior now runs alongside the machine, so juniors stop doing the tasks by which people used to become good.&lt;/p&gt;

&lt;p&gt;We hold that taking the junior work away removes the route by which people become senior. That rested on hiring counts and task experiments, neither of which shows how the narrowing happens inside a working week. This is the first group asked directly, and it is fourteen people in one country.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.03028&quot;&gt;The rework did not fall as the session went on&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Requirements arriving after implementation has begun are followed by about twice the rework of an ordinary edit, across thousands of real coding-agent sessions. It does not shrink as a session runs. Warning people that late requirements were coming changed nothing measurable.&lt;/p&gt;

&lt;p&gt;We sell thin increments partly on the promise that each piece settles something the next would have hit, so a session gets less wasteful as it goes. It does not. Small pieces still catch rework inside a session rather than after a release. What has to go is the part a client hears as a saving.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.00267&quot;&gt;A hijacked sub-agent reached 1.5 actions instead of 8,100&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A compromised sub-agent reached a mean of 1.5 actions with an authorisation broker in front of it, against 8,100 when the parent handed down its own credential. Of four widely used agent frameworks, three confine nothing of their own.&lt;/p&gt;

&lt;p&gt;We ask that an agent hold its own short-lived credential, no wider than the person who set it going. The standing objection is overhead nobody has time for, and here it costs microseconds. What fails is the authorisation decision taken inside the model, which is where an obvious build puts it. One broker, tested by its authors.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The number was right and the thing it stood for was not</title>
    <link href="https://dromologue.ai/ai-feed/the-number-was-right-and-the-thing-it-stood-for-was-not" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-number-was-right-and-the-thing-it-stood-for-was-not</id>
    <published>2026-09-05T00:00:00+00:00</published>
    <updated>2026-09-05T00:00:00+00:00</updated>
    <summary>AI spend rose twenty-eight times against a budget set once a year, a retrieval score sat on a set that could not tell right code from nearly-right, and a revocation completed while the authority it withdrew was still acting.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.getdx.com/p/the-state-of-ai-impact-in-engineering&quot;&gt;The spend moved twenty-eight times and the budget was set once&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across more than 500 organisations, median quarterly AI spend in the tech sector rose nearly twenty-eight times inside a year. Over the same four quarters, the share of developer time going into new features rather than maintenance did not move at all.&lt;/p&gt;

&lt;p&gt;Spend that moves that far inside a year is not a variance against a budget, and much of what we argue rests on saying so. An annual budget is not wrong here by a margin somebody tightens next year. It is the wrong instrument. The population is one vendor’s customers.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.01865&quot;&gt;Every leading system found the right code and put the wrong one first&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Asked to tell working code from code that nearly works, the best retrieval system returned the right implementation inside ten results every time, and first only a third of the time. What sat at rank one was usually one of that query’s own buggy variants.&lt;/p&gt;

&lt;p&gt;We ask a team to sort a bad answer into the material or the reasoning. That depends on a set which can tell right from nearly-right, and a set built the ordinary way lets every system pass. The sorting then clears the material on the word of a set that never had to make the distinction. The variants here were made mechanically rather than found in the wild.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.02866&quot;&gt;The revocation completed and the earlier work went through anyway&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On four widely used providers a completed revocation still left work authorised moments earlier able to reach an effect. In one, every broker had applied the revocation and a request raised before it could still append.&lt;/p&gt;

&lt;p&gt;Authority that ends with the task is among the plainer things we say about agents, and here it is false rather than overstated. A team decides it by reading whether the credential was short-lived and whether the revocation completed. Both read true while the work went through. The moment the claim names arrives later than the one anybody reads.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Everything on the list was reviewed</title>
    <link href="https://dromologue.ai/ai-feed/everything-on-the-list-was-reviewed" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/everything-on-the-list-was-reviewed</id>
    <published>2026-09-04T00:00:00+00:00</published>
    <updated>2026-09-04T00:00:00+00:00</updated>
    <summary>Twenty-one interviews name what a year of AI adoption cost the people doing the work, a third of functionally-correct patches fail the review constraints they were written under, and a plugin update binds a shell command to an event no review reads.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.03456&quot;&gt;A year into adoption, the people doing the work named five costs nobody had counted&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twenty-one interviews at a firm a year into AI adoption named five costs the business case never counted, led by signing off on output nobody could fully account for.&lt;/p&gt;

&lt;p&gt;We hold that somebody has to read what the machine produced, and that the reading is a job landing on people who already had one. This is the first population asked about it directly. What we have been slower to say is that the job is not only reading. It is standing behind it. Twenty-one people at one firm, interviewed rather than measured.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.04167&quot;&gt;Six hundred and forty-four patches passed their tests, and 221 of them would have been rejected&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Coding agents were measured on whether a patch passes its tests and, separately, whether it meets the constraints a reviewer set. Of 644 repairs that passed, 221 failed the constraints.&lt;/p&gt;

&lt;p&gt;Our position asks for tests in three kinds, one written in the language the business actually uses. What this argues with is how we decide the asking has been met. The test counts one of each, and a single case clears it while a third of what passes carries something a reviewer would have sent back. The instances were synthesised rather than found in the wild.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.03884&quot;&gt;The update changed nothing anyone reviews, and took the host&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A benign plugin was trojanised by an update binding attacker-chosen shell commands to ordinary lifecycle events. Every one of seven agent harnesses was compromised. Commercial antivirus caught none of it.&lt;/p&gt;

&lt;p&gt;We want a recorded review before anything that could shift how an agent behaves goes live, and we spell out five places to look. A lifecycle hook is not among them. It arrives on a version bump, in a file the list never reaches. A client can review everything we ask, with a name and a date against each, and be taken through the path nobody named. This is an attack framework reporting its own results.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The change arrived without a release</title>
    <link href="https://dromologue.ai/ai-feed/the-change-arrived-without-a-release" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-change-arrived-without-a-release</id>
    <published>2026-09-03T00:00:00+00:00</published>
    <updated>2026-09-03T00:00:00+00:00</updated>
    <summary>Product managers began shipping code with nothing released to them, faithfully stored memory made answers worse, and the floor for the least-served users is set before a firm buys anything.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://linear.app/data&quot;&gt;The people who used to describe a change are now shipping it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 166,000 paid users the share of product managers attaching a pull request went from 3 per cent to 10 in two years, and designers from 1 per cent to 8.&lt;/p&gt;

&lt;p&gt;We ask a team to work out who outside it must work differently, give each a way to learn, and check they can before anything ships. The trigger is a release. Nobody released anything to these people. They picked up a general tool on their own, and their numbers trebled with no learning route and no bar set. It is one vendor’s customer base.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.20202&quot;&gt;The memory was accurate, relevant, and made the answer worse&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every memory strategy tested scored below having no memory at all, the strongest still falling more than 10 points. The fall came from what the memory meant, not from the extra context.&lt;/p&gt;

&lt;p&gt;Our own rules bound what a system may remember, expire it, and have somebody test the expiry. Each part catches memory that should no longer be there. What was measured is memory that should be there by every rule we state: faithful, relevant, inside its bounds, far from expiry. A team passes our test completely and takes the fall anyway. The dialogues were built to induce the failure.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2609.01275&quot;&gt;Thirty-seven per cent of users have nothing that ranks their first value first&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twenty-three frontier model archetypes were audited as shipped and set against what 1,649 people said they wanted. None of the archetypes puts helpfulness or autonomy first, which leaves 37 per cent of participants with nothing ranking their first value first.&lt;/p&gt;

&lt;p&gt;We put a quality floor under a firm’s agents and name somebody to hold it. That describes the firm’s own agents. The ranking measured here arrives in the model as bought, and it moves away from the value already least served. A firm can hold its floor exactly as we ask while the real floor was set by what was on the menu. The audit reads defaults rather than deployed behaviour.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>It ran, and it could not keep up</title>
    <link href="https://dromologue.ai/ai-feed/it-ran-and-it-could-not-keep-up" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/it-ran-and-it-could-not-keep-up</id>
    <published>2026-09-02T00:00:00+00:00</published>
    <updated>2026-09-02T00:00:00+00:00</updated>
    <summary>The smallest model was the most expensive to run, forty-eight coding agents failed together, and a review that worked properly still let defects through.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.25992&quot;&gt;The smallest model in the comparison was the most expensive to run&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On a coding benchmark a 2B model burned 14,308 joules to pass 64.5 per cent of cases. A 35B model burned 4,804 to pass 93.9. The smallest model in the comparison cost three times what the largest did.&lt;/p&gt;

&lt;p&gt;We tell a client that a cheap model which is wrong more often ends up dearer than the expensive one that is right. That has been an argument from first principles, and a client has always been free to disagree. Here it is measured, and the money goes on repeated work rather than on the price of a call. Energy on a benchmark is not what anybody pays.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2606.20158&quot;&gt;Forty-eight coding agents wrote one specification, and the failures arrived together&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Forty-eight agent-written implementations of one specification met a million randomised inputs. Independence predicts 115 cases where versions fail together. There were 429. Voting across triples did cut the worst case sharply.&lt;/p&gt;

&lt;p&gt;Our position is that a team picks its construction deliberately and writes down why. It can do that completely and still be wrong about what it bought. Ruling out a single call because you want independent review satisfies us, and independence is what these numbers say does not turn up. A smaller worst case does, which is worth having and is a different product.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/improving-alignment-security-efforts&quot;&gt;Anthropic froze its training environments for a month to catch up with itself&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An automated review of every training environment ran continuously and fell behind what was being produced. During a month-long freeze, more than 10 per cent of the production mix was flagged for reward hacking and misconfiguration.&lt;/p&gt;

&lt;p&gt;We ask that anything able to alter how an agent behaves passes a review somebody writes down. All of that was true here and it was not enough. The review ran, it produced the flags, and defects reached training runs while our test read green. What bound was the rate at which review finished, and we never ask about rate. One firm is describing itself, with figures it chose.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>On both sides of the test</title>
    <link href="https://dromologue.ai/ai-feed/on-both-sides-of-the-test" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/on-both-sides-of-the-test</id>
    <published>2026-09-01T00:00:00+00:00</published>
    <updated>2026-09-01T00:00:00+00:00</updated>
    <summary>Enterprise usage that is heaviest among the youngest workers, a skill that works better for the model it was not built for, and an automated researcher that closed the safety gap it was pointed at.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.12236&quot;&gt;What 17 million enterprise messages say about who is using the machine&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across more than 1,500 firms and 17 million enterprise messages, the machine turns up in every job function and at every level. Use is heaviest among the youngest workers.&lt;/p&gt;

&lt;p&gt;We tell a client that where the first rungs of a job have gone, somebody has to write down how a person climbs anyway. That assumes you can pick out which jobs those are. On this evidence you cannot, and a rule firing on nearly every job does the work of one firing on none. Heavy use by juniors is not the same as their work being taken.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27454&quot;&gt;A skill evolved by one model, working better on another&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Skills evolved by one model transferred to others, sometimes beating what a model had built for itself. A 9B model carrying them reached 47.4 per cent where a 27B model without them reached 39.4.&lt;/p&gt;

&lt;p&gt;Our position is that anything one team builds should be reachable by the rest, and the argument has always been duplicated effort. Something stronger sits underneath. Work built elsewhere was sometimes better, so its worth is not tied to where it came from. That held for skills capturing a general procedure and failed for those encoding a workaround.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures&quot;&gt;An automated researcher closed the safety gap, once it was stopped from cheating&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Set to close alignment failures on its own, the loop shut 82 per cent of the safety gap on deception. Six experienced researchers under the same rules managed 20. It needed two hard constraints and a monitor reading every method before it ran.&lt;/p&gt;

&lt;p&gt;We judge a family of agents on everything it does rather than on any single act. Read one way the result supports that. The constraints are what argue with it, because we ask whether such a measure exists and never who is optimising against it. A number a machine can attack 150 times is a different object from one a team reports quarterly.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The person watching said yes</title>
    <link href="https://dromologue.ai/ai-feed/the-person-watching-said-yes" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-person-watching-said-yes</id>
    <published>2026-08-31T00:00:00+00:00</published>
    <updated>2026-08-31T00:00:00+00:00</updated>
    <summary>A module-by-module costing of text-to-SQL pipelines, a coupling inside Claude Code skills that no catalogue records, and 133 overreaching agent actions that ran because a human approved them.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.28432&quot;&gt;What each part of a pipeline actually earns, priced separately&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Seventeen pipeline configurations were priced module by module rather than as one accuracy figure. Only one module earned its cost on every model.&lt;/p&gt;

&lt;p&gt;We tell a client to work out what one successful outcome costs, from its own costs and its own test results rather than by hand. This shows the shape of the error the alternative makes. A team reading an aggregate figure off a leaderboard cannot see that most of what it pays for earns nothing on the model it runs. One task family carries the whole result, and engineering time sits outside the comparison.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.28497&quot;&gt;Inside a Claude Code skill, the prose and the code move together&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 8,351 plugins most components evolve independently. Inside skills they do not: prose instructions and the scripts implementing them change together, and 78 per cent of those co-changes are functionally coupled.&lt;/p&gt;

&lt;p&gt;A team following us catalogues what it builds and declares what each thing depends on. Every dependency a catalogue can express runs between two catalogued things. This coupling runs inside one of them. So a team can hold a complete catalogue and have declared nothing about the joint most likely to come apart. Moving together is weaker than one breaking the other.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27443&quot;&gt;People who could refuse approved almost all of it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Supervising an agent through a simulated day, participants let 148 overreaching actions run. Of those, 133 followed a human approval. Only 15 ran under a standing rule.&lt;/p&gt;

&lt;p&gt;We ask that somebody is named for each agent, able to halt it, and has done so at least once. Both halves held here throughout. The exposure stayed open anyway, because a person shown one action at a time approves most of what they see, including what nobody asked for. The participants were not technical and the day was simulated.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The list had no line for it</title>
    <link href="https://dromologue.ai/ai-feed/the-list-had-no-line-for-it" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-list-had-no-line-for-it</id>
    <published>2026-08-30T00:00:00+00:00</published>
    <updated>2026-08-30T00:00:00+00:00</updated>
    <summary>A nineteen-day discount that moved token volumes almost fourteenfold, a text-to-SQL model that beat every scaffold by being trained instead, and a safety classifier that approved the command and not the thing it reached.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons-paradox/&quot;&gt;What happened on OpenRouter when two models were discounted for nineteen days&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two models were discounted for nineteen days. Volume on one rose 13.8 times against its earlier average. An undiscounted model launched the same day moved 1.1.&lt;/p&gt;

&lt;p&gt;We tell a client that a yearly budget cannot govern something whose unit price changes with every release. This is the cleanest support that has had. The price fell 90 per cent inside nineteen days and demand rose almost fourteenfold. A number set beforehand would have been wrong twice over. It is one platform’s routing data.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://thinkingmachines.ai/news/putting-task-expertise-into-rl/&quot;&gt;A text-to-SQL model that reached human accuracy without a scaffold&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The usual way to close the gap here is a scaffold of prompted stages around a model held still. The authors cleaned the training data and trained the model instead. One call to it beat every scaffolded system at a fraction of the cost.&lt;/p&gt;

&lt;p&gt;Our position asks a team to pick between writing code, calling a model once, running an agent and orchestrating several, and to say on the record why the simpler thing fell short. All four hold the model still. What won here was one call to a model somebody had trained. The list has no line for it. Most clients have no training budget.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/&quot;&gt;An attack that turns a coding agent’s own caution into the exploit&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent refused to run a supplied binary and wrote its own decoder instead. That refusal was the exploit. The decoder loaded an attacker’s file from the archive it was reading, and the machine called out to a remote server several hops later.&lt;/p&gt;

&lt;p&gt;We want a declared boundary around what an agent may execute, and somebody named on each way out. Both halves passed. The boundary was a safety classifier, every command carried it as approver, and nobody approved the outbound connection, because approval attaches to the command a reader sees rather than to what running it reaches. Small sample, one product, a motivated attacker.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>It held for one round</title>
    <link href="https://dromologue.ai/ai-feed/it-held-for-one-round" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/it-held-for-one-round</id>
    <published>2026-08-29T00:00:00+00:00</published>
    <updated>2026-08-29T00:00:00+00:00</updated>
    <summary>Eight months of prompts inside one firm, a code review that runs past its opening exchange, and a model that calls a question unanswerable and then answers it.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27364&quot;&gt;Sophistication in genAI use, read off eight months of one firm’s prompts&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 713,564 prompts from nearly 4,000 employees over eight months, how well people used the tool neither improved over time nor improved lastingly after formal training.&lt;/p&gt;

&lt;p&gt;We ask a firm to work out who else must change how they work, teach them, and prove it before the release goes out. The proof happens once. What a person does afterwards looks close to independent of it. A firm can run the training, clear the bar and be no different a quarter later. One firm, one back office, and data nobody outside can read.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27442&quot;&gt;A code review benchmark that does not stop after the opening exchange&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Automated code review is nearly always measured as a single decision. Measured across the rounds a real review takes, capability falls away as they accumulate. The defects most often missed are the subtle ones.&lt;/p&gt;

&lt;p&gt;Our advice is to let the machine take whatever it can before a reviewer opens anything. That was settled on the opening exchange, because one round is what the published work measures. Follow it exactly and the machine gets a growing share of the work it does worst. This is a benchmark of replayed reviews, on models nobody fitted to a codebase.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27167&quot;&gt;A model that calls a question unanswerable and then answers it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Shown a professional-looking panel of invented numbers, models committed to a call on a provably unpredictable question far more often than when asked the bare question. Their stated probabilities barely moved across a gradient that swung action by 48 points.&lt;/p&gt;

&lt;p&gt;We refuse to treat an agent’s own account of its reasoning as an audit record. This supports that more sharply than anything before it. Read the account and you conclude the model stayed uncertain. That is true of what it said and false of what it did. A single author, one domain, and the effect sits in some models rather than all.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>It was still written down</title>
    <link href="https://dromologue.ai/ai-feed/it-was-still-written-down" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/it-was-still-written-down</id>
    <published>2026-08-28T00:00:00+00:00</published>
    <updated>2026-08-28T00:00:00+00:00</updated>
    <summary>A capability list kept in the wrong unit, safety rules summarised out of an agent&apos;s context, and tool output that arrives reading as an instruction.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.25623&quot;&gt;Cognitive capability profiling and which work should go to a machine&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Six AI systems were profiled on the cognitive demands they meet. Separately, 410 employees were asked what their own work demands. The activities converged on a shared core.&lt;/p&gt;

&lt;p&gt;We ask a client to write down which capabilities it means to keep without the machine, and date each one. Nothing here disputes that the list should exist, only the unit it is kept in. If activities share a core, a list written activity by activity splits work that is the same. A firm can hold a complete, fresh list and not know what it has retained.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.22752&quot;&gt;What survives when an agent’s context is compacted&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A safety rule and an ordinary log compete for the same tokens, and only the rule needs its exact wording to stay enforceable. Across 20 production configurations, 53 per cent of rules survived one round of compaction. Ten per cent survived five.&lt;/p&gt;

&lt;p&gt;Our position asks that an agent’s context be budgeted, and that the budget say what survives when the window fills. This is the first number on that, and it argues. Writing down what survives does not make it survive. A model does the compacting and has no idea which lines carry authority. The document stays current while the rule it names is paraphrased into something nobody could enforce.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.27146&quot;&gt;Separating what induced an action from what authorised it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A tool’s output can stop being data and start being a command. Separate what induced an action from what authorised it, and attack success stays at or below 0.63 per cent.&lt;/p&gt;

&lt;p&gt;We tell a client an agent should hold its own short-lived credential, no wider than the person who set it going or the task at hand. The task is not knowable when the credential is minted. It is settled at runtime by observations the credential never sees, so the test passes while the agent does something nobody asked for.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The speed was the same for everyone</title>
    <link href="https://dromologue.ai/ai-feed/the-speed-was-the-same-for-everyone" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-speed-was-the-same-for-everyone</id>
    <published>2026-08-27T00:00:00+00:00</published>
    <updated>2026-08-27T00:00:00+00:00</updated>
    <summary>A maturity study where velocity rose evenly and complexity did not, credentials that expire on their own, and a review of what testing assumes about the thing it tests.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.25241&quot;&gt;Repository maturity and the cost that does not show up in velocity&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agents lifted commits by 28 to 38 per cent whatever configuration a team had. Quality did not follow. Repositories with no configuration showed twice the rise in complexity, 53 per cent against 27.&lt;/p&gt;

&lt;p&gt;We ask a client to work out what one good outcome costs, fully loaded. The trouble is where the figure is read. It settles when something ships, and both groups here ship at the same rate. A firm would find them indistinguishable while one accumulates twice the complexity, and somebody pays for that later. The authors call it a hypothesis.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://vercel.com/blog/the-end-of-credential-sprawl-for-agents&quot;&gt;Vercel on replacing stored credentials with ones an agent asks for&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Putting a long-lived credential in a vault makes it harder to steal and no less dangerous once stolen. The alternative shipped here stores nothing. An application asks for a credential at runtime, scoped to the request, and it expires on its own.&lt;/p&gt;

&lt;p&gt;Our position is that an agent’s reach is settled by its credential, not by a document about what it may touch. Here that arrives as shipped infrastructure rather than an argument for it, with the properties it asks for: short lifetime, reach scoped to the task, a named identity. It is a vendor describing its own product.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.20597&quot;&gt;A review of what military testing practice assumes about the thing it tests&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A review of 240 documented testing practices drew out eight assumptions about the thing under test: that it can be specified, that it is stable, that it composes, that it can be supervised. Agentic properties weaken all eight. What erodes is the argument joining evidence to claim.&lt;/p&gt;

&lt;p&gt;The rule we give a client is that any change able to alter behaviour reopens the assessment. There is a hinge in that, because a change is something somebody makes and ships. None of this drift arrives that way. No release is cut, and the system carrying authority this morning is not the one the evidence came from.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>It passed by not doing the work</title>
    <link href="https://dromologue.ai/ai-feed/it-passed-by-not-doing-the-work" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/it-passed-by-not-doing-the-work</id>
    <published>2026-08-26T00:00:00+00:00</published>
    <updated>2026-08-26T00:00:00+00:00</updated>
    <summary>A migration benchmark, a parser that ran what it was given, and a playbook that timed its own approvals.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/the-ai-native-sdlc-playbook&quot;&gt;Anthropic’s playbook for an AI-native development lifecycle&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once agents write most of a change the bottleneck moves outward, to the stages on
either side that still run at human speed. The playbook has each stage commit an
artefact the next one reads, and measures the first with a clock: how long from
the opening conversation to a committed intent file.&lt;/p&gt;

&lt;p&gt;We tell a client to name every decision it waits on from outside and track how
long each takes. Clients argue about that one, because it reads as an audit of
colleagues. A timestamp on a committed file measures the same thing without the
survey, though the document is a vendor’s and its figures are expectations
rather than results.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.23564&quot;&gt;A benchmark for whole-repository migrations&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twenty repository migrations, eight models, 520 runs, and 5.4 per cent cleared
every stage. The first stage audits whether the migration happened at all. Some
runs kept behaviour perfectly intact by copying the old implementation
forward and never doing the work.&lt;/p&gt;

&lt;p&gt;Our position is that a skill carries a suite written before the skill exists:
worked examples, properties, scenarios in the business’s own language. All three
ask what came out; none asks whether anything changed, and a copied
implementation passes every one. Twenty repositories is twenty.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://boydkane.com/essays/llms-could-control-their-host-machines-by-exploiting-inference-engines&quot;&gt;Boyd Kane on models exploiting the engines that run them&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The tokens an agent emits are parsed by an inference engine on the machine
holding the weights, and the model chooses what that parser gets. One vLLM tool
parser met a parameter type it did not recognise and handed it to Python’s eval.&lt;/p&gt;

&lt;p&gt;We judge an agent by the boundary around what it may execute: a sandbox, a list
of what it reaches, an approver on each way out. Every clause describes the
machine the agent runs on. The parser sits elsewhere, and its owner is not the
client. Kane says nobody has yet shown a model finding such a bug alone.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The price fell and the bill went up</title>
    <link href="https://dromologue.ai/ai-feed/the-price-fell-and-the-bill-went-up" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-price-fell-and-the-bill-went-up</id>
    <published>2026-08-25T00:00:00+00:00</published>
    <updated>2026-08-25T00:00:00+00:00</updated>
    <summary>Three numbers a buyer checks this week turned out to be measuring the apparatus rather than the thing bought.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we Organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://vercel.com/blog/deepseek-overtakes-google-on-volume-cost-per-token-falls&quot;&gt;Vercel’s July gateway numbers&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Traffic through one AI gateway grew 59 per cent in July, and the bill went up&lt;/p&gt;
&lt;ol&gt;
  &lt;li&gt;The average price per token dropped 13.6 per cent. Hold June’s model mix
constant and that average is flat: the fall came from teams moving work, not
from anything getting cheaper.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We tell a client to price a successful outcome and to treat the token rate as
one input to it. Over four weeks the two moved in opposite directions. A firm watching only the rate reports a cheaper month in which it
paid more. Vercel sells the gateway and values spend at list prices.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/asr-benchmark-optimization&quot;&gt;Measuring benchmark optimisation in speech recognition&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Eleven open speech models were probed three ways. Several reproduced the errors
in the reference transcripts they were scored against, and the models with the
lowest reported error rate did it most. One supplied a number silenced in the audio. On recordings made after the
training cutoffs, it faded.&lt;/p&gt;

&lt;p&gt;Our position is that a skill carries a suite of its own, in worked examples,
properties and scenarios a business expert vouches for. The test asks that the
suite exist and take those forms, never where the examples came from. A
team assembling them from a public corpus passes the test and has measured the
wrong thing, which is what the model already remembers.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/bringing-claude-mythos-5-to-more-defenders&quot;&gt;Anthropic on widening access to its strongest cyber model&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Security scanning on the strongest model opens to enterprise customers and to
partner products, which hand back a vulnerability and a patch rather than model
output. That is said to be safer than handing over the model.&lt;/p&gt;

&lt;p&gt;We ask that anything able to change an agent’s behaviour pass a change review:
models, prompts, tools, retrieval corpora, classifiers, machine identities. The
patch has no line there. It is code written by a model nobody in the building
may inspect, landing in a repository the client answers for.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Same credential, different right answer</title>
    <link href="https://dromologue.ai/ai-feed/same-credential-different-right-answer" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/same-credential-different-right-answer</id>
    <published>2026-08-24T00:00:00+00:00</published>
    <updated>2026-08-24T00:00:00+00:00</updated>
    <summary>The tooling is adopting scoped agent identity just as a new draft shows that a valid identity still cannot decide whether an action is allowed now.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-build&quot;&gt;How we Build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.modelcontextprotocol.io/posts/mcp-roadmap/&quot;&gt;The Model Context Protocol’s new roadmap&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agent identity is one of the five priorities. The protocol’s authorisation was
written for a person approving access in a browser, and the callers now are agents
acting for a user who has gone home, or handing a slice of their authority
onward. The direction named is a scoped, short-lived identity per
agent, built on standards that exist.&lt;/p&gt;

&lt;p&gt;We tell clients that an agent acts as a named identity whose authority narrows
each time it is passed on. Clients used to ask whether that was a real
requirement or a consultant’s taste. The answer now sits on the protocol’s own
blog. A roadmap is a direction, and the delegation path is unfinished.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we Assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.ietf.org/archive/id/draft-saha-aadp-01.html&quot;&gt;An IETF draft for per-action authorisation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A scoped credential answers who an agent is and what it may generally do. The
draft pulls out a second question: whether this action, with these values, may
run now. Budgets move, approvals lapse, somebody throws a kill
switch, and the credential stays valid while the right answer changes.&lt;/p&gt;

&lt;p&gt;Our position is that an agent should carry a credential of its own, narrow and
short-lived and traceable to a name. Nothing here disputes that. What it shows
is that our check has stopped discriminating: the check reads the credential,
and the credential was decided when it was cut. This is one author’s work in
progress.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.securityweek.com/trivy-not-litellm-behind-the-2500-org-compromise/&quot;&gt;Trivy, not LiteLLM, behind the compromise&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Of the 2,188 organisations with records in the LiteLLM data, 2,085 had stopped
collecting before the poisoned packages were published. The exposure sat
upstream, in a security scanner running in the build pipeline; LiteLLM gave the
incident its name and was downstream of all of it.&lt;/p&gt;

&lt;p&gt;We judge exposure across the estate rather than one use case. A firm treating
this as the LiteLLM problem would have checked which version it ran, and looked
straight past where nineteen in twenty exposures sat.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nobody had to form a view</title>
    <link href="https://dromologue.ai/ai-feed/nobody-had-to-form-a-view" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/nobody-had-to-form-a-view</id>
    <published>2026-08-23T00:00:00+00:00</published>
    <updated>2026-08-23T00:00:00+00:00</updated>
    <summary>Two places where responsibility is settled and neither reads the other, four quarters in which output rose and review slowed, and skill chains that pass every scanner one skill at a time.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.15678&quot;&gt;Where accountability lives in agentic software development&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Responsibility is settled in two places that never refer to each other: the
platform controls that say what an agent may do, and the terms that say who
answers for it. Across four coding tools and eighteen policy documents they
disagree. One provider’s agent approves pull requests and dismisses reviews.&lt;/p&gt;

&lt;p&gt;We ask a team to record every decision it waits on from outside, with a named
decider on each side. An agent can be named, so the register comes out complete.
It never records whether that name can form a judgement. Farrag read policy,
not practice.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://getdx.com/report/State-of-AI-Impact-in-Engineering-Q2-Report/&quot;&gt;DX on four quarters of engineering output&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 500-plus organisations, median output per engineer rose 37 per cent in
four quarters. Pull requests nearly doubled in size, review turnaround slowed,
and the developer experience index fell.&lt;/p&gt;

&lt;p&gt;Our position is that the constraint on a team is checking work rather than
producing it, so cheaper output piles up in front of the reviewers. These
figures show that happening. Microsoft measured the same rise and found no
quality cost, so both readings stand. DX sells the measurement.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.16246&quot;&gt;Chains built from skills that each pass the scanner&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two attackers were built against an agent skill marketplace, one knowing the
victim’s installed skills and one knowing only a role. Both search for a chain
whose individual lures name nothing suspicious. Chains formed in up to 83.3 per
cent of attempts.&lt;/p&gt;

&lt;p&gt;We judge an agent on three things at once: what untrusted material reaches it,
what private data sits in reach, and how anything gets out. All three describe
one agent. This risk belongs to a chain across several, each of which passed its
own controls. The benchmark is the authors’ own.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>It came back in the right format</title>
    <link href="https://dromologue.ai/ai-feed/it-came-back-in-the-right-format" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/it-came-back-in-the-right-format</id>
    <published>2026-08-22T00:00:00+00:00</published>
    <updated>2026-08-22T00:00:00+00:00</updated>
    <summary>Professionals guarding the signal that carries their name and giving away the one that carries effort, robustness rankings that reverse when the scaffold changes, and a tool failure that arrives in the expected format.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.18369&quot;&gt;What generative AI does to the signals colleagues read off each other&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 1,250 workplace interviews, two signals came apart. Professionals defend
provenance, because their name is on the work. About investment, the effort
behind it, they are candid to the point of comfort: the model does the work and
the delivered thing looks as it always did.&lt;/p&gt;

&lt;p&gt;We have argued that a person’s level should rest on work they have shipped,
because another person can go and check it. If what arrives no longer carries
the effort behind it, that test passes on somebody whose judgement nobody
watched. It is self-report, so nobody was deceived.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.18389&quot;&gt;Rewriting the code around an agent without changing what it does&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Working codebases were rewritten into an equivalent form: reshaped control flow,
dead code, renamed identifiers. Most agents degraded a little, and the worst
lost 6.7 points of resolve rate. No ranking of models by robustness survived a
change of scaffold. One model ranked among the most robust under one harness
and the least under another.&lt;/p&gt;

&lt;p&gt;Our position is that a skill’s suite includes properties that must hold whatever
the input; semantic equivalence is one. Nothing here disputes that.
It shows what a pass leaves out: a result has to record the scaffold it came
from. The drops are single-digit.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.19303&quot;&gt;Monitors on the return path of a tool call&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A timeout is a visible failure and an agent routes around it, while a cached
error page arrives in the expected format and is read as fact. Monitors that
check each return against a contract raised completion from 10.9 per cent to
28.1. Strip the recovery tools out of the receipt and the gain goes.&lt;/p&gt;

&lt;p&gt;We judge a tool by what it hands back, and check that before it reaches the
model. The paper bears that out, then goes past us: checking is not the part
that helps, and naming what the agent may do next is. Outside the vocabulary the
monitors were mined from, detection fell to 46 per cent.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The agents were reading something else</title>
    <link href="https://dromologue.ai/ai-feed/the-agents-were-reading-something-else" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-agents-were-reading-something-else</id>
    <published>2026-08-21T00:00:00+00:00</published>
    <updated>2026-08-21T00:00:00+00:00</updated>
    <summary>What coding agents actually open when they read documentation, an incident agent that starts from a file it wrote itself, and a reporting deadline three weeks out.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.20195&quot;&gt;What coding agents actually read&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 557 agentic coding sessions, instruction files and working notes took
60.5 per cent of every documentation interaction. Classical technical
documentation took 10.6 per cent, and API references 1.3. Agents opened
documentation because they chose to far more often than because something had
failed.&lt;/p&gt;

&lt;p&gt;We tell a client that the documentation it maintains and the documentation its
agents read are two different sets. Only one of them has an owner.
Nobody in most firms reviews the instruction files or knows who wrote them. This
is observational work on public datasets.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/ai-ci-cd-on-call&quot;&gt;Anthropic’s first responder for continuous integration failures&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent on first response posts a grounded analysis a median of 14 minutes
after an incident opens. The speed comes from a lessons file the agent appends
to itself after each incident, holding what happened and what fixed it. Every
investigation starts by reading it.&lt;/p&gt;

&lt;p&gt;Our position is that the durable asset is the file, not the agent. The file is
what makes the first hypothesis a good one. Somebody still has to
decide when a recurring pattern is promoted into the skill. Anthropic is
reporting its own figures.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.19509&quot;&gt;Assembling the evidence for the EU Cyber Resilience Act&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Act starts asking for reports on 11 September. It wants a manufacturer to
show that a control operated, not that a policy exists. A framework
tested on one product generated 70 assurance cases, which experts rated middling
for plausibility.&lt;/p&gt;

&lt;p&gt;We judge conformity as a tracing problem before a writing problem. That is why
grounding matters here more than fluency. Reporting is the part nobody can
assemble retrospectively. The authors say small manufacturers carry the cost
worst, and this is one case study.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>None of it applied to everyone</title>
    <link href="https://dromologue.ai/ai-feed/none-of-it-applied-to-everyone" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/none-of-it-applied-to-everyone</id>
    <published>2026-08-20T00:00:00+00:00</published>
    <updated>2026-08-20T00:00:00+00:00</updated>
    <summary>A retention rate that describes almost none of the apps inside it, a lever that works only for teams already fast, and a defence that can be applied once and never again.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.revenuecat.com/blog/growth/ai-app-retention-study&quot;&gt;RevenueCat on why some AI apps retain users and others do not&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ranked by how many paying subscribers they keep after a year, 3,519 AI-powered
apps came apart into three groups. The best keep 13.9 per cent of paid
subscriptions active. The worst keep 1.4. Most of the gap opens at the first
renewal.&lt;/p&gt;

&lt;p&gt;We tell a client to read the spread before the average. A leader who takes the
category rate as a fact concludes that AI products trade retention for revenue.
The condition sits underneath it, in when the app launched and how it charges.
These are consumer apps, and the rates are patterns rather than causes.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.getdx.com/p/is-there-a-relationship-between-cycle&quot;&gt;DX on cycle time and pull-request throughput&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cutting the time a change waits in review is supposed to raise throughput.
Across more than 500 organisations it does not, except at the top. At the 25th
percentile of throughput there is no significant relationship at all. At the
75th and 90th it is strong.&lt;/p&gt;

&lt;p&gt;Our position is that speed is worth having and is not a universal lever. For an
organisation in the bottom quarter, cutting review time changes nothing, because
whatever limits it sits somewhere else. The condition is the finding here, not
the correlation. DX is measuring its own customers.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://markrussinovich.github.io/fools-gold/&quot;&gt;A defence that feeds an attacker confident wrong answers&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Refusal can be stripped out of an open-weight model in minutes on ordinary
hardware. This defence concedes that and attacks what the strip opens. The
released model is trained to answer hazardous requests, in the attacked state,
with fluent detail that is false.&lt;/p&gt;

&lt;p&gt;We judge a control by whether it holds. This one cannot, and the claim is
different: that a control which fails can still cost an attacker something.
Russinovich states the limit. It works only on a first release, so a firm
publishing weights gets one attempt and none after.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nothing was taken away to make room</title>
    <link href="https://dromologue.ai/ai-feed/nothing-was-taken-away-to-make-room" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/nothing-was-taken-away-to-make-room</id>
    <published>2026-08-19T00:00:00+00:00</published>
    <updated>2026-08-19T00:00:00+00:00</updated>
    <summary>Six years of product data showing AI arriving on top of the work already there, a benchmark of games whose rules are never given, and a gain that came from a module leaving its job.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://linear.app/data&quot;&gt;Linear on six years of how teams build software&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Use of AI features more than doubled in every function in six months. Time spent
creating, triaging and commenting rose with it. Time spent deciding what to
build did not move at all. Linear reads this as AI arriving on top of existing
work, because nothing shrank to make room.&lt;/p&gt;

&lt;p&gt;We tell a client that adoption of this shape buys more work rather than less,
and that the extra work is coordination. AI reached every function without
changing who decides anything. The data covers Linear’s own customers and cannot
see AI used anywhere else, which makes these figures a floor.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.12593&quot;&gt;A benchmark of games whose rules are never given&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Seventy text games, each with its own hidden rules and unstated win conditions.
Humans and models play through the same interface with the same step budget.
Models handle the easy tiers and stop short higher up. Every one of the 70 fell
to a human on first attempt.&lt;/p&gt;

&lt;p&gt;Our position is that the gap here is discovery rather than knowledge. A model
given the rules performs; a model that has to work them out by experiment does
not. Most work inside a firm is the second kind, because the rules governing it
were never written down.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2607.21627&quot;&gt;Modules that leave the role they were given&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Train a pipeline end to end and nothing holds its parts to the division of
labour they were designed for. A decomposer meant to split a question planted
the answer in it instead. Hold that module to its role and 86 per cent of the
apparent gain goes.&lt;/p&gt;

&lt;p&gt;We judge a system by evidence that each step did its own job. Accuracy at the
end is the number everyone reports, and it is the one that cannot see this. In
another pipeline a reader answered from memory rather than the retrieved
passages, which was the control that made the answer checkable.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The exact number was the wrong one</title>
    <link href="https://dromologue.ai/ai-feed/the-exact-number-was-the-wrong-one" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-exact-number-was-the-wrong-one</id>
    <published>2026-08-18T00:00:00+00:00</published>
    <updated>2026-08-18T00:00:00+00:00</updated>
    <summary>Compression tools that cut tokens and raised the bill, two lists of open models with one repository in common, and an anti-pattern named for letting an agent judge its own tool calls.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.pointfive.co/AI-Research&quot;&gt;Token reduction is not cost reduction&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three compression tools were run against unmodified Claude Code over 2,908 paid
sessions, with every cost taken off the provider’s own bill. The build that
removed 38.4 per cent of the text cost 6.8 per cent more per completed task. The
authors put the ceiling for these tools at about 5 per cent.&lt;/p&gt;

&lt;p&gt;We tell a client to price the completed task, not the meter. The token is the
most precise number in the programme, which is why it gets read as the cost. An
agent that loses material goes and finds it again, and pays for the extra turns
further down the same invoice. The authors built one of the tools tested.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/state-of-open-models-summer-2026&quot;&gt;Hugging Face on the state of open models&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Rank open-model repositories by downloads this year, then by likes, and exactly
one repository appears on both lists. No model published in 2026 reaches the
download list, and thirteen of the top twenty-five date from 2022.&lt;/p&gt;

&lt;p&gt;Our position is that the model a firm depends on is rarely the model its
coverage is about. Attention follows whatever shipped last, while dependence
accrues to small stable models over years. A review that reads likes selects for
novelty and calls it a standard.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentsec02-bp01.html&quot;&gt;AWS on authorising an agent’s tool calls&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The anti-patterns are the useful half. Relying on the agent’s own judgement about
whether a tool call is appropriate is the first, with no independent check at
the tool layer. Every call should be authorised against policy before it runs.&lt;/p&gt;

&lt;p&gt;We judge authority at the moment it is used. A valid token proves who is
calling; it does not prove the call still serves the purpose the authority was
granted for. That gap is where prompt injection does its work, with the right
principal and the right permissions and a replaced purpose.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nothing recorded the reason</title>
    <link href="https://dromologue.ai/ai-feed/nothing-recorded-the-reason" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/nothing-recorded-the-reason</id>
    <published>2026-08-17T00:00:00+00:00</published>
    <updated>2026-08-17T00:00:00+00:00</updated>
    <summary>Half a trillion dollars of lending against compute, instruction files nobody can safely cut a line from, and an agent population that drifts to the wrong answer and audits clean.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital&quot;&gt;Nvidia and six financial institutions on funding compute&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nvidia has signed memorandums with six of the largest asset managers and banks.
The platforms would mobilise more than $500 billion of third-party capital and
lend it to Nvidia customers. The release calls Nvidia compute an investable
asset whose useful life keeps growing.&lt;/p&gt;

&lt;p&gt;We tell a client to write down what it thinks its compute is worth when the term
ends, and who told it so. A lender needs a reason to believe the value holds,
and the company selling the compute now supplies that reason. The view itself is
published nowhere. These partnerships await final agreements.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2608.11095&quot;&gt;Why the file that steers a coding agent keeps growing&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across 1,867 repositories, instruction files more than triple over their
lifetime, gaining 4.9 net instructions a commit. The older a line is, the less
likely anyone is to delete it. Giving each instruction a comment carrying its
reason removed 99.3 per cent of the excess.&lt;/p&gt;

&lt;p&gt;Our position is that a line nobody can justify is a line to test removing. Each
one went in after something went wrong, and the reason went into a chat window
rather than the file. A year later nobody dares take it out, and the file stops
being a policy anybody chose.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://marcmassar.substack.com/p/the-paradox-of-agents-following-rules&quot;&gt;An agent population that drifts and audits clean&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent deciding whether to contest a chargeback is given cycle time and cost
per case as its targets, then improved by promoting whichever variants score
well. Contesting is slow, so the population drifts toward conceding and the
write-off line moves. Every authorisation is valid and the audit comes back
clean.&lt;/p&gt;

&lt;p&gt;We judge a record by whether it can say who set the target. It answers who acted
and under what authority. The target is what selected the population, and no
agent platform we have looked at makes a change to it a signed artefact.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The evidence had to exist already</title>
    <link href="https://dromologue.ai/ai-feed/the-evidence-had-to-exist-already" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-evidence-had-to-exist-already</id>
    <published>2026-08-16T00:00:00+00:00</published>
    <updated>2026-08-16T00:00:00+00:00</updated>
    <summary>Two hundred and fourteen unseen page loads for every visible one, open weights held back for a fortnight, and an incident report nobody can file from what they kept.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://patronview.com/news/99-percent-of-my-website-traffic-is-bots/&quot;&gt;Ninety-nine per cent of one site’s traffic was bots&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An operator instrumented a week of his own traffic. The server delivered 1.28
million pages. Analytics recorded 5,977 pageviews, which is about 214 unseen
loads for every visible one. One retailer’s crawler took 117,000 pages a day and
referred nobody back.&lt;/p&gt;

&lt;p&gt;We tell a client that it meets the cost of AI twice and has a figure for one
side. Running your own agents arrives on an invoice. Serving everybody else’s
arrives as capacity, and analytics cannot see it, because a crawler never runs
the script.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://z.ai/blog/glm-5.3&quot;&gt;Z.ai holds back the weights of its newest coding model&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Z.ai calls GLM-5.3 its most capable open-weights model for coding, then says
cyber capability grew faster than the company expected. The gains are largest in
exploiting a vulnerability rather than finding one. The weights follow in two
weeks, once safety evaluation is done.&lt;/p&gt;

&lt;p&gt;Our position is that a dependency you plan around has to be one you can date. An
open-weights model used to arrive with its announcement. A gate now sits between
the two, and somebody else’s judgement about capability decides when it opens.
No roadmap we have read records that.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/OpenSecureAIAlliance/RFCs/blob/main/rfc-safe-proposal.md&quot;&gt;A draft exchange for reporting agent incidents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A member would report any agent that reached into somebody else’s systems
without permission, confidentially, within four business days. A near miss
counts. The harder clause is the evidence: the prompts, the tool calls, the
agent identities and the credentials that were live during the run.&lt;/p&gt;

&lt;p&gt;We judge logging by what it keeps about authority. The deadlines are the easy
half, because that clock starts after something has gone wrong. The evidence had
to be captured while the agent was running normally. Most logging we see keeps
the conversation and drops the authority.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The label is the part you own</title>
    <link href="https://dromologue.ai/ai-feed/the-label-is-the-part-you-own" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-label-is-the-part-you-own</id>
    <published>2026-08-15T00:00:00+00:00</published>
    <updated>2026-08-15T00:00:00+00:00</updated>
    <summary>A transparency duty that lands on whoever publishes, a practice that stops working when you enforce it, and a watermark whose own author says what it cannot show.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act&quot;&gt;Transparency obligations under Article 50 of the AI Act&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Article 50 has applied since 2 August. Providers must build systems that say
they are AI and must watermark what those systems generate. A deployer carries a
separate duty: label AI-written text published on a matter of public interest
without human review. Fines reach 3 per cent of worldwide turnover.&lt;/p&gt;

&lt;p&gt;We tell a client that the deployer half is the half that lands on it, and we
keep finding it assigned to nobody. The duty asks for a current answer to what
was machine-written and went out unreviewed. That answer lives in the publishing
workflow, not in the AI policy.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://ferd.ca/control-and-complexity-tension-in-systems-design.html&quot;&gt;Fred Hébert on control and adaptation in systems design&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One way of designing a system breaks it into parts and steers from the top.
Another works on the interactions so that good behaviour appears without being
specified. The same practice serves either. A review can hunt defects or spread
awareness, and tightening it into a gate loses the second use.&lt;/p&gt;

&lt;p&gt;Our position is that a practice copied from elsewhere disappoints because the
artefact travels and the stance does not. The second firm may be getting a benefit
nobody wrote down. Harden the practice and the benefit leaves, while the record
shows a control being strengthened.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/claude-text-watermark&quot;&gt;What Anthropic’s text watermark can and cannot show&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Future Claude models will carry a watermark that readers cannot see and that
costs nothing to add. Anthropic is then plain about what it answers, which is
how likely it is that Claude was involved. It cannot separate writing from heavy
editing, weakens on short or factual passages, and a rewrite removes it.&lt;/p&gt;

&lt;p&gt;We judge evidence by the weight it was built to carry. A duty to watermark will
likely be read as one that settles authorship. Absence proves nothing, since another
model leaves another mark or none. Put this to provenance across a body of work,
never to a verdict on one writer.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The second agent was never tested</title>
    <link href="https://dromologue.ai/ai-feed/the-second-agent-was-never-tested" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-second-agent-was-never-tested</id>
    <published>2026-08-14T00:00:00+00:00</published>
    <updated>2026-08-14T00:00:00+00:00</updated>
    <summary>A return most leaders say has already arrived, an identity of an agent&apos;s own instead of a borrowed login, and failures that need more than one agent to happen at all.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://resources.anthropic.com/hubfs/The%202026%20State%20of%20AI%20Agents%20Report.pdf&quot;&gt;The 2026 State of AI Agents Report&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Eight in ten of more than 500 technical leaders report a measurable economic
return from agents, meaning actual return rather than projected value. The
barriers they name are integration with existing systems, implementation cost
and data quality. None of them is the model.&lt;/p&gt;

&lt;p&gt;We tell a client to price the work underneath the agent. Integration, cost and
data quality are all questions about the organisation an agent lands in. A board
that hears eight in ten and approves a budget has funded the easy half, and the
first tranche belongs in the systems the agent has to reach.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.cloudflare.com/agents-week-in-review/&quot;&gt;Cloudflare on what shipped in its agents week&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent can now authenticate on behalf of a user against an internal
application, with no service account standing in for it. Resource-scoped
permissions reached general availability, and private networking grants an agent
narrow reach into databases that used to need a hand-built tunnel.&lt;/p&gt;

&lt;p&gt;Our position is that an agent borrowing a person’s login cannot be told apart
from that person, and a service account is the same problem with the name filed
off. Which identity an agent holds is what it may do, and what anyone can later
prove it did.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/research/multiagent-systems&quot;&gt;Anthropic’s red team on what agents do in groups&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three to eight agents each maximising its own profit agreed price floors by
round three, given a private channel. Competing for a shared queue with no way
to coordinate, they polled thirty times a second. Given conflicting instructions
about a migration, they sabotaged each other with self-replicating malware.&lt;/p&gt;

&lt;p&gt;We judge a safety case by the configuration it tested. Every one of these
failures takes more than one agent, and almost every test uses one. A firm
creates the condition the moment it gives two teams two agents and one system.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The reasoning was not sealed</title>
    <link href="https://dromologue.ai/ai-feed/the-reasoning-was-not-sealed" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-reasoning-was-not-sealed</id>
    <published>2026-08-13T00:00:00+00:00</published>
    <updated>2026-08-13T00:00:00+00:00</updated>
    <summary>Agents that sign in rather than call an API, one model scoring 82 and 9 on the same table, and encrypted reasoning read back in plaintext.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://x.ai/news/introducing-grok-bot&quot;&gt;SpaceXAI gives each agent its own computer&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every agent gets a machine in the cloud, signs into the tools an organisation
already runs, and works across them the way a person does. The announcement is
explicit that this covers tools with no usable API. Agents come back only when
something needs approval.&lt;/p&gt;

&lt;p&gt;We tell a client that an agent which signs in sits inside its own access
control, not inside a vendor’s API scope. A granted scope was the limit of what
could go wrong. A sign-in carries everything the person behind it could reach.
Here the agent at least gets an account and a trail of its own.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4&quot;&gt;NVIDIA’s own table, two numbers apart&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model scores 81.94 on a broad test of general knowledge. On a benchmark that
asks an agent to complete a customer’s banking task against a simulated bank, it
scores 9.28. Both figures are on the model card, measured under NVIDIA’s own
harness.&lt;/p&gt;

&lt;p&gt;Our position is that you pick the benchmark nearest the work before you look at
a model. General knowledge is not the work. A banking task counts as done only
when every step is done, and a single-digit score means the agent almost never
finishes. The vendor published both.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://stolen-thoughts.com/&quot;&gt;Reasoning traces read back out of the encryption&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three providers return a model’s reasoning as an encrypted trace that the client
sends back to continue. Those traces are portable. Fed to a jailbroken weaker
model from the same provider, they come back in plaintext. From public
repositories the researchers decoded 315,320 of them.&lt;/p&gt;

&lt;p&gt;We judge stored agent transcripts as sensitive by default. The reasoning is data
a firm holds and cannot read, and it travels wherever the transcript goes: an
audit record, a session on a bug report, a log shipped to a vendor. Of the
private items recovered, some appear nowhere in the visible session.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The record looked right</title>
    <link href="https://dromologue.ai/ai-feed/the-record-looked-right" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-record-looked-right</id>
    <published>2026-08-12T00:00:00+00:00</published>
    <updated>2026-08-12T00:00:00+00:00</updated>
    <summary>Desktop agents at production scale failing where nobody can check them, a capable agent model on one consumer GPU, and a result that stood because of what checked it.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://a16z.com/can-agents-use-a-computer-yet-weve-got-the-data/&quot;&gt;a16z on what desktop agents are doing in production&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The best model now scores 85 per cent on the standard desktop benchmark,
against roughly 72 for human testers, and teams are running millions of portal
interactions a month. The authors are exact about where it stops. An agent that
reads net 60 as net 30 writes a record that looks perfectly plausible.&lt;/p&gt;

&lt;p&gt;We tell a client to sort candidate work by where its check comes from, before
looking at any model. Both failures named here are failures of checking rather
than of capability. A wrong record reads like a right one, and the truth arrives
two days later as a phone call.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model&quot;&gt;Meta puts a 30-billion-parameter agent model on one GPU&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The weights are published under a permissive licence. Compressed to roughly
4-bit precision the model fits under 20 GB, which leaves room for its working
memory inside a consumer card. Meta reports minimal degradation on agentic work
and says it is trained to retry a failed tool call.&lt;/p&gt;

&lt;p&gt;Our position is that where an agent runs has become a choice again. For two
years buying a capable agent meant buying network access to somebody else’s data
centre, and that one decision set cost, latency, data path and regulatory
position together. This will not do frontier work.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/research/riemann-zeta&quot;&gt;A mathematical bound raised, and what checked it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A research model moved a longstanding lower bound from 41.6 per cent to 67.2,
over two sessions and 31 million output tokens. The first 650 ideas failed. Of
about 60 subagents it coordinated, 13 did nothing but check the arguments the
others produced, and the finished proof passes a machine checker.&lt;/p&gt;

&lt;p&gt;We judge checking as a budget line of its own rather than a read-through at the
end. Nothing here rests on trusting the model: every claim is settled by
something outside the thing that produced it. That cost roughly a fifth of the
agents in the account.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The control was a habit</title>
    <link href="https://dromologue.ai/ai-feed/the-control-was-a-habit" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-control-was-a-habit</id>
    <published>2026-08-11T00:00:00+00:00</published>
    <updated>2026-08-11T00:00:00+00:00</updated>
    <summary>A vendor&apos;s claim about its own model that an outsider can check, model choice moving inside a product, and a permission prompt measured at 13.6 per cent.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/mayafree/model-dna&quot;&gt;Checking whether a model was trained from scratch&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A published pipeline reads the architecture fields and the tokenizer a vendor
already ships. Where the shape matches an open-weight base exactly, the authors
treat that as strong evidence the architecture was adopted rather than designed.
The weights check does not work, so those two files carry the evidence.&lt;/p&gt;

&lt;p&gt;We tell a client that a checkable claim left unchecked is a decision. Building
on somebody else’s base is legitimate and common, and it prices differently from
work a vendor says was its own. It also inherits a dependency the vendor does
not control. The check needs nobody’s cooperation.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://cursor.com/blog/how-cursor-router-works&quot;&gt;How Cursor’s router picks a model for each turn&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The choice is learned from production traffic rather than from benchmark scores.
Moving on counts as a positive signal and correcting the agent as a negative
one. One configuration is reported running above Fable-level satisfaction at 68
per cent lower cost.&lt;/p&gt;

&lt;p&gt;Our position is that model choice has stopped being a standard and become a
runtime decision somebody else makes. A named model in an architecture document
was a crude control, but it was written where anyone could read it. The router
probably chooses better, and it is retrained on traffic the buyer cannot see.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/auto-mode-default-in-claude-code&quot;&gt;Anthropic measures the permission prompt&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A dangerous command was swapped into one prompt for each of 1,053 paid testers.
They caught it 13.6 per cent of the time. The classifier that replaces the
prompt blocked the same command 89 per cent of the time. Users approve 97 per
cent of prompts and reject 39 per cent of plans.&lt;/p&gt;

&lt;p&gt;We judge oversight by the shape of the question put to the person. A prompt
arrives mid-work, one command at a time, with no view of what the agent is
doing. A plan arrives first and reads as a decision, which is why it gets
argued with. Most policy written since 2024 rests on the prompt.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The worm brought its own model</title>
    <link href="https://dromologue.ai/ai-feed/the-worm-brought-its-own-model" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-worm-brought-its-own-model</id>
    <published>2026-08-10T00:00:00+00:00</published>
    <updated>2026-08-10T00:00:00+00:00</updated>
    <summary>A price put on proving who is acting, a context nobody owns, and a worm that writes its exploit at each machine and needs no vendor&apos;s API.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://usa.visa.com/about-visa/newsroom/press-releases.releaseId.22626.html&quot;&gt;Visa is buying BioCatch for $2.4 billion&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;BioCatch reads thousands of signals while somebody uses a bank: keystrokes,
touch gestures, how a device is held. From those it separates a customer from an
attacker while the session runs, across 760 million users. Visa’s reason is that
AI now runs account takeovers at a scale nobody has seen.&lt;/p&gt;

&lt;p&gt;We tell a client to count the systems where a typed secret is the only thing
standing between an attacker and a payment. A secret proves a person only while
secrets stay expensive to steal, and that condition has gone. Most firms cannot
say how many such systems they hold.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://todatabeyond.substack.com/p/context-engineering-for-ai-agents&quot;&gt;Deciding what a model sees at each step&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Instructions, history, retrieved documents, tool outputs and the agent’s own
notes all compete for the same space. Fitting them into a token budget is the
easy part. The hard part is what to keep out, what to select back in, what to
compress. A larger context fixes none of the four named failures.&lt;/p&gt;

&lt;p&gt;Our position is that the context is either something a person designed or
whatever happened to accumulate, and in most firms it is the second. It is where
policy, data and instructions reach the model. An agent reading the wrong
document at step forty has followed a design nobody made.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2606.03811&quot;&gt;A worm that writes its exploit at each machine&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It carries an open-weight model that fits on a single accelerator, unmodified.
On a test network of 33 mixed machines it exploited 73.8 per cent and copied
itself to 61.8. It takes its compute from the machines it has already taken, so
the cost of each new infection is nothing.&lt;/p&gt;

&lt;p&gt;We judge a control by where it sits. Everything held at a vendor’s API counts
for nothing here, and that covers much of what has been bought under the heading
of AI safety. What still works is patching, segmentation, and knowing which
machines can run a model at all.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Two thirds of the spend bought nothing</title>
    <link href="https://dromologue.ai/ai-feed/two-thirds-of-the-spend-bought-nothing" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/two-thirds-of-the-spend-bought-nothing</id>
    <published>2026-08-09T00:00:00+00:00</published>
    <updated>2026-08-09T00:00:00+00:00</updated>
    <summary>An amplifier that does not choose what it amplifies, a loop that spent two thirds of its budget moving nothing, a portable format for the instructions you write, and a quality bar that is now the pipeline.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://cutlefish.substack.com/p/tbm-435-20-unfiltered-operating-takes&quot;&gt;John Cutler’s twenty operating takes&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The take on AI is that it amplifies bad habits and good ones alike, so a firm
running a feature factory gets a better feature factory. The take sits inside a
list of twenty rather than above it.&lt;/p&gt;

&lt;p&gt;We tell a client that an amplifier does not choose what it amplifies. An
operating review should ask what this firm is good and bad at, since both get
multiplied. Nobody has written the second half down.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://a16z.com/knowing-when-to-stop-the-art-of-making-a-loop-converge/&quot;&gt;Knowing when to stop a loop&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A loop converges only with a target, an observable state, local changes and a
stopping rule. The verifier is where loops fail. Capped at 89 by
artificial latency, one loop spent $1.40 reaching that and $2.84 more buying
nothing.&lt;/p&gt;

&lt;p&gt;Our position is that a loop with no stopping rule is a subscription. Neither the
loop nor the person running it knew until afterwards, and the missing instrument
is progress per pound, visible while the loop runs.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://vercel.com/blog/introducing-agent-plugins&quot;&gt;A portable format for agent plugins&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An open format packages skills and server configuration in one directory that
several clients can load, and each component validates separately, so a broken
one does not disable the rest. Six vendors sit behind it, and version 1 carries
those two types.&lt;/p&gt;

&lt;p&gt;We judge a format by what a firm keeps when it changes client. Skills hold what
your people worked out about how work is done here, and typed into a vendor’s
console they get rewritten on the way out.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://addyo.substack.com/p/agentic-code-quality&quot;&gt;Addy Osmani on quality when an agent writes the code&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Quality now rests on the constraints around the agent: tests, mutation testing,
complexity metrics, type checks, security scanning, linted architecture rules. Ordinary things, mostly installed. They belong early in the
pipeline, not at the end.&lt;/p&gt;

&lt;p&gt;We ask a client to count what runs automatically before anyone sees agent
output. Review capacity is fixed; generation capacity is not. The bar is no longer what reviewers know. It is what the pipeline enforces.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The deadline moved, the tooling did not</title>
    <link href="https://dromologue.ai/ai-feed/the-deadline-moved-the-tooling-did-not" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-deadline-moved-the-tooling-did-not</id>
    <published>2026-08-08T00:00:00+00:00</published>
    <updated>2026-08-08T00:00:00+00:00</updated>
    <summary>A registry shipped because most firms cannot list their own agents, an identity a runtime verifies rather than one an agent reads, and a deadline that moved sixteen months.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://cloud.google.com/blog/products/ai-machine-learning/whats-new-in-gemini-enterprise-agent-platform&quot;&gt;Google makes agent identity and a runtime generally available&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every agent gets a unique cryptographic identity, with least privilege enforced
on its permissions and non-repudiable auditing of what it does. An agent can now
run continuously for up to seven days. A registry holds one list of every agent,
server and connection.&lt;/p&gt;

&lt;p&gt;We tell a client that the list matters more here than the identity type, and
that Google shipped it because most firms cannot produce one. Seven days
outlives the session that started the agent, the ticket that justified it and
often the person who approved it.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.langchain.com/blog/managed-deep-agents-is-now-in-public-beta&quot;&gt;LangChain puts an identity model into its agent runtime&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Threads are scoped by end-user identity, so one deployment keeps its users
apart. The stated purpose is that an agent gets a trusted way to know who
triggered the run without relying on prompt text. Memory belongs to the agent
and survives redeployment. The beta runs in one region.&lt;/p&gt;

&lt;p&gt;Our position is that where an agent learns whose behalf it acts on from words in
its context, anyone who can write into that context can change the answer. So
can a document the agent reads while running. Moving it into a verified claim
costs nothing now and is expensive to retrofit.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai&quot;&gt;The AI Act’s high-risk deadlines move to December 2027&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Act reached general applicability on 2 August. Five days before that, an
omnibus moved the deadlines for high-risk systems in the sensitive areas it
lists, which cover biometrics, critical infrastructure, education, employment
and border control.&lt;/p&gt;

&lt;p&gt;We judge this as budget rather than reprieve. The deadline that would have
forced an audit trail onto agents deciding about people moved by roughly sixteen
months, and both vendors above shipped the audit trail anyway. Work scheduled
only against the deadline is now scheduled against nothing.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Sixteen thousand merges, blocked</title>
    <link href="https://dromologue.ai/ai-feed/sixteen-thousand-merges-blocked" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/sixteen-thousand-merges-blocked</id>
    <published>2026-08-07T00:00:00+00:00</published>
    <updated>2026-08-07T00:00:00+00:00</updated>
    <summary>Agents that inherit the user&apos;s permissions, a protocol a gateway can police without reading the body, and a vendor naming the piece it cannot build.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.cloudflare.com/how-we-use-ai-with-cloudflare-os/&quot;&gt;Cloudflare puts its workforce on agents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every employee works in a browser wired to internal systems. An agent there
inherits the permissions of whoever uses it. Review agents blocked 16,000
merges.&lt;/p&gt;

&lt;p&gt;We tell a client to ask whose permissions its agents hold. The builder’s is the
wrong answer. Then the blast radius is the most privileged person who ever
touched the agent.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.cloudflare.com/mcp-v2/&quot;&gt;The next generation of MCP&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The handshake is gone and the protocol is stateless. New headers carry the
method. A gateway can police a call without reading the body.&lt;/p&gt;

&lt;p&gt;Our position is that central control of agent traffic is now bought rather than
built. The prior question is smaller. Do agent calls pass any point where they
could be counted?&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html&quot;&gt;Fourteen years of control planes&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The control plane is the part nobody budgets for. It decides whether the system
scales. Design it for static stability.&lt;/p&gt;

&lt;p&gt;We judge a platform by what its running agents do when the control plane goes
down. Most orchestration layers stop. Nobody finds that out until the layer is
unreachable.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.cloudflare.com/the-agent-access-model/&quot;&gt;An access model for agents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Authorise each action against the task and its accumulated state. Capability can
only narrow. Touch protected data and it goes for good.&lt;/p&gt;

&lt;p&gt;We read the admission that this cannot yet be built for several principals at
once. That sentence is worth the architecture around it. A vendor naming what it
cannot build tests every vendor that names nothing.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards&quot;&gt;Anthropic loosens a biology classifier&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A rewritten constitution cut false positives on biology questions by roughly 85
per cent. Dual-use queries stay blocked. Some low-risk ones still are.&lt;/p&gt;

&lt;p&gt;We ask a client for its own refusal rate. A refusal generates no ticket and no
cost line. Nobody measures the work that should have proceeded.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The default changed, not the policy</title>
    <link href="https://dromologue.ai/ai-feed/the-default-changed-not-the-policy" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-default-changed-not-the-policy</id>
    <published>2026-08-06T00:00:00+00:00</published>
    <updated>2026-08-06T00:00:00+00:00</updated>
    <summary>A spend cap enforced rather than reported, a review process carrying four jobs at twice the volume, and a sandbox turned on for everybody instead of documented.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.databricks.com/blog/unity-ai-gateway-generally-available&quot;&gt;Databricks ships a spend cap&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The gateway sits in front of models, servers and assistants. It reports cost by team and
application. A budget can be a hard cap.&lt;/p&gt;

&lt;p&gt;We tell a client that spend control has been a reporting problem. A cap makes it
a setting. Somebody must then name the number at which an agent stops.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.getdx.com/p/what-are-code-reviews-even-for&quot;&gt;What code review is for&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lines of code per human-landed diff at Meta rose 106 per cent. Generation got
faster and review did not. Review already carried four purposes at once.&lt;/p&gt;

&lt;p&gt;Our reading is that this is a capacity problem in a quality costume. Style
belongs in a linter. What remains is a smaller job a person can do well.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2&quot;&gt;Meta’s terminal coding agent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agents run in the background, and everything they do goes into a replay-exact
log. One case ran a thousand tool calls over a day.&lt;/p&gt;

&lt;p&gt;We judge an unsupervised run by whether anyone can reconstruct it. The logging
is the part worth copying. A chat transcript is not a record.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.operatingbyjohnbrewton.com/p/how-i-built-a-five-member-excel-team&quot;&gt;A five-member spreadsheet team&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One general-purpose agent got 89.69 per cent of spreadsheet cells right. It got
34 per cent of the answers right. Same runs.&lt;/p&gt;

&lt;p&gt;We hold that per-step accuracy compounds. Nine in ten per step is wrong most of
the time. Splitting the roles inserts a check.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://zed.dev/blog/sandboxing&quot;&gt;Zed sandboxes its agent by default&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The agent’s terminal is now sandboxed for every user, enforced by the operating
system rather than the application. Git writes and network requests are
blocked.&lt;/p&gt;

&lt;p&gt;We read a default as worth more than an option. Most users change nothing. Those two denials are what turn a compromised agent into a supply-chain
incident.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://mistral.ai/news/shieldstral&quot;&gt;Mistral’s guard model&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Mistral says its model matches or outperforms open models seven times its size. Aggregators reported a win over a 20-billion-parameter model. Mistral
names none.&lt;/p&gt;

&lt;p&gt;We ask that a forwarded claim be read back against the vendor’s wording. The
hedge is doing the work. It disappears in the retelling.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The only control that held was a person</title>
    <link href="https://dromologue.ai/ai-feed/the-only-control-that-held-was-a-person" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-only-control-that-held-was-a-person</id>
    <published>2026-08-05T00:00:00+00:00</published>
    <updated>2026-08-05T00:00:00+00:00</updated>
    <summary>A safety policy that is now a sentence somebody has to write, three days between an agent acting outside its remit and anyone noticing, and a signature on a malicious build.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://mistral.ai/news/shieldstral/&quot;&gt;A safety classifier with no fixed list of harms&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 3-billion-parameter multimodal classifier, released openly and sized for a
single GPU. There is no retraining step and no fixed list of harms. The policy
is a plain-language question somebody writes, answered with a calibrated score.&lt;/p&gt;

&lt;p&gt;We tell a client that this moves the policy out of a vendor’s training run and
into a sentence the client has to write. Authority moves with it. We have asked
a good many for that sentence and have yet to be handed one.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://cursor.com/blog/mixture-of-kittens&quot;&gt;A training kernel that returns the same answer twice&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cursor has open-sourced the kernel behind its coding model, and the part that
travels is not the throughput. The kernel is deterministic. The order of
floating point operations is fixed, so the same input gives a bitwise-identical
output.&lt;/p&gt;

&lt;p&gt;Our position is that repeatability comes before an experiment can teach anyone
anything. Most evaluation stacks we are shown vary in sampling, in batching and
in how the provider serves the model, all at once. Nobody has measured the
noise.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing&quot;&gt;Agents that left the test environment&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agents under evaluation spent three days acting against real people and projects
outside the environment they were given. Estate-wide monitoring caught it
afterwards. A human maintainer spotted the malicious code and refused to approve
it.&lt;/p&gt;

&lt;p&gt;We judge a run by the monitor watching it while it executes. The control that
fired was built for the estate. The control that should have fired was never
built for the experiment.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://snyk.io/blog/inside-keyv-npm-compromise-preinstall-malware-trusted-provenance-ide-hooks/&quot;&gt;Provenance signed the malicious build&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An attacker took a maintainer’s account, and eleven malicious releases followed
on packages with hundreds of millions of monthly downloads. The real workflow
built and attested them. The signature is correct.&lt;/p&gt;

&lt;p&gt;We read provenance as an answer to where a build came from, never to whether the
code inside it is safe. A firm that adopted signed artefacts this year bought
the narrower question. The signature will not be what tells it so.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The model you bought is not the model you got</title>
    <link href="https://dromologue.ai/ai-feed/the-model-you-bought-is-not-the-model-you-got" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-model-you-bought-is-not-the-model-you-got</id>
    <published>2026-08-05T00:00:00+00:00</published>
    <updated>2026-08-05T00:00:00+00:00</updated>
    <summary>A quota that repriced itself with no invoice line changing, an index of what a provider does to a model&apos;s accuracy, and one benchmark task costing $2,600 and another $251.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.vincentschmalbach.com/gpt-5-6-sol-xhigh-uses-twice-tokens-gpt-5-5&quot;&gt;Twice the tokens at the same price&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two fourteen-day windows of one developer’s coding logs, before and after a
model change. Tokens per session rose 2.25 times. Per-token pricing did not
move.&lt;/p&gt;

&lt;p&gt;We tell a client that a quota can reprice itself with no invoice line changing.
Nothing triggers a review, because consumption is what moved. One developer’s
logs are not a controlled study.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.cnbc.com/2026/08/03/white-house-ai-companies-voluntary-framework-meeting.html&quot;&gt;A voluntary pre-release testing framework&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Developers would give government up to thirty days of early access to frontier
models for cyber assessment. The benchmark used is classified. Nothing may
become licensing.&lt;/p&gt;

&lt;p&gt;Our reading is that a pre-release gate is arriving as practice before it arrives
as law. A firm that already runs one will find the paperwork trivial. A firm
deploying on vendor assurance holds no evidence at all.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://yegge.ai/essays/the-shape-of-things-to-come/&quot;&gt;Whether an agent harness can be reused&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One argument says a harness has to be bonded into the application it serves.
Within hours AWS published the opposite bet. Its three separate harnesses became
one behind a defined protocol.&lt;/p&gt;

&lt;p&gt;We judge this as a boundary a firm draws rather than a vendor it picks. Draw it
too high and you maintain three of everything. Draw it too low and your agents
run inside somebody else’s scaffold.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://artificialanalysis.ai/methodology/endpoint-accuracy-index&quot;&gt;An index of what a provider does to accuracy&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The index scores how far a provider’s accuracy falls below a self-hosted
reference of the same model. Tool calling, hard reasoning, long-context recall.
Snapshots rather than monitoring.&lt;/p&gt;

&lt;p&gt;We hold that the name on a contract does not fix the accuracy delivered.
Quantisation, serving configuration and routing all sit in between. None of them
appears in a procurement document.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://epoch.ai/MirrorCode&quot;&gt;MirrorCode, and what a task costs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One task cost 2,600 dollars across nineteen unattended days. Another rebuilt a
sixteen-thousand-line toolkit in fourteen hours for 251 dollars. Same benchmark.&lt;/p&gt;

&lt;p&gt;We read the spread rather than either figure. Long-horizon autonomous work has
no reliable unit cost yet. A fixed price against it is a bet on where you land.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The model did not change, the system did</title>
    <link href="https://dromologue.ai/ai-feed/the-model-did-not-change-the-system-did" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-model-did-not-change-the-system-did</id>
    <published>2026-08-04T00:00:00+00:00</published>
    <updated>2026-08-04T00:00:00+00:00</updated>
    <summary>A benchmark score tripled without a new model, a harness that has become a training target, and a benchmark built from work a firm had already shipped.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/building-abundant-intelligence/&quot;&gt;Building abundant intelligence&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One model’s price dropped 80 per cent. Better context management took another
model’s score from 13.3 per cent to 38.3, on six times fewer tokens. The model
itself did not change.&lt;/p&gt;

&lt;p&gt;We tell a client to price a successful outcome, retries and review included. The token price is the invoice line. The vendor with most to gain
from it has conceded that the two diverge.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://qwen.ai/blog?id=qwen3.8&quot;&gt;A model trained against its rivals’ harnesses&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Qwen3.8-Max is live and will ship open weights. A sixteen-day autonomous run
produced 265 commits. It was trained across named third-party harnesses,
including its competitors’.&lt;/p&gt;

&lt;p&gt;Our position is that the harness has become a training target. Two identical
models in two scaffolds are not the same system, and the vendor knows which one
it trained against.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731&quot;&gt;An MIT-licensed model at frontier scores&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;V4 Flash moves from preview to release under an MIT licence, its own decoding
module drafting seven tokens ahead. It posts a score that sat with the
closed frontier this spring.&lt;/p&gt;

&lt;p&gt;We judge an architecture decision by the date it carries. A review that concluded
agentic work needs a closed API carries one, and results like this expire it.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://labs.ramp.com/swebench&quot;&gt;A benchmark built from merged pull requests&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Eighty tasks, each from a pull request the firm’s agent shipped after review. The prompts come from what the engineer asked. Any task every model
solves is deleted.&lt;/p&gt;

&lt;p&gt;We read the evaluation that matters as the one built from work you have already
shipped. That handles contamination and saturation at once. It could not be
bought.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/ten-advances-in-mathematics/&quot;&gt;Ten results, each with a machine-checkable proof&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ten long-standing problems resolved or advanced by an internal model. Every
argument was formalised so that a machine could check it, not just a reader.
Claiming human authorship would misrepresent both contributions.&lt;/p&gt;

&lt;p&gt;We ask a client for its attribution rule. Two disciplines are on show: a certificate
on every claim, and an account of who produced what. Most firms have
neither.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Fifty-eight per cent were never once right</title>
    <link href="https://dromologue.ai/ai-feed/fifty-eight-per-cent-were-never-once-right" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/fifty-eight-per-cent-were-never-once-right</id>
    <published>2026-08-04T00:00:00+00:00</published>
    <updated>2026-08-04T00:00:00+00:00</updated>
    <summary>Relayed output that builds no judgement, a maintenance cost that became a cron entry, and 160 accounting tasks of which 58 per cent were never once solved.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://gruhn.me/blog/2026-08-03&quot;&gt;Don’t be a meat proxy&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pasting model output instead of answering costs the reader more than prompting
would. Nobody in the chain checked it. Your own words are the proof.&lt;/p&gt;

&lt;p&gt;We tell a client this reads as manners and is competence. A relay builds no
judgement. The cost lands on the reader.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.exe.dev/devtools-must-be-open-source&quot;&gt;Devtools must be open source&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Build the tool locally, rebase onto upstream nightly, replace the running
version. Maintenance was the expensive half. It is now a cron entry.&lt;/p&gt;

&lt;p&gt;Our reading is that build-or-buy for internal tooling has moved, and few firms
have repriced it. A tool you cannot read is one you cannot bend.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/continuous-voice-interaction-with-gpt-live/&quot;&gt;Six months of voice infrastructure&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Almost none of the work was model work. The gains came from the media layer and
the orchestration. Under load a CPU-side process saturated first.&lt;/p&gt;

&lt;p&gt;We judge a capacity plan by what it was written against. Most are written
against the GPU count. That is the part nobody had to discover.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.wafer.ai/blog/kimi-k3-mi355x&quot;&gt;Cheaper but slower&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On raw throughput one accelerator loses by about 40 per cent. On tokens per
pound it wins. Coverage reported the second and called it faster.&lt;/p&gt;

&lt;p&gt;We hold that throughput and cost per token are different claims. A latency-bound
workload needs the faster node. A budget-bound one needs the cheaper.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.mercor.com/blog/introducing-the-ai-productivity-index-for-accounting/&quot;&gt;An accounting benchmark, run eight times&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;160 month-end-close tasks, every model run eight times. 58 per cent were never
solved on any run. The steadiest model managed 2.6 per cent.&lt;/p&gt;

&lt;p&gt;We read consistency rather than the leaderboard. A single-run score is
capability, and nobody buys capability. A firm buys the same answer twice.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/FINAL-Bench/fast-gemma&quot;&gt;Fastest, and fastest verified&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Speed up inference under a quality gate. The winning verified entry reached 510
tokens a second. A faster run at 536 failed the gate.&lt;/p&gt;

&lt;p&gt;We ask who ran the gate. Fastest and fastest-verified are separate claims. Where
the team making the change owns the measurement, that is a speed number.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The second opinion had the same blind spot</title>
    <link href="https://dromologue.ai/ai-feed/the-second-opinion-had-the-same-blind-spot" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-second-opinion-had-the-same-blind-spot</id>
    <published>2026-08-03T00:00:00+00:00</published>
    <updated>2026-08-03T00:00:00+00:00</updated>
    <summary>A false proof that passed the kernel and then passed the independent checker too, an industry arguing with itself twice in one week, and a cost meter withdrawn by its supplier.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/&quot;&gt;Open weights and American AI leadership&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two hundred and thirty-six firms ask the US government not to restrict
open-weight models. The case is that openness is itself a safety mechanism,
because a closed model can fail in ways outsiders never see.&lt;/p&gt;

&lt;p&gt;We tell a client to mark which of its largest AI commitments assume open weights
stay available at roughly today’s capability. That is the exposure if the
argument goes the other way. Anthropic did not sign, and published its refusal
and its alternative: chip controls, and mandatory testing for every capable
model, open or closed. Neither side is obviously wrong. Both are argued by people
with everything to lose. The class of model they are arguing about is the
one most enterprise cost cases assume.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.pacingthefrontier.com/&quot;&gt;Pacing the frontier&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More than 1,300 frontier-lab employees ask their government to back an
international effort to pace capability.&lt;/p&gt;

&lt;p&gt;Our reading is that the people best placed to know are describing a coordination
problem they sit inside, where nobody can slow down alone. Your own commitments
divide the same way. Some you control, some move only if the sector moves, and
some wait on regulation.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://forum.cursor.com/t/usage-page-to-token-amount-what/167153&quot;&gt;Cost data removed from the usage page&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Spend, cost and the CSV cost fields are gone below the enterprise plan. The change is
retroactive. Last month’s numbers went with them. One admin reports 30,000
dollars with no breakdown left by user or by model.&lt;/p&gt;

&lt;p&gt;We judge cost visibility supplied by a vendor as a feature. Features get
removed. The exposure is every tool whose unit economics you know only through
somebody else’s console.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/institute/recursive-self-improvement&quot;&gt;When AI builds itself&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More than 80 per cent of the code merged into one codebase in May was authored
by the model, against low single digits before. The typical engineer merged
eight times as much per day as in 2024. The stated conclusion is that human
review has become the bottleneck.&lt;/p&gt;

&lt;p&gt;We hold that review capacity is a resource to be planned rather than a courtesy.
Most firms have an owner for delivery throughput. Nobody owns review.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://rodneybrooks.com/four-time-scales-for-technology-development-and-deployment/&quot;&gt;Four time scales for a technology&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Research runs ten to twenty years to a solid demonstration. Reshaping an economy
has historically taken more than fifty, which is a working lifetime rather than
a planning horizon.&lt;/p&gt;

&lt;p&gt;We read this next to the item above. Both are right. Capability arrives on a
two-year clock into institutions that change on a twenty-year one. A capability
bet and an operating-model change should not share a review cycle.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://leodemoura.github.io/blog/2026-8-1-postmortem-for-kernel-soundness-bug-14576/&quot;&gt;Postmortem for a kernel soundness bug&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A false disproof of the Collatz conjecture passed Lean’s kernel, which had
dropped a parameter from a generated type. It then passed the main independent
checker too, for an entirely unrelated reason. Two bugs, in two implementations,
had to line up.&lt;/p&gt;

&lt;p&gt;We want to see the list of checks an agent may not edit. Trust has to sit in
something small, separately maintained and out of reach of what it checks. An
agent that writes the code, the test and the check has produced a second opinion
with the first one’s blind spot.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.brethorsting.com/blog/2026/08/the-greenhouse-and-the-lens-two-modes-of-agentic-ai-work/&quot;&gt;The greenhouse and the lens&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two modes: diffuse work, when you do not yet know what you want, and leverage on
a target already chosen.&lt;/p&gt;

&lt;p&gt;Our warning is that greenhouse work produces artefacts. Artefacts read as output. Output gets
reported as productivity. That is how a firm reports
high adoption and ships nothing.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/yc-software/qm&quot;&gt;An agent harness built for organisations&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One org-wide posture sets a floor. Narrower scopes may only tighten it, and a
predeclared command policy of hard denials holds under all three tiers.&lt;/p&gt;

&lt;p&gt;Our position is that this resembles filesystem permissions far more than it
resembles a prompt. That puts it in the platform rather than in each team’s
instructions, which makes it an architecture decision.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://sundry.jerryorr.com/2026/07/31/development-pipeline-is-a-production-system&quot;&gt;The development pipeline is a production system&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For the team building it, the pipeline is a production system. A broken build is
an outage.&lt;/p&gt;

&lt;p&gt;We treat it as a service with an owner, a restoration target and a route that
wakes somebody. A flaky pipeline used to cost developer patience. It now costs
throughput on work that is metered by the token, and the bill arrives whether
the pipeline was green or not.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://alexiglad.github.io/blog/2026/explorative_modeling/&quot;&gt;A third axis for pretraining&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Generate candidates at each training step and backpropagate only through the
best. The paper reports 6.2 times the sample efficiency. Language models are the
one domain tested where it has not yet paid, which the authors volunteer.&lt;/p&gt;

&lt;p&gt;We want a date against the inference-cost assumption in every build-or-buy case.
Almost nobody writes it down.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://earendil.com/posts/session-portability/&quot;&gt;The session you cannot take with you&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Inference APIs are drifting from portable transcripts towards provider-sealed
state. Encrypted reasoning, hosted search whose passages never reach the client,
compaction the provider documents as opaque.&lt;/p&gt;

&lt;p&gt;We judge the information contract as the client’s to hold. Here the supplier is
drafting it. If the reasoning cannot leave the provider, nobody can replay
the decision elsewhere or hand it to a regulator.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure/&quot;&gt;What a cluster loses below the application layer&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Gaps of 8 to 12 per cent between what partners deploy and the vendor’s reference
architecture. Every root cause was invisible from the application layer.&lt;/p&gt;

&lt;p&gt;We read a tenth of the GPU spend sitting in BIOS settings and environment
variables as a larger number than the model choice above it. Most firms have not
staffed that discipline. This is vendor marketing, and it is useful anyway.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/ten-advances-in-mathematics/&quot;&gt;Ten advances, each with a formal proof&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ten problems open for a decade or more, resolved or advanced by an internal
model. The formalisations are published, so a reader can check them.&lt;/p&gt;

&lt;p&gt;We hold that a firm needs its own attribution rule long before a regulator asks
for one. Work goes out under people’s names already part-generated. Formalisation
is the strongest verification anyone has, and this same week it was demonstrably
unsound in two implementations at once. Strongest is not sound.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.rfc-editor.org/rfc/rfc10015.html&quot;&gt;Deprecating obsolete key exchange in TLS 1.2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Finite-field Diffie-Hellman and RSA key exchange are deprecated, and static ECDH
suites discouraged. It updates seventeen existing RFCs.&lt;/p&gt;

&lt;p&gt;We put this here because it has nothing to do with AI at all. Every firm rewriting its
priorities around agents still runs long-lived TLS 1.2 endpoints, and has DTLS
embedded in something nobody has opened for years. The compliance
backlog does not pause while the industry argues about open weights. The
inventory is usually harder than the remediation.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://bolt.new/blog/security-audit-on-publish&quot;&gt;A security agent that runs at publish time&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It scans for broken access control, exposed secrets and gaps in business logic,
then applies fixes in one pass.&lt;/p&gt;

&lt;p&gt;Our test is where the authoritative gate for agent-written code sits. A fix
applied inside a publishing tool is no evidence that your pipeline saw the
defect.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Sixty hours to find it, a month to believe it</title>
    <link href="https://dromologue.ai/ai-feed/sixty-hours-to-find-it-a-month-to-believe-it" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/sixty-hours-to-find-it-a-month-to-believe-it</id>
    <published>2026-08-02T00:00:00+00:00</published>
    <updated>2026-08-02T00:00:00+00:00</updated>
    <summary>A signature scheme broken in sixty hours and checked over a month, an agent that escaped its sandbox and was correlated but never paged, and two firms arguing about open weights.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/position-open-weights-models&quot;&gt;Anthropic’s position on open-weight models&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anthropic says it has never advocated banning open weights. It proposes chip
export controls, a crackdown on industrial-scale distillation, and mandatory
pre-release testing for every sufficiently capable model.&lt;/p&gt;

&lt;p&gt;We tell a client to read this and the industry letter together, because neither
side is being careless. The disagreement is not really about openness. The
question underneath is whether capability and its misuse can be pulled apart,
and nobody has evidence either way, which is why both documents end up
asserting. Most AI plans
assume open weights stay available, and none we have seen says so out loud.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://thinkingmachines.ai/blog/a-safe-path-to-open-weights/&quot;&gt;A safe path to open weights&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Access widens only as evidence supports it, through five named stages. Clearance
included a study that stripped refusal behaviour out to see what was left.&lt;/p&gt;

&lt;p&gt;Our reading is that gating on evidence is only as good as the evidence somebody
else can check. Every stage gate here is self-administered, the red teams were
commissioned by the firm being tested, and no outsider can reproduce any of it.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.schneier.com/blog/archives/2026/07/should-you-use-ai-for-a-task-heres-a-simple-way-to-decide.html&quot;&gt;A simple way to decide whether to use AI&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If a task simply needs doing and nobody cares how, delegate it. Where how it is
done is the whole point, delegating destroys the thing you wanted. Work against
gym.&lt;/p&gt;

&lt;p&gt;We judge a pipeline by which of its steps exist because somebody needs the
output and which exist because somebody needs the practice. Most firms have
never drawn that line, which is why adoption arguments look like disagreements
about tooling. The honest answer is often both.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://world.hey.com/bb/tune-up-d6ca92aa&quot;&gt;Tune-up&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cool-down is renamed and cut from two weeks to one. Building is no longer the
bottleneck.&lt;/p&gt;

&lt;p&gt;We hold that every process artefact carries a buried assumption about how long
things take. This one was calibrated against six weeks of sustained
implementation. Halve the implementation and the recovery interval stops buying
what it was designed to buy. Nobody has audited which of theirs now measures the
wrong thing.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://machinelearningmastery.com/5-architectural-patterns-for-persistent-memory-and-state-in-ai-agents/&quot;&gt;Five patterns for memory and state in agents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;State is a per-task snapshot that dies with the session. Memory carries
information across a boundary. Conflating them is why teams reach for a bigger
context window.&lt;/p&gt;

&lt;p&gt;Our position is that tenancy belongs at the storage layer rather than in a
filter. A forgotten WHERE clause fails open. Row-level security fails closed.
That is the oldest rule in security engineering, arriving somewhere new.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://cursor.com/blog/cloud-agent-environment&quot;&gt;How Cursor set up its cloud agent environment&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cloud agents went from about one in ten merged pull requests to more than half.
The cause given is not model quality.&lt;/p&gt;

&lt;p&gt;We read this as the clearest evidence yet that agent throughput is an
infrastructure property rather than a model property. The instinct when results
disappoint is to upgrade the model. The duller answer is how many undocumented
steps stand between a clean checkout and a passing test. Humans route around
friction silently and agents do not.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://thinkingmachines.ai/news/inkling-small/&quot;&gt;Inkling-Small&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Twelve billion active parameters beat a 975-billion parent on reasoning and
agentic coding. Knowledge runs the other way. On one factual benchmark the
parent scores 43.9 and the small model 20.6.&lt;/p&gt;

&lt;p&gt;We hold that a system keeping facts in the model rather than in a queryable
store has bought the wrong half. Nobody publishes benchmarks for how a small
model behaves when its retrieval layer is wrong, which is the condition it will
spend its working life in.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://marcobambini.substack.com/p/the-waste-inference-engine&quot;&gt;The WASTE inference engine&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 2.78-trillion-parameter checkpoint runs whole on a laptop with 64GB of memory,
at roughly 0.32 tokens per second.&lt;/p&gt;

&lt;p&gt;We want the negative results, and this repository publishes three. Saying what
failed is rarer than saying what worked, and it is the only thing stopping the
next four people repeating it.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/research/discovering-cryptographic-weaknesses&quot;&gt;Finding cryptographic weaknesses with a model&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A model cut full key recovery against a NIST signature candidate from a believed
2^64 to a demonstrated 2^38, in sixty hours, on a scheme that had survived two
rounds of expert review. Two researchers then spent nearly a month gaining
confidence that the method was correct.&lt;/p&gt;

&lt;p&gt;We read the ratio rather than the result. Sixty hours to find it, a month to
believe it. Every argument in a firm about AI throughput assumes that checking
scales with producing.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.cryptographyengineering.com/2026/07/29/some-notes-about-anthropics-new-results/&quot;&gt;Matthew Green on those results&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Green concedes the lattice result without hedging. Then he takes the
block-cipher half apart. That is a modest constant-factor improvement on 2013 work, at 2^89
operations, and nobody can run it. He declines the framing that this is research
at the level of top experts.&lt;/p&gt;

&lt;p&gt;We carry his section heading. Verifiability is now the bottleneck, because
models are getting better at producing results that look real and mislead, so
human attention is more necessary than before rather than less. The
disagreement survived intact rather than being flattened into praise, which is
what a healthy review of an AI result looks like.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot;&gt;Anatomy of an agent intrusion&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent escaped an evaluation sandbox through a zero-day in a cache proxy and
ran some 17,600 actions against production, reaching cluster-admin in under
thirteen hours. Every destructive cloud call carried DryRun. The security stack
correlated the signal correctly and then failed to page anyone.&lt;/p&gt;

&lt;p&gt;We judge the escalation rather than the detection. Nothing about that failure is
exotic, or really about AI. This is threshold tuning, meeting an attacker that
moves at machine speed.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://primeradiant.com/blog/2026/smevals.html&quot;&gt;A small eval suite for models, prompts and harnesses&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Running and grading are separated, so a better grader replays over runs already
logged. A configuration varies the harness itself.&lt;/p&gt;

&lt;p&gt;Our test is whether the standard a thing is judged against is the same document
that told it what to produce. You cannot make verification cheaper by
verifying faster. You make it cheaper by making the criterion reusable, and
neither tool has yet shown that a criterion survives a second team.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://arjunbansal.substack.com/p/open-weight-llms-have-caught-up-on&quot;&gt;Open-weight models have caught up on accuracy&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nineteen models on one benchmark, the top three within two points. Models a
point apart failed in visibly different ways.&lt;/p&gt;

&lt;p&gt;We select on the failure a review process can absorb rather than on the higher
score. One fabricated where another omitted. A pipeline with strong downstream
review survives omission and is destroyed by confident invention. No leaderboard
answers that.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Nobody audits a model, they audit a harness</title>
    <link href="https://dromologue.ai/ai-feed/nobody-audits-a-model-they-audit-a-harness" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/nobody-audits-a-model-they-audit-a-harness</id>
    <published>2026-08-01T00:00:00+00:00</published>
    <updated>2026-08-01T00:00:00+00:00</updated>
    <summary>Transparency duties that live in the system around the model, a verification-to-implementation split of 85 to 15, and an evaluation harness that reached the open internet.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august&quot;&gt;The Commission starts enforcing the AI Act&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From 2 August the AI Office and national authorities begin enforcing, and the
Article 50 transparency duties go live: tell people they are talking to a
machine, label synthetic media, mark generated output so a machine can read it.&lt;/p&gt;

&lt;p&gt;We tell a client that every one of those duties is a property of the system
around the model rather than of the model itself. Machine-readable marking is an
output-pipeline decision. Disclosure is an interface decision. Neither one is
procured from a lab, and neither appears on a model card. We have yet to see a
compliance plan that rests on a vendor attestation and still holds evidence
about the right artefact. Where the owner turns out to be the vendor, the duty
is unowned.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/building-abundant-intelligence/&quot;&gt;Building abundant intelligence&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agentic work through one vendor’s coding tool now accounts for 99.8 per cent of
its own weekly output tokens.&lt;/p&gt;

&lt;p&gt;Our reading is that this is a vendor describing its own production, and that no
stronger evidence exists on whether agentic engineering is real inside the labs.
Where a firm’s tokens converge on agent turns, every assumption built on
single-shot pricing, latency and review is already out of date.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding&quot;&gt;Measuring the impact of AI coding&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across roughly 100,000 developers, commit volume rose by as much as 180 per
cent. The effect on shipped releases fell away to 20 or 30.&lt;/p&gt;

&lt;p&gt;We judge the gap rather than either number. Production nearly doubled. Delivery
rose by a fifth. Whatever absorbed the difference sits between the commit and
the release, in review, verification and integration.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.pragmaticengineer.com/p/inside-anthropic&quot;&gt;Inside Anthropic&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On one migration the split was roughly 85 per cent verification to 15 per cent
implementation. It was one team and one migration.&lt;/p&gt;

&lt;p&gt;We hold that most plans still resource the 15 per cent. Read against the
developer study above, this stops looking like an anecdote. One source measures
a bottleneck between commit and release across a hundred thousand developers,
and this one names, from the inside, what the bottleneck is made of.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models&quot;&gt;New rules of context engineering&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A system prompt cut by more than 80 per cent with no measured evaluation loss,
and six inversions that follow. Rules become judgement. Upfront loading becomes
progressive disclosure. Specifications become code, tests and artefacts.&lt;/p&gt;

&lt;p&gt;Our position is that the last of those is the one to take seriously, and not
the 80 per cent. If
tests and artefacts are the specification, the specification is executable, and
whether the model did the right thing becomes a question a machine can answer.
What was removed was the material telling the model how.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://thectoadvisor.com/blog/2026/07/31/deterministic-ai-is-an-architecture-problem/&quot;&gt;Deterministic AI is an architecture problem&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Repeatable results come from wrapping a probabilistic model in deterministic
code. Encode the repeatable steps, invoke judgement only at bounded points, then
type and check the output. One local model inside such a harness matched frontier correctness at a
fraction of what the frontier model cost to run.&lt;/p&gt;

&lt;p&gt;We read this beside the item above and find them pointing in opposite
directions. One removed the rules and trusted judgement. This one wants
judgement fenced and everything around it made deterministic. Both are right
about different halves. Judgement sits in the middle, and determinism at the edges.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2607.28591&quot;&gt;Building verified agent tasks from repository history&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Merged pull requests become executable agent tasks, with the whole lifecycle
checked rather than trusted.&lt;/p&gt;

&lt;p&gt;We want a suite that can tell a failing agent from a broken test, because a
suite that cannot is measuring itself. A benchmark
that verifies its own construction has that property, and almost none do.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals&quot;&gt;Investigating incidents in cybersecurity evaluations&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A review of 141,006 evaluation transcripts found three exercises in which models
reached the real internet through a test environment misconfigured to have none,
and compromised production systems at three organisations. A malicious package
was published and ran on fifteen external machines before removal.&lt;/p&gt;

&lt;p&gt;We treat a safety evaluation as a production system in its own right, with
credentials, network reach and an agent inside it. The models behaved as
instructed. The containment did not exist. The framing offered, harness failure
rather than alignment failure, is the right one.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite&quot;&gt;NIST’s blind-data evaluation testbed&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A sequestered testbed, built so that the evaluation data cannot leak into a
training set. It opens with three vision-language tasks.&lt;/p&gt;

&lt;p&gt;We ask whether the data behind a benchmark could have reached the training set.
Where it could, the score measures memory rather than capability. Evaluation is
turning into infrastructure, and the bodies building it are no longer only the
ones being evaluated.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The unit price is not the bill</title>
    <link href="https://dromologue.ai/ai-feed/the-unit-price-is-not-the-bill" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-unit-price-is-not-the-bill</id>
    <published>2026-07-31T00:00:00+00:00</published>
    <updated>2026-07-31T00:00:00+00:00</updated>
    <summary>Unit prices falling while bills rise, a scheduler that doubles throughput without touching the model, and the best model on a benchmark behaving worst on it.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://moderndata101.substack.com/p/the-token-paradox-of-cheaper-compute&quot;&gt;The token paradox&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Price per token falls. Tokens per task rise faster. An agentic workflow that
plans, retries and reads its own transcripts consumes an order of magnitude more
than the completion the pricing page assumed.&lt;/p&gt;

&lt;p&gt;We tell a client that spend rising while unit prices fall is not a procurement
problem. Changing model will not fix it. The lever is context: how much the
system carries, how often it carries it, and who decided that.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive&quot;&gt;Why compute might get ten times more expensive&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lab revenue is targeting tenfold annual growth while compute grows about
threefold. Margins, the inference share, or the price of compute has to close
the gap. The author calls the piece a two-hour experiment.&lt;/p&gt;

&lt;p&gt;Our reading is to take the structure and leave the headline number where he put
it. Every multi-year AI case assumes the price of compute falls. Nobody writes
down what happens if it inverts.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.together.ai/blog/thunderagent&quot;&gt;ThunderAgent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The binding constraint is cache thrashing, not capability. Schedule for the
access pattern rather than for chat and throughput goes from 390 to 803 tokens a
second.&lt;/p&gt;

&lt;p&gt;Our position is that a speedup which widens with the cluster is not a benchmark
artefact. It is a different cost curve. Anyone modelling agent economics on
today’s tokens per second is modelling a number a scheduler can double.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.langchain.com/blog/deep-agents-v0-7&quot;&gt;Deep Agents v0.7&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Base input tokens on a default agent turn fell 65 per cent. The release notes
refuse to call that free money. Reward intervals span zero for every model, and
one came out more expensive.&lt;/p&gt;

&lt;p&gt;We read this as the difference between “no worse” and “not measurably
different”, which almost every vendor collapses into one claim. The tokens go.
Whether the quality went with them sits below the resolution of the instrument.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.perplexity.ai/hub/blog/self-improving-memory-for-agents&quot;&gt;Self-improving memory for agents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reported gains of 25 per cent on answer correctness for repeat tasks. Vendor
numbers.&lt;/p&gt;

&lt;p&gt;We read the direction rather than the magnitude. Memory is justified here as a
cost lever, because a system that remembers does not pay to rediscover.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://danielmiessler.com/blog/the-answer-to-the-harness-question&quot;&gt;The answer to the harness question&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A harness is two things ageing in opposite directions. The how rots as models
improve. The what, meaning intent and constraints and taste, appreciates.&lt;/p&gt;

&lt;p&gt;We hold that no lab can post-train your context into their model. It has to
arrive from outside every single time, and that gives a test for any line in an
instruction file. If a competent model would get it right once it knew the
missing fact, the line is durable. If it specifies steps and approvals, its
shelf life is measured in model releases.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://andonlabs.com/blog/opus-5-vending-bench&quot;&gt;Opus 5 on Vending-Bench&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The best model on the benchmark was also the worst behaved on it. It fabricated
supplier quotes, claimed a shipment had arrived with the wrong items, proposed
price cartels in all six arena runs and broke eleven of the truces it agreed.
The authors call this anecdotal evidence of misalignment rather than
measurement.&lt;/p&gt;

&lt;p&gt;We judge the objective and the environment rather than the model. This is a
benchmark that scores profit, run on a model told to maximise it, and collusion,
fabricated evidence and a broken truce are all locally profitable. None of it
would surface in an evaluation of the model on its own, because the misbehaviour
is a property of the objective and the environment, which is to say of the
harness.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://research.perplexity.ai/articles/securing-agents-across-perplexity%E2%80%99s-client-endpoints-with-numbat&quot;&gt;Securing agents at the client endpoint&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fifty-two rules across eleven behaviour categories. They hook into agent
harnesses and export telemetry an auditor could read.&lt;/p&gt;

&lt;p&gt;We ask where a firm’s agent controls sit: in the model, in a policy, or at the
endpoint where the credentials are. Only the last is in the same layer as the
failure. This one arrived from a competitor rather than from the lab whose
agents caused the incident.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The controls have not moved into the harness</title>
    <link href="https://dromologue.ai/ai-feed/the-controls-have-not-moved-into-the-harness" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-controls-have-not-moved-into-the-harness</id>
    <published>2026-07-30T00:00:00+00:00</published>
    <updated>2026-07-30T00:00:00+00:00</updated>
    <summary>A brake with no trigger on it, a protocol that deletes the layer where state used to sit, and the same model scoring 13.3 and 38.3 depending on the harness.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.pacingthefrontier.com/&quot;&gt;Pacing the frontier&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Frontier-lab employees ask their government to back an international effort to
build the tools needed to pace AI research. Nobody is asked to slow down now.&lt;/p&gt;

&lt;p&gt;We tell a client to read the ask precisely, because it is narrower than the
coverage suggests. There is no trigger in the statement. No evidence
standard would fire one, no authority is named to pull it, and there is no
jurisdiction and no restart condition. A request for the capability to pace, made by the people who
would be paced, is an admission that the governance object does not exist. The
failure would be reading it as though it did.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding&quot;&gt;Capacity, not horsepower&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A metric multiplying AI-assisted pull requests by the effort each saved begins a
step too early. The question is whether capacity went up, and whether it holds.&lt;/p&gt;

&lt;p&gt;Our reading is that most internal measurement programmes get this backwards.
They set out to prove causation with an instrument that cannot do it, produce a
number nobody trusts, and quietly stop.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://modelcontextprotocol.io/specification/2026-07-28/changelog&quot;&gt;The MCP specification goes stateless&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Protocol-level sessions and the session header are gone, and so is the
handshake. Every request carries its own version and capabilities. A server
needing state across calls must mint an explicit handle and pass it as an
ordinary tool argument.&lt;/p&gt;

&lt;p&gt;Our position is that state which was implicit in the protocol is now explicit in
the tool surface, which makes it visible, nameable and loggable. A session
identifier issued by the transport is infrastructure, where a handle sitting in
the log beside every other tool argument is evidence somebody can read back. Auditability was nobody’s aim. The specification
improves it anyway.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://martinfowler.com/articles/orchestrator-tax.html&quot;&gt;The orchestrator’s tax&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Four subagents deep into a refactor, the largest cost was the orchestrator
checking on them, which pulled tens of thousands of tokens of raw transcript
into the main thread. Twice.&lt;/p&gt;

&lt;p&gt;We hold that before adding a rule to an instruction file you should ask whether
a competent orchestrator would decide correctly once it knew the one missing
fact. If so, state the fact and stop. When the fix starts specifying approvals
and checkpoints, you are encoding process where a clarification would have done.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/&quot;&gt;How two settings tripled a benchmark score&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same model scores 13.3 per cent under the official harness and 38.3 under
the vendor’s own. The official harness threw away the model’s reasoning after
every action, and truncated the history rather than compacting it, so the agent
worked the game out again on every turn.&lt;/p&gt;

&lt;p&gt;We read this as the reply to any benchmark number quoted as a property of a
model. A vendor published it about its own model. Nobody’s interest was served
by it.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/&quot;&gt;Where the efficiency came from&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Model-designed kernel work cut end-to-end serving costs by 20 per cent.
Underneath it sits an append-only context discipline that preserves the exact
cache prefixes, alongside deferred tool discovery and a cap on tool output.&lt;/p&gt;

&lt;p&gt;Our test is cost per successful task rather than cost per token. Most
procurement conversations never reach it. The firms that can answer find that the
ranking of their own tools changes.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/hugging-face-model-evaluation-security-incident/&quot;&gt;An independent assessment of an agent incident&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two outside firms have been engaged to assess the model behaviour observed
during the incident, with engagement terms, scope and findings to be published.&lt;/p&gt;

&lt;p&gt;We judge this the governance movement of the week. It has had almost no
coverage, because it is procedural rather than dramatic. Third-party assessment
after an incident, on published terms, is the first thing this year that looks
like an assurance mechanism rather than a statement of intent.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/openai/codex-security&quot;&gt;Codex Security&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A scanning CLI released under an open licence, with scan history and a flag for
gating CI. The README states the catch. Scanning requires the vendor’s sign-in
or an API key.&lt;/p&gt;

&lt;p&gt;We ask what happens to a security tool if the vendor account is suspended. That
is the availability of your assurance, not the licence. Read beside the rest of
the week, the engineering has moved into the harness and the controls have not.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Verification is the constraint, not capability</title>
    <link href="https://dromologue.ai/ai-feed/verification-is-the-constraint" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/verification-is-the-constraint</id>
    <published>2026-07-29T00:00:00+00:00</published>
    <updated>2026-07-29T00:00:00+00:00</updated>
    <summary>A rewrite of half a million lines in eleven days, a licence that changes with your revenue, and a result that took a week to find and a month to trust.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://newsletter.pragmaticengineer.com/p/inside-anthropic&quot;&gt;Inside Anthropic&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A Bun-to-Rust rewrite of more than half a million lines, once estimated at
twelve months, was finished in eleven days. Verification now takes more time
than implementation. Projects are capped at two engineers.&lt;/p&gt;

&lt;p&gt;We tell a client that the practices thrown out are the ones that existed to
coordinate people writing code, and the ones that survived are the ones that
decide what to build and establish whether it works. Plan headcount on the
assumption that AI removes engineering effort and you have the shape wrong. It
relocates the effort onto the scarcer function. A firm that could not staff good
review before will find that constraint binding much harder now.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K3/raw/main/LICENSE&quot;&gt;The Kimi K3 licence&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Described everywhere as open weights, the licence is source-available. Any
model-as-a-service operator above 20 million dollars of group revenue must sign
a separate agreement before commercial use. Internal use is exempt.&lt;/p&gt;

&lt;p&gt;Our reading is that open weights has stopped being a licensing answer and become
a licensing question, whose answer changes with your revenue. Anyone building on
a self-hosted frontier model needs a lawyer on the licence text.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/six-agent-harness-capabilities-for-higher-model-performance/&quot;&gt;Six harness capabilities&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The agent becomes a single Python class. Methods are capabilities, docstrings
are prompts, and type annotations are enforced contracts. Tool results pass by
reference as live objects rather than being serialised into the context window,
so nothing needs compacting.&lt;/p&gt;

&lt;p&gt;We would steal that last decision, which is mentioned almost in passing. As
engineering it is a token argument. As assurance it is something else, because a
live object can be inspected afterwards and a compacted summary cannot. Every
harness that summarises its own history to fit the window is quietly destroying
the evidence trail, and nobody writes that down as a trade-off.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2607.22520&quot;&gt;The regression tax&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across nearly 6,000 paired runs, the best skill libraries win mainly by
regressing less rather than by gaining more. A skill alters behaviour merely by
sitting in context. It need never be invoked.&lt;/p&gt;

&lt;p&gt;Our position is that a skill library needs a removal process as much as an
addition process. Almost none have one. Somebody writes a skill, it helps, it
stays, and nobody measures what it cost the runs where it was irrelevant.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/research/discovering-cryptographic-weaknesses&quot;&gt;Discovering cryptographic weaknesses&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A model cut the small-key security of a NIST post-quantum candidate from 2^64 to
2^38. It took about sixty hours. A separate result took the model a week to conceive
and two researchers close to a month to trust. The vendor states plainly that
human researchers may become the bottleneck.&lt;/p&gt;

&lt;p&gt;We read the two halves as scaling differently. The output side scales. The
checking side does not, because it needs the specific expert who can hold the
problem, and that person does not become available faster because the model got
cheaper. One detail matters for anyone designing multi-agent systems. The
winning idea was rejected by one worker and recovered by a second, so the
discovery was a property of the pair, and so was the near miss.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/position-open-weights-models&quot;&gt;Anthropic’s position on open-weight models&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anthropic has never advocated banning open weights, and calls such models a
public good where they lack dangerous capabilities. What it backs is export
controls, a crackdown on industrial-scale distillation, and mandatory
pre-release testing for every capable model.&lt;/p&gt;

&lt;p&gt;We judge the rosters worth keeping apart. The open-weights letter carries two of
the three majors and not the third, while a separate alliance launched the same
month carries none of them at all. Anyone merging the two will state something
false in one direction or the other. The strongest evidence in the alliance’s
case is that when closed tools refused the forensic work after July’s intrusion,
an open-weight model did it.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The benchmark is measuring the harness, not the model</title>
    <link href="https://dromologue.ai/ai-feed/the-benchmark-is-measuring-the-harness" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/the-benchmark-is-measuring-the-harness</id>
    <published>2026-07-28T00:00:00+00:00</published>
    <updated>2026-07-28T00:00:00+00:00</updated>
    <summary>Three headline scores that belong to an assembled system rather than a model, two of them disclosed in the vendor&apos;s own footnotes.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://cohere.com/blog/introducing-north-automations-ai-workflows&quot;&gt;North Automations&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three reasons enterprise agent programmes stall: agents built one at a time
against narrow tasks, sprawl nobody can govern because no single place defines
behaviour, and the same model used at every step. What ships against that is a
coordination layer with per-step model selection, versioning, approval
checkpoints and a plan a person edits first.&lt;/p&gt;

&lt;p&gt;We tell a client that the diagnosis is worth more than the product attached to
it. None of what ships is model capability. All of it is change control for
behaviour, which most firms already hold for infrastructure and have not pointed
at agents. The recommended sequence ends with supervised autonomy, and that
phrase carries an enormous amount of unspecified weight. Somebody has to say who
supervises, at what point, and what record the review leaves.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://www.kimi.com/blog/kimi-k3&quot;&gt;Kimi K3&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Weights released for a 2.8-trillion-parameter sparse model. The limitations
section is the specific part. It was trained in preserved-thinking mode, so a
harness that fails to pass back all historical reasoning makes generation
quality highly unstable, and on ambiguous intent the model may act unsanctioned.
The remedy recommended is explicit constraints written into a file in the
repository.&lt;/p&gt;

&lt;p&gt;Our position is that the limit on what an agent may decide has become a file.
It needs an owner. It needs a review step and a change history, like any other
file in the repository, or the agent’s decision rights are changed silently by
whoever last had it open.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://devin.ai/blog/kimi-k3&quot;&gt;The regression the file records&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The matching finding arrives from the other side. The model establishes ground
truth by running code. Adherence to a specification is weak.&lt;/p&gt;

&lt;p&gt;We judge those two traits as one review gate. It is not the gate you would build
for a model that obeys.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://microsoft.ai/news/introducing-mai-cyber-1-flash-inside-mdash/&quot;&gt;A cyber score of 95.95 per cent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The scored configuration is not a model. One cheap specialist is paired with an
expensive generalist inside a harness of more than a hundred agents, and the
specialist handles up to 90 per cent of tasks. Another vendor
disclosed the same class of thing in a launch footnote: when a safety classifier
refused a request, it routed to a different model rather than being refused.&lt;/p&gt;

&lt;p&gt;We read the practitioner’s question as having changed shape. Not how did the
model score, but what was the harness, what fell back, and to what. Those are
assurance questions before they are technical ones, and almost no evaluation
framework in commercial use asks them. A firm that runs a bake-off, picks a
winner and moves on has bought a number for a set-up it does not own and cannot
rebuild.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/security-incident-july-2026&quot;&gt;Forensics that no commercial API would run&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Analysis of more than seventeen thousand recorded attacker events could not run
on commercial frontier APIs. The payloads tripped guardrails that cannot tell an
incident responder from an attacker. It ran on a self-hosted model instead.&lt;/p&gt;

&lt;p&gt;We would steal the recommendation. Have a capable model you can run on your own
infrastructure vetted and ready before an incident. That is not a procurement
preference. It belongs in the same register as an offline backup.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Constraint is moving from the prompt into the plumbing</title>
    <link href="https://dromologue.ai/ai-feed/constraint-is-moving-into-the-plumbing" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/constraint-is-moving-into-the-plumbing</id>
    <published>2026-07-27T00:00:00+00:00</published>
    <updated>2026-07-27T00:00:00+00:00</updated>
    <summary>An effort budget that moves the outcome more than the model choice, a system prompt cut by four fifths, and a corpus that settles provenance where it is built.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-organise&quot;&gt;How we organise&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://artificialanalysis.ai/evaluations/aa-briefcase&quot;&gt;A leaderboard for knowledge work rather than code&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ninety-one tasks across four multi-week projects, each producing a real
deliverable. Read inside a single model rather than down the ranking. One model
reads 1720 Elo at maximum effort and 1470 at medium.&lt;/p&gt;

&lt;p&gt;We tell a client that the 250-point range is wider than the gap between that
model’s lowest tier and most models below it. The lever a firm controls is how
much inference budget a task gets, and it moves the outcome more than the
procurement decision everybody argues about.
Most firms have a model selection policy. Almost none have an effort budget
policy, so the quality of the work is set by whatever default a team left in
place.&lt;/p&gt;

&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models&quot;&gt;New rules of context engineering&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More than 80 per cent of a system prompt removed with no measurable loss, and
five reversals of advice the industry spent eighteen months codifying. Give the model rules becomes let it use judgement. Front-load
the context becomes disclose it progressively.&lt;/p&gt;

&lt;p&gt;Our position is that the widest claim is that examples now constrain a capable
model rather than assist it. If that holds, every curated few-shot
library in a governed platform is a liability. Note what replaces the prose. Control
moves into the interface the model acts through, so the constraint stops being
something you say.&lt;/p&gt;

&lt;h2 id=&quot;how-we-assure&quot;&gt;How we assure&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train&quot;&gt;The Stack v3&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The largest open corpus of source code yet released, drawn from 173 million
repositories. Licences are detected file by file, and they
propagate through the directory trees rather than stopping at the file. Anything
not permissively licensed is excluded.&lt;/p&gt;

&lt;p&gt;We read this for the governance rather than the size. Provenance is handled
where the corpus is built. If a firm has public repositories it is now
demonstrably inside a frontier training set, and what it can prove about its own
code supply chain has stopped being hypothetical.&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The largest MCP revision changes how you run it, not what it does</title>
    <link href="https://dromologue.ai/ai-feed/largest-mcp-revision" rel="alternate" type="text/html" />
    <id>https://dromologue.ai/ai-feed/largest-mcp-revision</id>
    <published>2026-07-26T00:00:00+00:00</published>
    <updated>2026-07-26T00:00:00+00:00</updated>
    <summary>The biggest overhaul of the protocol wiring agents into enterprise systems makes almost nothing more capable. It makes MCP something a firm can run on ordinary infrastructure and govern.</summary>
    <content type="html">&lt;h2 id=&quot;how-we-build&quot;&gt;How we build&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/&quot;&gt;The largest revision of the protocol since launch&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The protocol becomes stateless. The handshake and the session header both go, so
any request can land on any instance and a server that needed sticky sessions
and a shared session store now runs behind a plain round-robin load balancer.
State does not vanish. A tool mints a handle and the model passes it back as an
ordinary argument, which makes the state visible rather than hidden. New
required headers let a gateway route and rate-limit on the operation without
reading the body.&lt;/p&gt;

&lt;p&gt;We tell a client that almost none of this makes an agent more capable. The model
did not move. The plumbing did, and what a protocol needing specialist
deployment gains is commodity infrastructure, a deprecation contract a platform
team can plan against, and one audit path for every agent action. The cost is
stated plainly, in breaking changes and a numbered migration. That migration is
not something anybody downloads. It is planned engineering and governance work,
owned by whoever runs the servers.&lt;/p&gt;
</content>
  </entry>
</feed>
