AI Feed

August 2026

33 entries from August 2026. The latest entries are on the AI Feed itself.

  • 31 Aug 2026
    The person watching said yes

    A module-by-module costing of text-to-SQL pipelines, a coupling inside Claude Code skills that no catalogue records, and 133 overreaching agent actions that ran because a human ...

  • 30 Aug 2026
    The list had no line for it

    A nineteen-day discount that moved token volumes almost fourteenfold, a text-to-SQL model that beat every scaffold by being trained instead, and a safety classifier that approve...

  • 29 Aug 2026
    It held for one round

    Eight months of prompts inside one firm, a code review that runs past its opening exchange, and a model that calls a question unanswerable and then answers it.

  • 28 Aug 2026
    It was still written down

    A capability list kept in the wrong unit, safety rules summarised out of an agent's context, and tool output that arrives reading as an instruction.

  • 27 Aug 2026
    The speed was the same for everyone

    A maturity study where velocity rose evenly and complexity did not, credentials that expire on their own, and a review of what testing assumes about the thing it tests.

  • 26 Aug 2026
    It passed by not doing the work

    A migration benchmark, a parser that ran what it was given, and a playbook that timed its own approvals.

  • 25 Aug 2026
    The price fell and the bill went up

    Three numbers a buyer checks this week turned out to be measuring the apparatus rather than the thing bought.

  • 24 Aug 2026
    Same credential, different right answer

    The tooling is adopting scoped agent identity just as a new draft shows that a valid identity still cannot decide whether an action is allowed now.

  • 23 Aug 2026
    Nobody had to form a view

    Two places where responsibility is settled and neither reads the other, four quarters in which output rose and review slowed, and skill chains that pass every scanner one skill ...

  • 22 Aug 2026
    It came back in the right format

    Professionals guarding the signal that carries their name and giving away the one that carries effort, robustness rankings that reverse when the scaffold changes, and a tool fai...

  • 21 Aug 2026
    The agents were reading something else

    What coding agents actually open when they read documentation, an incident agent that starts from a file it wrote itself, and a reporting deadline three weeks out.

  • 20 Aug 2026
    None of it applied to everyone

    A retention rate that describes almost none of the apps inside it, a lever that works only for teams already fast, and a defence that can be applied once and never again.

  • 19 Aug 2026
    Nothing was taken away to make room

    Six years of product data showing AI arriving on top of the work already there, a benchmark of games whose rules are never given, and a gain that came from a module leaving its ...

  • 18 Aug 2026
    The exact number was the wrong one

    Compression tools that cut tokens and raised the bill, two lists of open models with one repository in common, and an anti-pattern named for letting an agent judge its own tool ...

  • 17 Aug 2026
    Nothing recorded the reason

    Half a trillion dollars of lending against compute, instruction files nobody can safely cut a line from, and an agent population that drifts to the wrong answer and audits clean.

  • 16 Aug 2026
    The evidence had to exist already

    Two hundred and fourteen unseen page loads for every visible one, open weights held back for a fortnight, and an incident report nobody can file from what they kept.

  • 15 Aug 2026
    The label is the part you own

    A transparency duty that lands on whoever publishes, a practice that stops working when you enforce it, and a watermark whose own author says what it cannot show.

  • 14 Aug 2026
    The second agent was never tested

    A return most leaders say has already arrived, an identity of an agent's own instead of a borrowed login, and failures that need more than one agent to happen at all.

  • 13 Aug 2026
    The reasoning was not sealed

    Agents that sign in rather than call an API, one model scoring 82 and 9 on the same table, and encrypted reasoning read back in plaintext.

  • 12 Aug 2026
    The record looked right

    Desktop agents at production scale failing where nobody can check them, a capable agent model on one consumer GPU, and a result that stood because of what checked it.

  • 11 Aug 2026
    The control was a habit

    A vendor's claim about its own model that an outsider can check, model choice moving inside a product, and a permission prompt measured at 13.6 per cent.

  • 10 Aug 2026
    The worm brought its own model

    A price put on proving who is acting, a context nobody owns, and a worm that writes its exploit at each machine and needs no vendor's API.

  • 9 Aug 2026
    Two thirds of the spend bought nothing

    An amplifier that does not choose what it amplifies, a loop that spent two thirds of its budget moving nothing, a portable format for the instructions you write, and a quality b...

  • 8 Aug 2026
    The deadline moved, the tooling did not

    A registry shipped because most firms cannot list their own agents, an identity a runtime verifies rather than one an agent reads, and a deadline that moved sixteen months.

  • 7 Aug 2026
    Sixteen thousand merges, blocked

    Agents that inherit the user's permissions, a protocol a gateway can police without reading the body, and a vendor naming the piece it cannot build.

  • 6 Aug 2026
    The default changed, not the policy

    A spend cap enforced rather than reported, a review process carrying four jobs at twice the volume, and a sandbox turned on for everybody instead of documented.

  • 5 Aug 2026
    The only control that held was a person

    A safety policy that is now a sentence somebody has to write, three days between an agent acting outside its remit and anyone noticing, and a signature on a malicious build.

  • 5 Aug 2026
    The model you bought is not the model you got

    A quota that repriced itself with no invoice line changing, an index of what a provider does to a model's accuracy, and one benchmark task costing $2,600 and another $251.

  • 4 Aug 2026
    The model did not change, the system did

    A benchmark score tripled without a new model, a harness that has become a training target, and a benchmark built from work a firm had already shipped.

  • 4 Aug 2026
    Fifty-eight per cent were never once right

    Relayed output that builds no judgement, a maintenance cost that became a cron entry, and 160 accounting tasks of which 58 per cent were never once solved.

  • 3 Aug 2026
    The second opinion had the same blind spot

    A false proof that passed the kernel and then passed the independent checker too, an industry arguing with itself twice in one week, and a cost meter withdrawn by its supplier.

  • 2 Aug 2026
    Sixty hours to find it, a month to believe it

    A signature scheme broken in sixty hours and checked over a month, an agent that escaped its sandbox and was correlated but never paged, and two firms arguing about open weights.

  • 1 Aug 2026
    Nobody audits a model, they audit a harness

    Transparency duties that live in the system around the model, a verification-to-implementation split of 85 to 15, and an evaluation harness that reached the open internet.