AI Feed

August 2026

33 entries from August 2026. The latest entries are on the AI Feed itself.

  • 31 August 2026 The person watching said yes

    A module-by-module costing of text-to-SQL pipelines, a coupling inside Claude Code skills that no catalogue records, and 133 overreaching agent actions that ran because a human ...

  • 30 August 2026 The list had no line for it

    A nineteen-day discount that moved token volumes almost fourteenfold, a text-to-SQL model that beat every scaffold by being trained instead, and a safety classifier that approve...

  • 29 August 2026 It held for one round

    Eight months of prompts inside one firm, a code review that runs past its opening exchange, and a model that calls a question unanswerable and then answers it.

  • 28 August 2026 It was still written down

    A capability list kept in the wrong unit, safety rules summarised out of an agent's context, and tool output that arrives reading as an instruction.

  • 27 August 2026 The speed was the same for everyone

    A maturity study where velocity rose evenly and complexity did not, credentials that expire on their own, and a review of what testing assumes about the thing it tests.

  • 26 August 2026 It passed by not doing the work

    A migration benchmark, a parser that ran what it was given, and a playbook that timed its own approvals.

  • 25 August 2026 The price fell and the bill went up

    Three numbers a buyer checks this week turned out to be measuring the apparatus rather than the thing bought.

  • 24 August 2026 Same credential, different right answer

    The tooling is adopting scoped agent identity just as a new draft shows that a valid identity still cannot decide whether an action is allowed now.

  • 23 August 2026 Nobody had to form a view

    A reading of four agent tools against eighteen policy documents finds the platform controls and the contracts disagreeing about who answers for what an agent ships, and one prov...

  • 22 August 2026 It came back in the right format

    A study of 1,250 workplace interviews finds professionals guarding the signals that carry identity and freely obscuring the ones that carry effort. Coding agents lose up to 6.7 ...

  • 21 August 2026 The agents were reading something else

    A study of 557 agentic coding sessions found instruction files and working notes take 60.5 per cent of everything agents do with documentation, against 1.3 per cent for API refe...

  • 20 August 2026 None of it applied to everyone

    RevenueCat ranked 3,519 AI-powered apps by retention and found the category rate describes almost none of them. DX measured pull-request cycle time against throughput and found ...

  • 19 August 2026 Nothing was taken away to make room

    Linear published six years of its own product data and found AI work added on top of existing work, with the time teams spend deciding what to build unmoved. DiG-bench gave huma...

  • 18 August 2026 The exact number was the wrong one

    PointFive ran 2,908 paid Claude Code sessions and found that the harder a tool compressed the prompt, the more the work cost. Hugging Face set the most downloaded open models ag...

  • 17 August 2026 Nothing recorded the reason

    Nvidia has signed memorandums of understanding with six of the largest asset managers and banks to mobilise over $500 billion behind compute. A study of 1,867 repositories finds...

  • 16 August 2026 The evidence had to exist already

    One operator instrumented his own traffic and found 214 unseen page loads for every visible one, with the largest crawler referring nobody at all. Z.ai is holding GLM-5.3's open...

  • 15 August 2026 The label is the part you own

    The EU's transparency duty has applied since 2 August, and the half that lands on an ordinary organisation is labelling what it publishes rather than watermarking what it genera...

  • 14 August 2026 The second agent was never tested

    Eight in ten leaders say agents have already delivered a return, and the barriers they name are integration, cost and data quality rather than the model. Cloudflare is shipping ...

  • 13 August 2026 The reasoning was not sealed

    Researchers replayed encrypted reasoning traces from three frontier providers into weaker models and read the hidden text back in plaintext, along with hundreds of private items...

  • 12 August 2026 The record looked right

    Agents are running real back-office work at scale, and both failures a16z found in production are failures of checking rather than of capability. Meta has put a capable agent mo...

  • 11 August 2026 The control was a habit

    Anthropic measured the permission prompt that most organisations count as their control on coding agents. Testers caught 13.6 per cent of dangerous commands and the classifier r...

  • 10 August 2026 The worm brought its own model

    A research worm writes a fresh exploit for every machine it meets, runs on compute stolen from the machines it has already taken, and makes every control held at a vendor's API ...

  • 9 August 2026 Two thirds of the spend bought nothing

    A measured agent loop spent two thirds of its budget on turns that moved the score by nothing. Neither the loop nor the person running it knew until the trace was read afterward...

  • 8 August 2026 The deadline moved, the tooling did not

    Google made agent identity generally available and gave every agent a unique cryptographic identity. LangChain shipped an identity model so an agent learns who triggered it from...

  • 7 August 2026 Sixteen thousand merges, blocked

    Cloudflare's own code-review agents blocked 16,000 merges in four months and flagged nearly a quarter of a million problems. Anthropic cut false positives on biology questions b...

  • 6 August 2026 The default changed, not the policy

    Zed turned sandboxing on for every user rather than documenting how to enable it. Databricks shipped hard spend caps as a product. A spreadsheet agent got 89.69 per cent of its ...

  • 5 August 2026 The only control that held was a person

    npm's provenance system signed a malicious build of a package with 619 million monthly downloads, and it signed it correctly, because the attacker held the maintainer's GitHub a...

  • 5 August 2026 The model you bought is not the model you got

    Artificial Analysis published an index measuring how much accuracy a model loses depending on which provider serves it. A developer found his coding model burning 2.25 times the...

  • 4 August 2026 The model did not change, the system did

    OpenAI cut a model's price by 80 per cent and tripled an agent benchmark score without touching the model. Qwen is open-weighting a 2.4 trillion parameter model trained against ...

  • 4 August 2026 Fifty-eight per cent were never once right

    Mercor and Ramp ran 160 real month-end-close tasks past every frontier model, eight times each. 58 per cent were never solved correctly on any run, and the most consistent model...

  • 3 August 2026 The second opinion had the same blind spot

    A false disproof of the Collatz conjecture passed Lean's kernel, then passed the independent external checker too, because two unrelated bugs lined up. Cross-checking survived, ...

  • 2 August 2026 Sixty hours to find it, a month to believe it

    Anthropic's model broke a NIST candidate signature scheme in sixty hours. Two researchers then spent nearly a month establishing that the method was correct. Matthew Green, revi...

  • 1 August 2026 Nobody audits a model, they audit a harness

    The EU starts enforcing the AI Act tomorrow. This week Anthropic disclosed that its own evaluation harness, not its models, reached the open internet and compromised three real ...