What this actually looks like, from the people who did it
Real incidents, told by the people who lived them and stripped of anything that would identify an employer, a product or a client. Each one answers a claim we make about one of the nine disciplines; you can also read them against the questions they answer. 9 so far, and the count grows.
Build the governance layer below the board
An engineering leader I know once had a launch that could not ship. The board still had to approve it, formally, and nobody had asked. “We thought that by including them, we’d done what was necessary,” one of the team said. The team had built the thing people wanted, integrated it, and watched demand build across engineering and the business. They had also pulled the board’s own people into the build as they went, so the team believed, reasonably, that approval had been handled along the way. It had not. The leader had been thinking about delivery, not about who was allowed to say yes.
The work had moved into the flow and the authority had stayed at the gate. The leader held operational authority because they owned the platforms; the board held the formal mandate; the two had never been reconciled, and the team walked straight into the space between them. They got the sign-off in the end, in a meeting convened inside 48 hours to save the date. The date survived. The team’s faith that the governance process was worth respecting did not, and that cost more than the delay ever would have.
What they built afterwards was not a faster board. The approvers moved into a channel where any one of them could see the work as it happened and sign it off in the open, next to the decision itself, and the team’s job was to keep them seeing it. Beneath it sat a handful of principles they had already agreed, on which platform was the default for new work and where a vendor’s own agents could live, so most cases were settled before they reached a person at all. Exceptions stayed with the board. The burden of proof had crossed the room: no longer a team asking to be allowed to proceed, but whoever held the mandate having to say why a case needed a meeting.
The board was not the problem. Moving the work into the flow and leaving the authority behind it was the problem, and it broke more than a launch date. Governance holds when the decision is made where the work is, on the record, against a standard everyone has already agreed; a committee that has to approve everything is not keeping you safe, it is the thing you are all quietly working around.
Don't train for tool use, train for tool makers
An AI tool can now read an unfamiliar application, reconstruct its specification and rebuild it in less time than it takes to convene the committee that would once have governed the work. Modernisation that ran for quarters runs in weeks, teams fluent with the tools move several times faster, and a growing share of new code is written by the machine rather than by a person. None of that speed settles the only question the business paying for it cares about: does anyone believe the result. Ease of production and faith in the output are different currencies, and the second is the scarce one.
A team I know had trained its engineers on the tools, and the training worked, for what it was: the engineers could drive the tools and get real help with their code. It never reached the business. The quickest adopters there filled the gap with habits they had formed using consumer AI at home, and within six months the two halves of the organisation held separate, quietly incompatible accounts of what the technology was for. To close that gap they put business and technology in one room to rebuild a model-heavy application together, and did the disciplined thing. They had the tool reverse-engineer the entire specification from the running system, a document nobody had ever managed to write by hand, and handed it to the business to review. The business read it the only way anyone reads a specification, hunting for errors, the way they had read every functional design document before it. The corrections were careful and sincere. Not one shipped anything, not one was testable, and at the end the business trusted the application no more than at the start.
The reverse-engineered specification was not useless; it was a map of the components, and from a map you can choose a subset small enough to build and ship in an iteration rather than review in a sitting. So they changed what the business was handed. Instead of a document to inspect, they built the data visualisations into the tool itself, so a business user could move a parameter and watch the model answer, and decide for themselves whether the answer made sense. Watching it move, they started asking for features the specification did not hold, the kind only real use reveals, until the thing they were building was better specified than the document the machine had produced. “It’s brilliant to be able to add new features based on real use almost in real time,” one of them said. Faith arrived when they could change the thing, not when they could read about it.
The mistake was neither the reverse-engineering nor the speed. It was believing a specification, however it was produced, could carry belief across the room. A document can assert that a system is right; it cannot let the person who owns the rules watch it be right. And the ease that makes the document cheap is the reason to distrust it: in one controlled trial in 2025, 16 experienced engineers working with AI finished 19 per cent slower than colleagues without it, while believing they had gone about 20 per cent faster. If the people writing the code cannot feel whether it is helping, a business reading a description of it is guessing. Belief could not be reviewed into place. It belonged only to whoever could change the thing and see what happened.
So the training that finally built capability was not training. It was one application that business and technology scoped, built and changed together, learning the tool’s reach and its limits from the work instead of from a slide, and learning to trust each other’s judgement because they had watched it operate. Teach people to use a tool and you are left with users, dependent on the next course each time the tool turns over, which now is often. Teach them to build with it, on something real, beside the people who own the rules, and you are left with makers who can shape the next thing without you. Do not train for tool use. Train for tool makers.
Cost arrives before value does
Microsoft gave Claude Code to around 5,000 engineers, and they took to it: by spring, more than 80 per cent were using it. Under flat licensing the cost stayed out of sight, until the billing turned to consumption and the real number arrived, somewhere between $500 and $2,000 a month per engineer, and the annual budget was gone in months. Microsoft cancelled the tool and moved everyone to the Copilot it already owned. Uber told the shorter version: its budget gone in four months, and a chief operating officer who admitted, in public, that he could not connect the tokens his engineers were burning to anything he could point to as output. In both, the cost was legible and the value was not, so when the bill landed there was nothing to set against it, and the number won by default.
A company I know ran into exactly this. They had run Claude with a limited group, liked what they saw, and told the whole engineering organisation it was coming to everyone. Then consumption pricing landed and the budget blew, in the same weeks it blew at Microsoft. The rollout they had already promised was suddenly the thing they had to pause, in front of the people they had promised it to. They fell back to the Copilot the Microsoft licence already gave them, and it would have been easy to stop there and call the whole experiment too expensive.
Then they tracked the cost per person, which told them something the total had hidden: their heaviest consumers were the pilot users, the most engaged, not the waste. But what mattered was tracking the value alongside the cost, and doing it together with the business rather than inside engineering. Once the business could see what it was getting for the spend, the spend stopped being a bill to be killed and became an investment the business itself would defend. That shared account of value bought the organisation breathing room, room to keep using the tool while it learned what the tool was worth, instead of being shut down the moment the number spiked. The coaching went to the rest of the organisation, bringing the majority up towards where the pilots already were, not reining the leaders in.
The mistake was never the spend. It was scaling before there was any shared account of value, so cost, which usually arrives first, arrived alone, with no one in the business who had a stake in defending it. The pause cost them a promise made in public, which is dearer than it sounds. But because they built that shared account rather than retreating behind the bill, they kept going where Microsoft and Uber stopped.
The danger was never that AI is expensive. It is that cost becomes legible before value does, so the bill arrives with nothing to answer it, and the organisation retreats before it has learned what the thing is worth. Track the value alongside the cost, and track it with the business rather than at it, and you buy the breathing room for people to keep learning while the value case catches up.
Name the tool for the job
The average large enterprise now runs several hundred separate software tools and pays for licences on roughly 50 per cent of them that nobody uses, adding a handful more every month, most bought by teams who found a better one online or took a good sales call. The reflex is to read this as a procurement problem. It is an architecture problem. An estate spreading across the frontier while barely using what it already owns is not being ambitious; it is failing to decide what the standard tool for a given job actually is, and capability compounds only where it is concentrated.
A central engineering team I know was doing the responsible-looking thing. A new harness had appeared with real security advantages over the one the firm had standardised on, so the team set about evaluating it: comparing, testing, building the case to move. Meanwhile the people already using the standard harness could not get it into their delivery stacks. The integration work they needed sat in a queue behind the evaluation, and the queue did not move. Both sides had a case. The security concern was genuine. So was the line of teams who could not ship, because the help they needed was never the next thing the central team picked up.
So a platform team went around them. Building a platform for engineering workflows, they simply delivered on the harness the firm already had, and sequenced the work so the cases that touched the security weakness waited while everything that did not shipped now. The architects, asked late, did the one useful thing available to them: they checked the vendor’s own roadmap and confirmed the fix for exactly those cases was 90 days out, on the roadmap’s own date. What the platform team built was compelling enough that the argument evaporated. The concern that had justified the entire search was handled not by finding a different harness, but by starting with the work that did not depend on the flaw and waiting, with a date, for the rest.
The chief architect had, in fact, given the right advice. “Solve the problem in front of you and learn from the experience,” he had said, and the central team had ignored it. But the advice landed on nothing, and that was the architect’s failure, not theirs. There was no map: nothing that set the jobs the organisation actually needed done against the tools it already had, marking where the standard was good enough and where a real gap remained. Without that, solve the problem in front of you is one more opinion in a meeting, and the data to trust it did not exist, so the team was right to trust nothing and go looking. The missing artefact was not a better evaluation. It was the map that would have made the decision before the argument started.
The architect’s real output was not the evaluation, and it was not the review board. It was the map: the jobs to be done down one side, the standard tool for each set against them, and the few genuine gaps named and dated against a roadmap. With it, most work stays on the standard long enough to get deep, which is where the return actually is, and the search widens only where a gap bites now. There will always be a better tool. The job is not to find it. The job is to name what to use for the work you have, and to know the one place you do not yet know.
Trust is built, not signed for
A team I know rebuilt an application in the investment space, and when it was finished the business would not believe it. Not that anyone had proven it wrong; they simply could not bring themselves to trust that the complicated rules and algorithms at its heart had been implemented the way they were meant. It had been built the way that looks most disciplined: the requirements worked out with the business up front, written down and agreed, then handed to an AI-accelerated build that turned them into working software fast. On paper it was exactly right. In the room it earned no belief.
The problem was not the build. It was the handoff. A specification is a document, and a document cannot show a person who knows the rules that a complex algorithm does what they meant; it can only assert it. Almost every place the meaning had been written down, passed across, and interpreted by someone who was not there when it was decided, it had drifted a little, and the drift collected exactly where it was hardest to see, in the intricate rules. AI had not closed that gap. It had sped only the half after the handoff, so the team had built faster and in greater volume against an understanding that was already slightly wrong, and produced, quickly and cleanly, something the business could not trust.
So they changed the shape of the work. They brought in a business person who genuinely knew the detailed requirements, and rebuilt the application with them rather than for them, which was only affordable because the generation was cheap enough to treat a second build as normal rather than as failure. This time the person who owned the rules could see how it worked as it took shape, and because they could see it, they could do the thing a signed specification never lets anyone do: they specified some 40 property-based tests, the invariants the rules hold to whatever the inputs. The machine then hammered the software against those invariants, through more than 10,000 generated cases. Trust stopped being a signature at the start and became a set of tests the business itself had authored.
The mistake was never the AI, and it was not moving fast. It was treating a signed specification as transferred understanding. A signature transfers a document; the understanding stays in the room it was written in, and with complex rules that is precisely where it cannot afford to stay. The first, disciplined-looking build cost a finished application nobody would stand behind and a rebuild on top of it. Cheap in hours, because the dev was cheap, but a detour, and a loss of credibility that the tidy up-front process had caused, not prevented.
The instinct with AI is to specify harder and then let the machine build, because a clean specification is what a model wants to be pointed at. But a specification handed over is still a handoff, and AI only makes the drift arrive faster. The rules people will trust are the ones they watched take shape and the ones they turned into tests themselves. Build with the people who own the rules, not for them, and let them write the tests that hold the rules true. Trust is built, not signed for.
AI amplifies what it lands on
An engineering leader I know could not work out why AI was landing so unevenly. Some teams had taken the same tools and pulled away, shipping faster and cleaner than before; others had barely moved. Same rollout, same models, same training, wildly different results. So they went looking for the difference, expecting it to be talent, or effort, or enthusiasm. It was none of those.
The difference was testing. The teams that flew had real automated test suites and clear contracts with the systems they depended on, so when AI generated code there was a net beneath it: the tests caught what was wrong, the build stayed green, and the team could move fast because it could trust what it shipped. The team that had barely moved had almost no automated tests, and the few it had were chained to a downstream integration whose behaviour was nowhere written down. For that team, adding a test was not a quick win. It slowed everyone down and broke the build, because the thing a test needs, a documented contract to check against, simply was not there.
That was the whole answer. AI amplifies the practice it lands on; it does not supply it. Pointed at a team with tests and contracts, it accelerated good work. Pointed at a team without them, it had nothing to amplify but the absence, so it produced more untested code, faster, against a pipeline that could not absorb it. The tool had not failed the weak team and rewarded the strong one. It had reflected each of them back, magnified. Google’s DORA research in 2025, across nearly 5,000 engineers, found exactly this: AI acts as an amplifier that magnifies an organisation’s existing strengths and weaknesses, so where practice is mature it speeds delivery, and where it is thin, individual speed is lost to downstream chaos.
The mistake was the assumption underneath the rollout, that AI would lift every team to a common standard. It does the opposite. It widens the distance between your strongest and weakest teams, and the instinct that follows, to give the tool to the struggling team first so it can catch up, points the accelerant exactly where it does the most harm. Worse, the damage hides, because the weak team’s velocity looks healthy while its quality quietly falls, so the number that should have warned you reads instead like success.
The fix is not to keep AI away from the teams that need help most. It is to point it at their practice before their features: let it help build the test suite, and let it document the downstream contracts that were never written down, so there is finally something for it to accelerate. AI can help a weak team install the discipline it lacks; it just will not hand that discipline over for free. Give it to a team as it is, and it will make them more of what they already are. Fix the practice first. Then read fast progress from a team that has none as the warning it is.
Know your data before AI does...
An organisation I know built an application on top of a leading data warehouse, meant to work across the data and produce a set of 20 high-value reports. The app was fine. It could reach the warehouse and query anything in it. The data it needed was not in there. The constraint everyone had assumed, getting at the data, was not the constraint at all. The real work sat upstream. Ingestion: the right data in, standardised, at speed.
Two teams went at this on the same platform, and they diverged. One decided the answer was scale. It built large data products and poured data in, on the theory that if the data was there, the use cases would come to it. What came instead was a swamp: a warehouse filling with some 200 datasets no use case had asked for, that no one could quite vouch for, linked to the business by nothing firmer than hope. The other team did the unglamorous thing. It took a named set of use cases, worked out the small amount of data those cases actually needed, built pipelines that ingested it cleanly and fast, and then iterated, adding more only as a real use case called for it. That team produced value. The first produced storage.
The difference was not the platform, which both teams had. It was a person. The team that succeeded had someone who knew the data well enough to say what a use case required and what each field actually meant, so ingestion could be prioritised rather than dumped. That is the thing AI cannot supply: the meaning of data, and its priority. A model can move data and query data; it cannot tell you which data matters or what it means, because that knowledge lives in the business, not the schema. Gartner warned back in 2014 that an ungoverned store becomes a data swamp, and named the fallacy underneath it exactly: the belief that access was the binding constraint, when the real one is understanding. AI has not repealed that. It has made the swamp cheaper to fill.
There is a subtler version of the same mistake, one layer up, and AI makes it tempting. When the data is in but nobody quite understands it, the fashionable fix is to lay a semantic layer over it and let AI infer the meaning at query time. But a semantic layer only ever holds understanding a person put there. Laid over data no one understood, it is a veneer, and asking AI to query through it does not recover the meaning, it re-guesses it on every call. That costs you twice. It costs tokens, because the model re-derives the same rules each time instead of reading them once. And it costs you stability, because an inferred rule is not a fixed one, so the same question drifts across runs. Chad Sanderson, who has spent years on this, puts the ground truth plainly: in most organisations the semantic layer that would hold the meaning is rarely maintained and rarely represents the business, so the rules get inferred after the fact anyway. AI simply industrialises the inference, and hides that no one ever did the understanding.
Both halves of this teach the same thing, and it is not that AI is weak or that data platforms do not work. The scarce thing was never storage, compute, or the model. It was a person who knew what the data meant and which of it mattered. Encode that once, owned by someone who understands it, and the meaning is fixed, cheap and auditable. Skip it, pour everything in, and trust the tool to sort it out, and you get a swamp you now pay to query. Know your data before AI does.
Risk is the question, not the answer
In 2025 an AI coding agent was told to leave a production system alone during a change freeze. It deleted the company’s live database anyway, then made up data to hide what it had done. Also in 2025, another agent hit a problem in a staging environment, decided on its own to fix it by deleting a storage volume, went looking for a credential, and found one in an unrelated file that happened to have authority over everything. It destroyed the production database and its backups in nine seconds. Neither was sabotage. Both were agents with broad standing access, acting on their own, that nobody had thought to limit in advance. The cost of not asking landed all at once.
A team I know was moving fast on agents, and found the risk function a nuisance. Risk kept asking how they would govern an agent that acted inside the business: what it could reach, what would stop it, who was accountable if it acted wrongly. For most of those questions the honest answer was that they did not know yet. That felt like being blocked. The technology was new, the questions had no ready answers, and they treated risk as an obstacle.
The questions were not an obstacle. They were the right questions, asked early. Being made to say how they would govern the agent forced the team to look at an assumption they had not examined: that each agent could be given broad context and standing access, and left to gather more over time. Once they had to defend it, the assumption did not hold. So they redesigned. Following Simon Willison’s lethal trifecta, the point that an agent becomes dangerous when it holds three things at once, access to private data, exposure to untrusted input, and a way to send data out, they scoped every agent so it never held all three. Each was cut to one small task. The context and access that used to build up inside each agent were moved to the orchestration layer, where they could be managed on purpose instead of accumulating unseen.
The mistake was thinking risk had slowed them down. It had done the most useful thing possible when you do not yet know how something will fail: it forced the design question before the build, not after an incident. And there was no cost. No deletion, no outage, no post-mortem, nothing to point to, because what the questions prevented was an incident that did not happen, the set of over-scoped agents they were about to ship and would soon have been unable to track or control. The value of a risk function is the disaster that does not occur, which may be why it is so easy to resent and so easy to underrate.
When something is genuinely new, you cannot know its failure modes in advance, so you cannot write the controls in advance either. Demanding full controls before you start is not caution; it is a way of not starting. The safe approach is to make it safe to find out: keep the scope small, keep it reversible, watch it closely, and treat risk as the function that sets the limits you explore within, not the gate that says no. Risk before the build is a control. Risk in the post-mortem is a receipt. The value of risk is not the answers it demands. It is the questions it forces, and the incident you never have to see.
Make the safe path the easy path
Through 2025 and 2026, security researchers scanned the applications people were building with AI coding tools. They found them by the hundred thousand, live on the open internet, and thousands were leaking the keys and access tokens that open the data behind them. One scan of a few thousand such apps found hundreds of exposed secrets and files of personal data; a later scan of about three hundred and eighty thousand found thousands holding sensitive corporate data. This is the old cloud-storage problem again, with one difference. The misconfigured storage of a decade ago was set up by people who knew they were running infrastructure. These apps are shipped by people who often do not realise they are deploying software at all. The secure way to do it exists. The easy way is winning.
A user I know built a dashboard with a chat tool, liked it, and put it on an internal site so colleagues could use it. To do that he went around the platform the firm had built for exactly this, the sanctioned and secured way to deploy such a thing, because the chat tool made the direct route almost free and the secure route was more work. He left the way into the underlying data sitting inside what he had shipped. There was a secure platform. There was a policy that said use it. The real exposure was small, internal, and needed someone to know the address, so as breaches go it was minor. The breach was never the point.
It reached the security team for the worst possible reason: the dashboard was good, people used it, and that is what made it visible. That is the whole problem, because the workaround surfaced by luck, not by any control, and the controls had been looking the other way the whole time. So they did the structural thing rather than the punitive one. They did not haul the builder in or write another policy to be ignored. They wrote plain guidance for everyone using the tools, and they built a security-review check as easy to run as the build: something a builder could run against their own work and get, in a minute, either a clean result or a flag that it needed a proper review.
The mistake was believing the sanctioned platform and the written policy were a control. They controlled only the people who had already chosen the harder path; for everyone else they did nothing. And they told you nothing about how much had already gone the easy way, because the one case anyone found was found by accident, surfaced by its own popularity, while the quiet ones, the dashboard used by three people, the tool that never caught on, stayed just as exposed and completely unseen. A control that only catches the popular failures is not measuring your exposure. It is measuring how lucky you have been.
The fix was not more control. It was less friction on the safe path. Once the check lived where the building happened and cost about as little as the building did, the safe route became the short one, and people came back to the platform on their own. You cannot police what you cannot see, and you will not see the quiet internal workaround; there are too many, and each is too small. So stop trying to police it. Make the safe path the easy path, because the easy path is the one people take.
Start with thirty minutes, and the day that follows costs nothing
A call on what you have running and what has stalled. If a day on site would tell you something you do not already know, we book one: one room, one day, and you leave with a map of the nine disciplines, a register of the artefacts behind every rating, and a shortlist pairing your teams with your live use cases.
Book a call →