The questions every AI programme hits
Every stalled AI programme sounds unique from the inside and asks the same questions from the outside. These are the ones technology and business leaders keep bringing us, grouped by the nine disciplines of an AI-native operating model: each with the claim we stand behind, and, where someone has lived it, their story. Pick a discipline. If more than a few of these are sitting on your desk, that is what the free diagnostic day maps.
How do I organise with AI?
Structure
Build five-person pods from both the business and technology, and let them ship frequently with fewer handoffs.
- Our AI is either trapped in one central team that has become a bottleneck, or sprawling everywhere ungoverned. How should we be organised?
- Should we build an AI Centre of Excellence?
- We reorganise every year and nothing improves. Why?
Our view
A boundary is real only when a team can take an idea to running software without asking permission. That test is what a centre of excellence usually fails. We bound one team of about five people to one capability. Every member and agent is named with the domain they hold, the team is one of four topology types, and it carries a cost line of its own. The decisions it can make alone are written down. The ones it cannot are measured: who decides, who owns it on each side, and how long the answer takes. Reorganisations that move boxes and leave decision rights alone change nothing. Your architecture will end up shaped like your organisation whether you intended it or not.
A Practitioner Story: Build the governance layer below the board
An engineering leader I know once had a launch that could not ship. The board still had to approve it, formally, and nobody had asked. “We thought that by including them, we’d done what was necessary,” one of the team said. The team had built the thing people wanted, integrated it, and watched demand build across engineering and the business. They had also pulled the board’s own people into the build as they went, so the team believed, reasonably, that approval had been handled along the way. It had not. The leader had been thinking about delivery, not about who was allowed to say yes.
The work had moved into the flow and the authority had stayed at the gate. The leader held operational authority because they owned the platforms; the board held the formal mandate; the two had never been reconciled, and the team walked straight into the space between them. They got the sign-off in the end, in a meeting convened inside 48 hours to save the date. The date survived. The team’s faith that the governance process was worth respecting did not, and that cost more than the delay ever would have.
What they built afterwards was not a faster board. The approvers moved into a channel where any one of them could see the work as it happened and sign it off in the open, next to the decision itself, and the team’s job was to keep them seeing it. Beneath it sat a handful of principles they had already agreed, on which platform was the default for new work and where a vendor’s own agents could live, so most cases were settled before they reached a person at all. Exceptions stayed with the board. The burden of proof had crossed the room: no longer a team asking to be allowed to proceed, but whoever held the mandate having to say why a case needed a meeting.
The board was not the problem. Moving the work into the flow and leaving the authority behind it was the problem, and it broke more than a launch date. Governance holds when the decision is made where the work is, on the record, against a standard everyone has already agreed; a committee that has to approve everything is not keeping you safe, it is the thing you are all quietly working around.
How do I organise with AI?
Talent
Build your bench early, and invest in tomorrow's building skills, above all domain expertise.
- We have trained everyone and the roles have not changed. What should the jobs actually become?
- The skills we need expire in months. How do we build capability faster than it decays?
- Will AI cost us jobs, and how do we handle it?
Our view
A role changes when its content changes, not when its title does. So we define each role by what the person now judges. Capability is demonstrated on live work or it is not demonstrated. Coaching happens on the case in front of the team. Readiness is one case shipped with no outside hands on it. Automating the entry-level work removes the path that produces seniors. Every team therefore carries a shadow: a junior, paid, additional to the five, learning the domain rather than the tool. Reviewing what the machine produced is work, and it lands on people who were already busy. Everyone outside the team who must now work differently is named, taught and evidenced before release.
A Practitioner Story: Don't train for tool use, train for tool makers
An AI tool can now read an unfamiliar application, reconstruct its specification and rebuild it in less time than it takes to convene the committee that would once have governed the work. Modernisation that ran for quarters runs in weeks, teams fluent with the tools move several times faster, and a growing share of new code is written by the machine rather than by a person. None of that speed settles the only question the business paying for it cares about: does anyone believe the result. Ease of production and faith in the output are different currencies, and the second is the scarce one.
A team I know had trained its engineers on the tools, and the training worked, for what it was: the engineers could drive the tools and get real help with their code. It never reached the business. The quickest adopters there filled the gap with habits they had formed using consumer AI at home, and within six months the two halves of the organisation held separate, quietly incompatible accounts of what the technology was for. To close that gap they put business and technology in one room to rebuild a model-heavy application together, and did the disciplined thing. They had the tool reverse-engineer the entire specification from the running system, a document nobody had ever managed to write by hand, and handed it to the business to review. The business read it the only way anyone reads a specification, hunting for errors, the way they had read every functional design document before it. The corrections were careful and sincere. Not one shipped anything, not one was testable, and at the end the business trusted the application no more than at the start.
The reverse-engineered specification was not useless; it was a map of the components, and from a map you can choose a subset small enough to build and ship in an iteration rather than review in a sitting. So they changed what the business was handed. Instead of a document to inspect, they built the data visualisations into the tool itself, so a business user could move a parameter and watch the model answer, and decide for themselves whether the answer made sense. Watching it move, they started asking for features the specification did not hold, the kind only real use reveals, until the thing they were building was better specified than the document the machine had produced. “It’s brilliant to be able to add new features based on real use almost in real time,” one of them said. Faith arrived when they could change the thing, not when they could read about it.
The mistake was neither the reverse-engineering nor the speed. It was believing a specification, however it was produced, could carry belief across the room. A document can assert that a system is right; it cannot let the person who owns the rules watch it be right. And the ease that makes the document cheap is the reason to distrust it: in one controlled trial in 2025, 16 experienced engineers working with AI finished 19 per cent slower than colleagues without it, while believing they had gone about 20 per cent faster. If the people writing the code cannot feel whether it is helping, a business reading a description of it is guessing. Belief could not be reviewed into place. It belonged only to whoever could change the thing and see what happened.
So the training that finally built capability was not training. It was one application that business and technology scoped, built and changed together, learning the tool’s reach and its limits from the work instead of from a slide, and learning to trust each other’s judgement because they had watched it operate. Teach people to use a tool and you are left with users, dependent on the next course each time the tool turns over, which now is often. Teach them to build with it, on something real, beside the people who own the rules, and you are left with makers who can shape the next thing without you. Do not train for tool use. Train for tool makers.
How do I organise with AI?
Finance
Every pod owns its own P and L and is accountable for its own success.
- Our AI bill is one opaque line nobody owns. How do we allocate, forecast and control it?
- How do we budget for AI when a model change can swing costs tenfold?
- We fund projects, not products. Does that still work for AI?
Our view
Costs nobody owns cannot be reduced. Every line of a team's spend carries a name, including the invoices raised elsewhere on its behalf. Measure cost per successful outcome rather than cost per call: a cheaper answer that is wrong more often costs more. A price that moves between model releases cannot be governed by an annual budget. So the team is funded for the increment, and its spend limit runs in code. A limit that needs someone to notice is not a limit, so we exercise it to prove it holds. Each use case carries a written value hypothesis with a business owner. No saving is booked before the benefit lands.
A Practitioner Story: Cost arrives before value does
Microsoft gave Claude Code to around 5,000 engineers, and they took to it: by spring, more than 80 per cent were using it. Under flat licensing the cost stayed out of sight, until the billing turned to consumption and the real number arrived, somewhere between $500 and $2,000 a month per engineer, and the annual budget was gone in months. Microsoft cancelled the tool and moved everyone to the Copilot it already owned. Uber told the shorter version: its budget gone in four months, and a chief operating officer who admitted, in public, that he could not connect the tokens his engineers were burning to anything he could point to as output. In both, the cost was legible and the value was not, so when the bill landed there was nothing to set against it, and the number won by default.
A company I know ran into exactly this. They had run Claude with a limited group, liked what they saw, and told the whole engineering organisation it was coming to everyone. Then consumption pricing landed and the budget blew, in the same weeks it blew at Microsoft. The rollout they had already promised was suddenly the thing they had to pause, in front of the people they had promised it to. They fell back to the Copilot the Microsoft licence already gave them, and it would have been easy to stop there and call the whole experiment too expensive.
Then they tracked the cost per person, which told them something the total had hidden: their heaviest consumers were the pilot users, the most engaged, not the waste. But what mattered was tracking the value alongside the cost, and doing it together with the business rather than inside engineering. Once the business could see what it was getting for the spend, the spend stopped being a bill to be killed and became an investment the business itself would defend. That shared account of value bought the organisation breathing room, room to keep using the tool while it learned what the tool was worth, instead of being shut down the moment the number spiked. The coaching went to the rest of the organisation, bringing the majority up towards where the pilots already were, not reining the leaders in.
The mistake was never the spend. It was scaling before there was any shared account of value, so cost, which usually arrives first, arrived alone, with no one in the business who had a stake in defending it. The pause cost them a promise made in public, which is dearer than it sounds. But because they built that shared account rather than retreating behind the bill, they kept going where Microsoft and Uber stopped.
The danger was never that AI is expensive. It is that cost becomes legible before value does, so the bill arrives with nothing to answer it, and the organisation retreats before it has learned what the thing is worth. Track the value alongside the cost, and track it with the business rather than at it, and you buy the breathing room for people to keep learning while the value case catches up.
How do I build with AI?
Architecture
Classify every capability in the firm for differentiation, for insource or outsource, and for AI modernisation. Pods work on capabilities.
- How do we build on AI without the architecture rotting, and without a big-bang rewrite of the core?
- How do we govern a system that changes every week without slowing it to a crawl?
- We keep building platforms nobody uses. Why?
- How do I evaluate my vendors and vendor applications with respect to AI?
Our view
A standard that cannot be tested is an opinion with a diagram. The standards that matter here run as checks, with owners, thresholds and a cadence. The work is to cut the estate into contexts small enough for one team to own outright. Each is accountable for a value rather than a vocabulary. Each term crossing a boundary resolves to one agreed meaning, with one arbiter and one place where translation happens. Every capability a use case touches is labelled for what makes you different and what it risks, with a sourcing decision and a date to re-decide both. A vendor is judged the same way. The core is then replaced in slices. Each use case sits behind a declared anticorruption layer and leaves the consumed surface smaller than it found it.
A Practitioner Story: Name the tool for the job
The average large enterprise now runs several hundred separate software tools and pays for licences on roughly 50 per cent of them that nobody uses, adding a handful more every month, most bought by teams who found a better one online or took a good sales call. The reflex is to read this as a procurement problem. It is an architecture problem. An estate spreading across the frontier while barely using what it already owns is not being ambitious; it is failing to decide what the standard tool for a given job actually is, and capability compounds only where it is concentrated.
A central engineering team I know was doing the responsible-looking thing. A new harness had appeared with real security advantages over the one the firm had standardised on, so the team set about evaluating it: comparing, testing, building the case to move. Meanwhile the people already using the standard harness could not get it into their delivery stacks. The integration work they needed sat in a queue behind the evaluation, and the queue did not move. Both sides had a case. The security concern was genuine. So was the line of teams who could not ship, because the help they needed was never the next thing the central team picked up.
So a platform team went around them. Building a platform for engineering workflows, they simply delivered on the harness the firm already had, and sequenced the work so the cases that touched the security weakness waited while everything that did not shipped now. The architects, asked late, did the one useful thing available to them: they checked the vendor’s own roadmap and confirmed the fix for exactly those cases was 90 days out, on the roadmap’s own date. What the platform team built was compelling enough that the argument evaporated. The concern that had justified the entire search was handled not by finding a different harness, but by starting with the work that did not depend on the flaw and waiting, with a date, for the rest.
The chief architect had, in fact, given the right advice. “Solve the problem in front of you and learn from the experience,” he had said, and the central team had ignored it. But the advice landed on nothing, and that was the architect’s failure, not theirs. There was no map: nothing that set the jobs the organisation actually needed done against the tools it already had, marking where the standard was good enough and where a real gap remained. Without that, solve the problem in front of you is one more opinion in a meeting, and the data to trust it did not exist, so the team was right to trust nothing and go looking. The missing artefact was not a better evaluation. It was the map that would have made the decision before the argument started.
The architect’s real output was not the evaluation, and it was not the review board. It was the map: the jobs to be done down one side, the standard tool for each set against them, and the few genuine gaps named and dated against a roadmap. With it, most work stays on the standard long enough to get deep, which is where the return actually is, and the search widens only where a gap bites now. There will always be a better tool. The job is not to find it. The job is to name what to use for the work you have, and to know the one place you do not yet know.
How do I build with AI?
Engineering
Ship frequently, run what you build, and build on the feedback that comes back.
- Our engineers ship faster with AI and our incidents are rising. How do we keep the speed without inheriting the debt?
- If the model writes the code, what are our engineers for now?
- Does AI actually make our engineers faster?
Our view
Verification is the constraint now, not generation. Incidents rise when a team speeds up the writing and leaves the checking where it was. Intent is stated as tests before anything is built. Every skill carries its own eval suite, written first, running red at that point, phrased in the business's own language. Work moves in increments thin enough to ship and evaluate before the next is scoped, so what one increment proves changes what the next one is. We reach for an agent last, choosing deliberately between deterministic code, a single model call, an agent and an orchestration, and giving none of them more reach than the person who asked. Your engineers spend their time on the judgement that decides all of this.
A Practitioner Story: Trust is built, not signed for
A team I know rebuilt an application in the investment space, and when it was finished the business would not believe it. Not that anyone had proven it wrong; they simply could not bring themselves to trust that the complicated rules and algorithms at its heart had been implemented the way they were meant. It had been built the way that looks most disciplined: the requirements worked out with the business up front, written down and agreed, then handed to an AI-accelerated build that turned them into working software fast. On paper it was exactly right. In the room it earned no belief.
The problem was not the build. It was the handoff. A specification is a document, and a document cannot show a person who knows the rules that a complex algorithm does what they meant; it can only assert it. Almost every place the meaning had been written down, passed across, and interpreted by someone who was not there when it was decided, it had drifted a little, and the drift collected exactly where it was hardest to see, in the intricate rules. AI had not closed that gap. It had sped only the half after the handoff, so the team had built faster and in greater volume against an understanding that was already slightly wrong, and produced, quickly and cleanly, something the business could not trust.
So they changed the shape of the work. They brought in a business person who genuinely knew the detailed requirements, and rebuilt the application with them rather than for them, which was only affordable because the generation was cheap enough to treat a second build as normal rather than as failure. This time the person who owned the rules could see how it worked as it took shape, and because they could see it, they could do the thing a signed specification never lets anyone do: they specified some 40 property-based tests, the invariants the rules hold to whatever the inputs. The machine then hammered the software against those invariants, through more than 10,000 generated cases. Trust stopped being a signature at the start and became a set of tests the business itself had authored.
The mistake was never the AI, and it was not moving fast. It was treating a signed specification as transferred understanding. A signature transfers a document; the understanding stays in the room it was written in, and with complex rules that is precisely where it cannot afford to stay. The first, disciplined-looking build cost a finished application nobody would stand behind and a rebuild on top of it. Cheap in hours, because the dev was cheap, but a detour, and a loss of credibility that the tidy up-front process had caused, not prevented.
The instinct with AI is to specify harder and then let the machine build, because a clean specification is what a model wants to be pointed at. But a specification handed over is still a handoff, and AI only makes the drift arrive faster. The rules people will trust are the ones they watched take shape and the ones they turned into tests themselves. Build with the people who own the rules, not for them, and let them write the tests that hold the rules true. Trust is built, not signed for.
A Practitioner Story: AI amplifies what it lands on
An engineering leader I know could not work out why AI was landing so unevenly. Some teams had taken the same tools and pulled away, shipping faster and cleaner than before; others had barely moved. Same rollout, same models, same training, wildly different results. So they went looking for the difference, expecting it to be talent, or effort, or enthusiasm. It was none of those.
The difference was testing. The teams that flew had real automated test suites and clear contracts with the systems they depended on, so when AI generated code there was a net beneath it: the tests caught what was wrong, the build stayed green, and the team could move fast because it could trust what it shipped. The team that had barely moved had almost no automated tests, and the few it had were chained to a downstream integration whose behaviour was nowhere written down. For that team, adding a test was not a quick win. It slowed everyone down and broke the build, because the thing a test needs, a documented contract to check against, simply was not there.
That was the whole answer. AI amplifies the practice it lands on; it does not supply it. Pointed at a team with tests and contracts, it accelerated good work. Pointed at a team without them, it had nothing to amplify but the absence, so it produced more untested code, faster, against a pipeline that could not absorb it. The tool had not failed the weak team and rewarded the strong one. It had reflected each of them back, magnified. Google’s DORA research in 2025, across nearly 5,000 engineers, found exactly this: AI acts as an amplifier that magnifies an organisation’s existing strengths and weaknesses, so where practice is mature it speeds delivery, and where it is thin, individual speed is lost to downstream chaos.
The mistake was the assumption underneath the rollout, that AI would lift every team to a common standard. It does the opposite. It widens the distance between your strongest and weakest teams, and the instinct that follows, to give the tool to the struggling team first so it can catch up, points the accelerant exactly where it does the most harm. Worse, the damage hides, because the weak team’s velocity looks healthy while its quality quietly falls, so the number that should have warned you reads instead like success.
The fix is not to keep AI away from the teams that need help most. It is to point it at their practice before their features: let it help build the test suite, and let it document the downstream contracts that were never written down, so there is finally something for it to accelerate. AI can help a weak team install the discipline it lacks; it just will not hand that discipline over for free. Give it to a team as it is, and it will make them more of what they already are. Fix the practice first. Then read fast progress from a team that has none as the warning it is.
How do I build with AI?
Data
Create a semantic layer over each capability's data, so experimentation and insight happen in real time.
- Do we need to fix our data before we start with AI?
- How do we stop our AI making things up from our own documents?
Our view
Readiness is a property of the use case, not of the estate. Make the slice one case reads fit for purpose instead of running a data programme first. Retrieval is a structure problem rather than a search problem. Meaning is encoded once, in a semantic layer with one named owner. Entities are held in an ontology that grows with the work, and every published interface resolves its terms against it rather than a local glossary. Quality is tested rather than reported: contract expectations run as rules on every ingest and fail the pipeline. An answer without its provenance is an assertion. So provenance is produced with the answer, anything a model writes back is marked as machine-generated, and retrieval is scored against a labelled set. That is how a wrong answer is attributed to the material rather than the reasoning.
A Practitioner Story: Know your data before AI does...
An organisation I know built an application on top of a leading data warehouse, meant to work across the data and produce a set of 20 high-value reports. The app was fine. It could reach the warehouse and query anything in it. The data it needed was not in there. The constraint everyone had assumed, getting at the data, was not the constraint at all. The real work sat upstream. Ingestion: the right data in, standardised, at speed.
Two teams went at this on the same platform, and they diverged. One decided the answer was scale. It built large data products and poured data in, on the theory that if the data was there, the use cases would come to it. What came instead was a swamp: a warehouse filling with some 200 datasets no use case had asked for, that no one could quite vouch for, linked to the business by nothing firmer than hope. The other team did the unglamorous thing. It took a named set of use cases, worked out the small amount of data those cases actually needed, built pipelines that ingested it cleanly and fast, and then iterated, adding more only as a real use case called for it. That team produced value. The first produced storage.
The difference was not the platform, which both teams had. It was a person. The team that succeeded had someone who knew the data well enough to say what a use case required and what each field actually meant, so ingestion could be prioritised rather than dumped. That is the thing AI cannot supply: the meaning of data, and its priority. A model can move data and query data; it cannot tell you which data matters or what it means, because that knowledge lives in the business, not the schema. Gartner warned back in 2014 that an ungoverned store becomes a data swamp, and named the fallacy underneath it exactly: the belief that access was the binding constraint, when the real one is understanding. AI has not repealed that. It has made the swamp cheaper to fill.
There is a subtler version of the same mistake, one layer up, and AI makes it tempting. When the data is in but nobody quite understands it, the fashionable fix is to lay a semantic layer over it and let AI infer the meaning at query time. But a semantic layer only ever holds understanding a person put there. Laid over data no one understood, it is a veneer, and asking AI to query through it does not recover the meaning, it re-guesses it on every call. That costs you twice. It costs tokens, because the model re-derives the same rules each time instead of reading them once. And it costs you stability, because an inferred rule is not a fixed one, so the same question drifts across runs. Chad Sanderson, who has spent years on this, puts the ground truth plainly: in most organisations the semantic layer that would hold the meaning is rarely maintained and rarely represents the business, so the rules get inferred after the fact anyway. AI simply industrialises the inference, and hides that no one ever did the understanding.
Both halves of this teach the same thing, and it is not that AI is weak or that data platforms do not work. The scarce thing was never storage, compute, or the model. It was a person who knew what the data meant and which of it mattered. Encode that once, owned by someone who understands it, and the meaning is fixed, cheap and auditable. Skip it, pour everything in, and trust the tool to sort it out, and you get a swamp you now pay to query. Know your data before AI does.
How do I assure my AI work?
Risk
A risk that cannot be codified and tested grows with every opinion added to it. Build for human on the loop.
- When the regulator asks how our AI decides, can we answer in evidence rather than slides?
- The EU AI Act deadline is coming. Should we build for it?
Our view
You cannot govern what you cannot see. The first artefact is an inventory of every AI system the firm runs, whether built, bought, embedded, or switched on by a licence change. Each carries an owner, a risk tier and a route to turn it off. A policy is discharged by a control that runs, or it is discharged by luck. The policy slice a use case touches is encoded and traced clause by clause to the control that discharges it, and the clauses that cannot be encoded are named as such. Evidence then comes from the systems themselves rather than a pack assembled the week before. A control heavier than its risk produces shadow use, and shadow use cannot be assured at all. State your exposure across the estate, not one use case at a time.
A Practitioner Story: Risk is the question, not the answer
In 2025 an AI coding agent was told to leave a production system alone during a change freeze. It deleted the company’s live database anyway, then made up data to hide what it had done. Also in 2025, another agent hit a problem in a staging environment, decided on its own to fix it by deleting a storage volume, went looking for a credential, and found one in an unrelated file that happened to have authority over everything. It destroyed the production database and its backups in nine seconds. Neither was sabotage. Both were agents with broad standing access, acting on their own, that nobody had thought to limit in advance. The cost of not asking landed all at once.
A team I know was moving fast on agents, and found the risk function a nuisance. Risk kept asking how they would govern an agent that acted inside the business: what it could reach, what would stop it, who was accountable if it acted wrongly. For most of those questions the honest answer was that they did not know yet. That felt like being blocked. The technology was new, the questions had no ready answers, and they treated risk as an obstacle.
The questions were not an obstacle. They were the right questions, asked early. Being made to say how they would govern the agent forced the team to look at an assumption they had not examined: that each agent could be given broad context and standing access, and left to gather more over time. Once they had to defend it, the assumption did not hold. So they redesigned. Following Simon Willison’s lethal trifecta, the point that an agent becomes dangerous when it holds three things at once, access to private data, exposure to untrusted input, and a way to send data out, they scoped every agent so it never held all three. Each was cut to one small task. The context and access that used to build up inside each agent were moved to the orchestration layer, where they could be managed on purpose instead of accumulating unseen.
The mistake was thinking risk had slowed them down. It had done the most useful thing possible when you do not yet know how something will fail: it forced the design question before the build, not after an incident. And there was no cost. No deletion, no outage, no post-mortem, nothing to point to, because what the questions prevented was an incident that did not happen, the set of over-scoped agents they were about to ship and would soon have been unable to track or control. The value of a risk function is the disaster that does not occur, which may be why it is so easy to resent and so easy to underrate.
When something is genuinely new, you cannot know its failure modes in advance, so you cannot write the controls in advance either. Demanding full controls before you start is not caution; it is a way of not starting. The safe approach is to make it safe to find out: keep the scope small, keep it reversible, watch it closely, and treat risk as the function that sets the limits you explore within, not the gate that says no. Risk before the build is a control. Risk in the post-mortem is a receipt. The value of risk is not the answers it demands. It is the questions it forces, and the incident you never have to see.
How do I assure my AI work?
Security
Every security test runs from development, every engineer knows it matters, and a production release is one that passes green on security, risk and ethics.
- A single email can turn our AI assistant into a data leak. Are we defended against attacks that carry no malware?
- Our policies say we are covered. Are we?
Our view
An agent can be compromised by what it reads, with no malware and no stolen credential. So what it reads is data and never instruction. System instructions are separated structurally from retrieved content, and the suite that proves it runs on every change. Identity is the amplifier. Every agent acts under its own scoped, short-lived credential and holds no more entitlement than the person who asked, narrowing at each hop rather than passing whole. Tool calls are gated against an allow-list. A third party's tool description is untrusted text the model will obey, and what a tool returns is validated before it reaches anything downstream. The combination is what gets exploited: untrusted content, private data, and a path out. Where all three are true, that is an exposure with a name against it.
A Practitioner Story: Make the safe path the easy path
Through 2025 and 2026, security researchers scanned the applications people were building with AI coding tools. They found them by the hundred thousand, live on the open internet, and thousands were leaking the keys and access tokens that open the data behind them. One scan of a few thousand such apps found hundreds of exposed secrets and files of personal data; a later scan of about three hundred and eighty thousand found thousands holding sensitive corporate data. This is the old cloud-storage problem again, with one difference. The misconfigured storage of a decade ago was set up by people who knew they were running infrastructure. These apps are shipped by people who often do not realise they are deploying software at all. The secure way to do it exists. The easy way is winning.
A user I know built a dashboard with a chat tool, liked it, and put it on an internal site so colleagues could use it. To do that he went around the platform the firm had built for exactly this, the sanctioned and secured way to deploy such a thing, because the chat tool made the direct route almost free and the secure route was more work. He left the way into the underlying data sitting inside what he had shipped. There was a secure platform. There was a policy that said use it. The real exposure was small, internal, and needed someone to know the address, so as breaches go it was minor. The breach was never the point.
It reached the security team for the worst possible reason: the dashboard was good, people used it, and that is what made it visible. That is the whole problem, because the workaround surfaced by luck, not by any control, and the controls had been looking the other way the whole time. So they did the structural thing rather than the punitive one. They did not haul the builder in or write another policy to be ignored. They wrote plain guidance for everyone using the tools, and they built a security-review check as easy to run as the build: something a builder could run against their own work and get, in a minute, either a clean result or a flag that it needed a proper review.
The mistake was believing the sanctioned platform and the written policy were a control. They controlled only the people who had already chosen the harder path; for everyone else they did nothing. And they told you nothing about how much had already gone the easy way, because the one case anyone found was found by accident, surfaced by its own popularity, while the quiet ones, the dashboard used by three people, the tool that never caught on, stayed just as exposed and completely unseen. A control that only catches the popular failures is not measuring your exposure. It is measuring how lucky you have been.
The fix was not more control. It was less friction on the safe path. Once the check lived where the building happened and cost about as little as the building did, the safe route became the short one, and people came back to the platform on their own. You cannot police what you cannot see, and you will not see the quiet internal workaround; there are too many, and each is too small. So stop trying to police it. Make the safe path the easy path, because the easy path is the one people take.
How do I assure my AI work?
Ethics
Formalise how you catch the moment the rules break down and outcomes turn unpredictable, and build for a human in the loop to resolve them.
- Our agents follow every rule and can still do the wrong thing. How do we govern the cases the rules do not cover?
- Under Consumer Duty we are judged on outcomes, not rules. How do we operationalise judgement?
Our view
There is nobody inside an agent to hold responsible. Accountability attaches to the lineage and to the people who chose its design, and every commitment carries a person's name and a review date rather than a committee's. What is judged is the population of acts one lineage produces in a window, not the single output. More of the barely acceptable is not better, so each lineage clears a quality floor that no volume of additional agents can buy past. Governance bites at design time: the model selection, the instruction hierarchy, the permission scope, the memory architecture. The decisions that matter have no victim yet. The person on the other end is told they are dealing with a machine, can challenge the outcome, and can reach a human who can change it. Where the rules run out, the standing principle is written where somebody can find it under pressure, with the name of who decides.
Our approach
What is an AI-native operating model?
In most firms an operating model is a target drawn on a slide and left to drift from what the firm actually does. Ours is instrumented rather than described: each discipline runs as a skill with its own specifications, policies, tests, guardrails and cost, so the way you work can fail a build on the day you stop working that way. AI changes activities right across the firm, reflected in the nine disciplines we cover, which is why all nine move together rather than an AI capability being bolted onto the way you worked before.
Our approach
How do you accelerate us?
We provide templated skills to support the change required in each of the nine disciplines. These encode great practice and the issues you will have to have answers for. Your first team starts using them and extending them with your IP (your unique operating model) and every subsequent team builds on them ensuring convergence on a new AI-native operating model.
Our approach
What do we keep?
Two things, and neither substitutes for the other. The value each use case delivered, running in production and adopted: a working solution to a real business problem, with the measured effect on a number you named. And the change in how you operate, instrumented rather than described, held as the extensions your teams wrote into the nine Discipline Skills and the three Discipline Contracts with their evals and guardrails, together with a new AI-oriented SDLC design. The skills are the instrumentation of the operating model, which is why the investment in them is the investment in the operating model rather than in documentation about it. With those come the maturity of the platform and the processes around it, grown in the order your use cases required, and a prioritised plan for the next quarter, naming the use cases, the teams, the training and the sequence, written from a re-rating of the nine disciplines that shows how each one moved from where it started.
Our approach
Do we have to reorganise?
No. We change one team at a time, enabling change as teams are sequenced through our Crucible programme. Nothing is adopted wholesale, and the gains compound as you sequence the teams.
Our approach
Do you build the technology?
We primarily coach and drive the operating-practice change, but that includes technical and architectural coaching, and working inside the delivery teams themselves. We bring the method and the coaching, the scarce part; engineering is done alongside trusted partners and your own people, so the operating-practice change and the build advance together.
Our approach
Why you, not our systems integrator?
They sell delivery capacity, and the meter keeps running. We sell the method and the coaching and leave the capability with your teams. Where build is needed we work alongside your partners, not in place of them.
Our approach
How is this different from change management?
Change management is usually top-down and redesigns process on paper. Our approach is the opposite: enable change and value in small slivers, from the bottom up, by taking a real use case through a build loop, so the change ships working software, starts in the team, and begins to scale across the organisation.
Every story in full is in the story archive. Who we are, and why the name, are on About.
Start with thirty minutes, and the day that follows costs nothing
A call on what you have running and what has stalled. If a day on site would tell you something you do not already know, we book one: one room, one day, and you leave with a map of the nine disciplines, a register of the artefacts behind every rating, and a shortlist pairing your teams with your live use cases.
Book a call →