AI transformation is a leadership decision. Everything else is tactics and tools.
We start from the operating model; the technology arrives afterwards. Most of what goes wrong is decided before anyone writes code.
Why pilots die.
Four failure modes account for most of it. None of them is technical.
It was built beside the business.
The pilot runs on clean data, with a friendly user group, in a sandbox where nothing else depends on the outcome. Production has none of those conditions. A system that has never met the real queue, the real exceptions and the real edge cases has not been tested — it has been demonstrated.
The gap shows up as a set of conditions nobody wrote down: the record that arrives incomplete, the case that belongs to two departments, the applicant already in the system under a different spelling. None of it is exotic. It is the ordinary shape of the work, and a sandbox is built to exclude it.
Nobody owned it who could change anything.
The pilot sits with innovation, or digital, or a vendor. The process it would change sits with an operating department that was not in the room. When the pilot succeeds, there is no one with the authority to make the department adopt it.
Ownership here means something specific: authority over the process, over the people who run it, and over the target it is measured against. A sponsor who can fund the work but cannot change the standing procedure is not an owner. The test is whether one named person can switch the old process off.
There was no path to production before the build started.
Integration, identity, data agreements, procurement, security review and change management are discovered afterwards, one at a time, each taking months. By the time the path is clear, the sponsor has moved and the budget cycle has closed.
Each of those has a queue of its own, and the queues do not run in parallel unless somebody makes them. Sequenced after the build, they add a year to work that took eight weeks. Sequenced before it, most of them resolve while the build is still going on.
The measure was not one the executive committee recognised.
Accuracy improved. Satisfaction scores moved. None of it appears in cost per case, cycle time or backlog, so nobody at the top can tell whether it worked — and what cannot be evaluated does not get scaled.
The remedy is agreeing the baseline before the work begins, in the units the organisation already uses to run itself — and accepting the reading afterwards whether or not it flatters the project.
The four stages — each one ends with something you can act on.
1 Assess
We map where the work actually sits: volume, cost per case, cycle time, exception rate, and which decisions carry legal or safety consequence. We identify the two or three processes where agentic work would change the number, and the ones where it would not.
- You receive
- A written assessment your executive committee can act on, whether or not you continue with us.
2 Design
Target operating model first: who decides what, where the human sits, what the record must contain. Agent architecture second. A named owner inside your organisation is agreed before design begins — without one, we do not proceed.
3 Build
The production path is defined before the first line of code — integration, identity, data agreements, security review. We do not build anything that cannot graduate out of the sandbox.
4 Run and transfer
Monitoring, drift detection, escalation review, and a contracted handover to your own team. Our success condition is the day you run it without us.
What we measure.
Targets are agreed before the engagement begins and reported against afterwards. We propose these five as the standing set:
- Cost per case
- The number the executive committee recognises.
- Cycle time
- From arrival to resolved, not from assignment to closed.
- Escalation rate
- What share of work reaches a person, and whether that share is falling for the right reasons.
- First-pass accuracy
- How often the system is right without correction.
- Autonomous share
- How much volume completes without a human touching it.
If a proposed system cannot move at least one of these, we will say so during the assessment.
Human escalation by design.
An agentic system that never escalates is not efficient. It is unsupervised.
Every process we design names, in advance, the decisions that must reach a person: those with legal consequence, those affecting an individual’s rights or entitlements, those where the system’s own confidence falls below an agreed threshold, and those where the case does not resemble anything the system has handled before.
For each of those, three things are fixed before deployment: who holds the authority, what they see when the case arrives, and what record the decision leaves behind. The record is timestamped, attributable and linked to the case, so the decision can be reconstructed for an auditor, a regulator or a court.
A system that cannot reconstruct its own decisions will eventually be switched off. We build for the audit that comes two years later.
When we say no.
We turn down work. It is worth knowing when, before you brief us.
When automation would move risk onto people who cannot carry it.
If the design works by pushing consequences down to a frontline officer or a citizen who has no way to contest the outcome, the efficiency is not real. It has been relocated.
When the data is not ready and the project is really a data project.
Some processes do not need agents. They need a resolved source of record, and calling that an AI programme will waste a year and a sponsor.
When the process should be redesigned rather than automated.
Automating a bad process makes it faster and more permanent. Occasionally the honest answer is that four of the eleven steps should not exist, and no technology is required to remove them.
Saying this at the assessment stage costs us work. It costs you considerably less than finding out in month nine.
How we work with your team.
One named owner on each side. Our people work inside your operating department, not alongside it. Knowledge transfer is a contracted deliverable with defined artefacts — documentation, runbooks, and training against them — not a goodwill gesture at the end.
We are not trying to become permanent.
Two weeks. A fixed price. A document your executive committee can act on.
We look at where the work actually sits, what it costs per case, and which decisions carry real risk. You get an assessment you can use whether or not you work with us — measured in the five numbers above.
