Why the Pilot Always Works, and the Rollout Never Does

Why the Pilot Always Works, and the Rollout Never Does
A pilot doesn’t test a method. It tests an architecture, temporarily installed, and architectures only hold under two conditions.
Every organisation has run this sequence.
Someone proposes a new way of working. Rather than impose it everywhere, sensible people suggest testing it first: one team, one site, one department. The pilot runs. It works, sometimes remarkably well. The team is enthusiastic, the numbers move, and there is a presentation with a slide showing the before and after.
So the decision is taken to roll it out.
And it underperforms. Not catastrophically, usually. It simply fails to reproduce what the pilot produced, and after a year or so it becomes another thing the organization does, with no one able to say whether it works.
The explanations offered afterwards are familiar. The wider organisation was less committed. Change fatigue. Communication wasn’t good enough. Middle managers didn’t get behind it.
All of those may be true in particular cases, but none explains why the pattern repeats so reliably. Something that happens almost every time is not a series of unfortunate executions.
A second, less discussed thing also happens, and any honest account has to explain it. The pilot team itself usually reverts. Six months after the sponsor’s attention has moved on, the site that made it work is doing roughly what it did before. If the problem were that the rest of the organisation failed to adopt the change, the original team should still be a demonstration of it. Often they aren’t.
Both failures have the same cause, and it isn’t really about pilots.
What a pilot actually changes
The presentation says the pilot tested a method. It didn’t.
Alongside the method, a pilot almost always alters the conditions the team works under. Approvals that ordinarily take three weeks take a day, because someone senior is interested. Reporting requirements are relaxed, because it’s a pilot. Resources that would normally be contested are simply provided. And, most importantly, the ordinary consequences of failure are suspended. An experiment that doesn’t work is an interesting result rather than a black mark.
Every one of those is an architectural change. Decision rights move. Information flows differently. Consequence is reallocated.
So the pilot tested a bundle: the method, plus a temporary architecture installed around it. The team then behaves differently, and the difference is usually attributed to the method, sometimes to a change in culture, occasionally to the quality of the people involved.
Then the rollout transfers the method and leaves the architecture behind. The method was written down and the architecture wasn’t. Nobody recorded the cleared approvals or the suspended consequences because, at the time, they looked like support rather than the intervention itself.
The rollout is therefore not a scaled-up version of the pilot. It is a different intervention, and the pilot produced no evidence about it.
Why the architecture didn’t hold even where it worked
That explains the rollout. It doesn’t explain the reversion, and the reversion is the more useful finding.
Architecture is not free-standing. It is culture made operational: the organisation’s real answer to what pays off here, expressed in who decides, who carries the consequence, and what happens to the person who raises a problem. An architecture that expresses something the culture doesn’t hold is an imposition, and it survives only as long as something holds it in place.
This highlights two situations under which an architectural change will last. A change that meets neither will not.
The first: the architecture never expressed the culture correctly.
Structures arrive from many places: a predecessor organisation, a merger, a regulator’s requirement, a template someone brought with them. A control gets introduced because one executive wanted visibility on one occasion, and stays because removing it is nobody’s job. None of that was ever the organisation’s real answer to what matters here. It was simply what got built.
In that case, an architectural change corrects a mismatch that was there from the beginning. It holds without anyone’s beliefs having to change, because it brings the structure into line with what the organisation already valued.
The second: the culture has changed, and the architecture hasn’t caught up.
This is commoner than it sounds, because culture shifts without anyone deciding it should. What pays off in an organisation moves whenever experience moves: a chief executive who reliably backs people who escalate problems, a downturn that makes cost the only thing anyone is praised for, three rounds of redundancy after which everyone has drawn their own conclusions about who gets kept. None of it is announced. All of it changes the answer.
Meanwhile the architecture stays where it was, expressing an answer the organisation has stopped giving. An architectural change here is the structure catching up with something that has already happened, and it holds for the same reason: it fits.
Meet neither condition, and the change is an architecture imposed on a culture that doesn’t support it. It will work while sponsorship holds it up, and stop when sponsorship moves.
Which is precisely what a pilot looks like.
The people who volunteered were not a sample
There is a second reason pilots mislead, and it concerns the consequence half of the architecture.
Pilots run with volunteers, or with a team chosen for enthusiasm. Sensible, and it is where the reasoning quietly goes wrong.
Consider who volunteers for something unproven and visible. Disproportionately, people who can afford to: someone with standing, an established record, a secure position, or a manager who will protect them. For that person a failed experiment is an interesting attempt.
Now consider the person who didn’t volunteer. Newer, less established, working in a function where being wrong is expensive, or reporting to someone who wouldn’t have shielded them. For them, the same experiment carries real risk, and declining it was an accurate reading of their own exposure rather than a lack of enthusiasm.
So willingness and exposure are correlated, and the pilot was run by the people for whom the change was cheapest. It generates no information about what the change costs anyone else, and when the rollout reaches them, they meet the same change at a different price and respond accordingly.
That gets recorded as resistance. It is the first honest measurement of the cost, arriving too late to inform the decision.
The test nobody runs
Sponsorship doesn’t distinguish between the cases. A pilot can be sponsor-backed and correcting a genuine mismatch, in which case removing the sponsor changes nothing, because the architecture now fits. Or sponsor-backed and imposing something the culture doesn’t hold, in which case it decays as soon as attention moves.
So the presence of sponsorship tells you nothing. Its removal tells you everything.
Which suggests one test, and it is not complicated. Before scaling anything, withdraw the sponsorship from the pilot team and leave the new arrangement running. If it holds for three or four months with nobody senior watching, one of the two conditions was met, and it will transfer. If it decays, the pilot demonstrated what sponsorship can hold in place, which is not the same thing and will not survive contact with the rest of the organisation.
Almost nobody does this, and for the same reason as everything else here. The moment a pilot succeeds is the moment to announce it. Deliberately leaving it alone for a quarter to find out whether it survives, is a delay that produces no result anyone is rewarded for. And it risks discovering that what you have just presented to the board doesn’t hold.
And two questions worth asking before the pilot starts
Did this structure ever fit? If the approvals, controls, and reporting lines being changed were inherited, imposed, or built for a reason that has passed, the change is correcting a mismatch, and there is good reason to expect it to hold.
Has what people believe pays off here moved? If something substantial has happened – a change at the top, a restructuring, a period that taught everyone a lesson nobody wrote down – the architecture may be expressing an answer the organisation has already abandoned.
If neither question has a clear answer, the pilot is proposing to install an architecture that nothing underneath it supports, and the result will need sponsorship indefinitely.
Two things worth recording while it runs
What had to be cleared. Every approval waived, every rule relaxed, every resource found because someone asked. These are the most valuable data a pilot produces, and they are almost never written down. A list of what had to be unblocked is a list of what the rollout will hit.
What consequence was suspended. What would ordinarily have happened to someone who tried this and failed, and what happened instead. If failure was costless during the pilot and won’t be afterwards, the rollout is asking people to do something materially different from what was tested.
The finding underneath
Pilots are not a mistake, and the alternative – imposing an untested change everywhere – is worse.
But most pilots are designed to produce a result rather than information, and the result they produce is reliably misread. A pilot that works under protection has shown that the method is sound and that the architecture around it was temporarily different. Those are two findings, and organisations consistently take delivery of the first while discarding the second.
Which leads somewhere more uncomfortable than a methodology problem. If a change only works where the ordinary structure has been suspended, the pilot isn’t telling you that the change is good and the rollout was mishandled.
It is telling you what the structure does to that change, which is more useful than the result on the slide, and considerably less welcome.