The Proof-of-Concept Graveyard

Perspective·Giovanni Leonardi·January 2023·9 min read

A proof of concept proves that something can work; it says almost nothing about whether the organisation can operate it, and operating it was always the entire problem.

The Graveyard Nobody Walks Through

Every organisation that has taken artificial intelligence seriously over the last few years is quietly maintaining a graveyard. It is not on any slide. It is the accumulated remains of proofs of concept — dozens of them in the larger firms, sometimes hundreds — each of which worked, each of which was demonstrated to an appreciative audience, and almost none of which is running in production today. The pattern is so consistent across sectors that it has stopped being an anomaly and become the norm. The POC succeeds; the system is never born.

What makes this striking is that the failure does not look like failure at the moment it occurs. Quite the opposite. The proof of concept is a triumph. The model predicts, the demo runs, the room nods, the sponsor is pleased. Everyone leaves the session believing that the hard part is behind them and that production is now a matter of scaling up what has been shown to work. Then the initiative enters the space between the demo and the operating floor, and it dies there, silently, alongside all the others. No one holds a post-mortem, because nothing appears to have gone wrong. It simply never quite made it.

I have watched this happen enough times, and across enough organisations, to be confident that it is not a run of bad luck or a shortage of talent. It is structural. And the structure that produces it begins with a misunderstanding of what a proof of concept is actually for.

The Proof Proves the Wrong Thing

A proof of concept answers a single question: can this work? Can a model, given clean and curated data, in a controlled setting, produce a useful result? That is a real question, and for genuinely novel techniques it is worth answering. But it is almost never the question that determines whether an initiative reaches production, and here is the trap: in the overwhelming majority of cases, we already knew the answer was yes before we started. The technique had worked elsewhere. The feasibility was never seriously in doubt. We spent the effort proving the thing that was least uncertain.

The question that actually governs whether an AI capability survives is a different one entirely: can this organisation operate it? Can it feed the model reliable data every day, not once, in the sanitised extract prepared lovingly for the demo? Can it deploy the thing into a real environment, monitor it, retrain it as the world drifts underneath it, and hold someone accountable when it misbehaves at two in the morning? Can the people whose work it changes actually absorb it into how they operate? These questions have almost nothing to do with the model and almost everything to do with the organisation, and the proof of concept is designed in a way that systematically avoids asking any of them.

“The POC answers the question we were least worried about and stays carefully silent on the questions that will actually kill the initiative.”

This is why the graveyard fills up with successful pilots. They succeeded at the wrong test. A demonstration that a model can produce a good answer under ideal conditions is not evidence that the organisation can run it under real ones. We keep treating the first as a down payment on the second, and it is not. It is a different thing altogether.

Production Is Not a Bigger Pilot

The deepest error underneath the graveyard is the belief that production is simply a scaled-up proof of concept — the same thing, but larger and more robust. It is not. Production is a different discipline, with different demands, most of which are entirely absent from the pilot and none of which scale down into it.

Consider what a pilot quietly does without to make itself easy. It uses a one-off data extract, hand-cleaned, rather than a live pipeline that must deliver quality data continuously and survive every upstream change to a source system. It runs in a notebook or a sandbox rather than an environment with the security, resilience, and integration that operational systems require. It has no monitoring, because it runs once for an audience rather than continuously against a shifting world. It has no retraining strategy, because it never has to cope with the model degrading as reality moves away from the data it learned on. And it has no answer to the human question — who now works differently, who owns the output, who is accountable — because for the length of a demonstration nobody has to.

  • Data as a living supply, not a one-off extract. The pilot borrows clean data once; production must be fed reliable data forever, through pipelines that withstand every change upstream.
  • An environment built to run, not to impress. Integration, security, resilience, and observability are the substance of production and entirely optional in a demo.
  • A response to drift. Models decay as the world moves; a capability with no plan to detect and correct that decay is not production-ready, however well it performed on the day.
  • A change to how people work. An output nobody has been prepared to trust, own, or act on changes nothing, no matter how good it is.

Each of these is real work, and it is the work that the proof of concept, by its very design, defers. So when an initiative crosses from pilot to production, it does not meet a slightly harder version of what it has already done. It meets an entirely new body of work that was never scoped, never funded, and never staffed — because everyone believed the demo meant it was nearly finished. That is the moment the initiative enters the graveyard.

The Funding Pathology

There is an organisational logic that keeps the graveyard filling, and it is worth naming plainly because it is rational at every individual step and disastrous in aggregate.

Proofs of concept are cheap, fast, and visible. They can be funded from a modest innovation budget, delivered in weeks, and shown to a leadership team hungry for evidence that the organisation is doing something about AI. They generate exactly the kind of artefact — a working demonstration — that is easy to celebrate and easy to point to. Production, by contrast, is expensive, slow, and invisible. Building the data pipelines, the deployment environment, the monitoring, the retraining discipline, the change management — this is a long, costly, unglamorous programme of work that produces no photogenic moment and touches parts of the organisation that resist being touched.

We fund the part of the journey that is cheap, fast, and flattering, and we starve the part that is expensive, slow, and decisive. Then we are puzzled that so many initiatives complete the first part and none complete the second.

So the incentives point relentlessly at the first mile and away from the last. Sponsors get their demonstration and their credit; the harder, later work is left unfunded because it is harder and later, and because by the time its absence becomes visible, attention has already moved to the next pilot. The organisation optimises for the production of proofs of concept and calls it an AI programme. What it has actually built is a machine for generating graveyard occupants.

Design for the Last Mile First

If the diagnosis is right — that the POC answers the wrong question and that production is a distinct discipline the pilot systematically avoids — then the corrective is not to run better pilots. It is to invert the order of the work.

The first question of any AI initiative should not be can this work? It should be if this worked, could we operate it — and what would we have to build to do so? That question, asked at the outset, changes everything downstream. It forces the data pipeline, the deployment path, the monitoring, and the change-management burden into view while there is still time to scope and fund them. It reframes the proof of concept from a demonstration of feasibility into a test of operability: not merely can the model produce a good answer, but can we feed it, deploy it, watch it, correct it, and get people to act on it. A pilot designed to answer that question looks very different, and it fails — usefully — for very different reasons.

  1. Start from the operating model, not the algorithm. Before proving a technique works, establish what operating it would require and whether the organisation is willing to build that. If the answer is no, the pilot is a waste whatever it demonstrates.
  2. Fund the last mile as the main event. Treat the data pipeline, the production environment, the monitoring, and the change work as the substance of the programme and the model as a component within it — not the reverse.
  3. Make the POC test operability, not feasibility. Design pilots to surface the operational and human obstacles early, so that the ones which will never reach production die cheaply and quickly, before the investment rather than after it.
  4. Count production systems, not proofs of concept. Measure the programme by capabilities that are actually running and being used, and treat a graveyard of successful pilots as the failure it is, rather than the portfolio of achievement it is so often presented as.

None of this is a rejection of experimentation. Some questions genuinely need a pilot to answer, and cheap, fast failure is a virtue when the uncertainty is real. The argument is narrower and, I think, harder to dismiss: we are running pilots to answer a question we had mostly already answered, and declining to fund the work that actually decides the outcome. The graveyard is not evidence that AI does not work in our organisations. It is evidence that we keep proving it can work and keep refusing to do what it takes to make it run.

The organisations that will break the pattern are the ones that stop being impressed by the demo. A working proof of concept is not a milestone on the way to production; it is the easy part, dressed up as an achievement. The hard part — the only part that was ever going to determine whether any of this mattered — begins exactly where the applause stops.


More from Programme