Somewhere in your company there is a slide deck from a demo that went well. Maybe it was an assistant that answered questions about your policy library. Maybe it drafted a first pass of a client report in eleven seconds. People in the room reacted the way people react when something works. Someone said we should roll this out.
That was a while ago. The pilot still exists, in the sense that a link to it still exists. Nobody outside the original group uses it. No customer has ever touched it. The project has not been cancelled, because cancelling it would require a decision, and no decision is pending. It is simply not going anywhere.
This is the most common outcome in enterprise AI, and it is not because the technology failed. The demo was real. What it demonstrated was narrower than anyone in the room understood at the time.
A demo and a system are different objects
A demo has one job: show that the interesting part is possible. It runs on inputs someone chose. It is operated by the person who built it, who knows which questions it answers well. When something goes wrong, that person says let me try that again, and everyone laughs, and the demo continues.
A system has a different job. It runs on inputs nobody chose, operated by people who were not in the room, at volumes nobody rehearsed, on a Tuesday when the person who built it is on leave. When something goes wrong there is no let me try that again. There is a customer, a record, and a question about who is accountable.
The distance between those two objects is where pilots die. It is mostly not model work. It is the work of turning a capability into something an organisation can rely on, and it tends to be invisible at demo time because a demo is designed to skip it.
The five gaps that stop a pilot
Nobody defined what good means
Ask the team what accuracy the pilot achieved and you usually get a version of it seemed pretty good. That is not a criticism of the team. Nobody gave them a target, because nobody knew how to set one.
Without an agreed standard there is no way to say the system is ready, so it never is. Every stakeholder gets to hold a private threshold, and at least one of them will always be unconvinced. Somebody senior eventually asks whether it can be trusted with real clients, the honest answer is we think so, and that is the end of it.
A system that ships has acceptance criteria written before launch: which tasks it must handle, at what quality, measured against a fixed set of real cases including the awkward ones. That set is not a formality. It is the thing that lets you say yes.
It never saw your actual documents
The pilot ran on a folder. Fifty clean files, recent, well named, no duplicates. Your real environment is twelve years of documents across four systems, three of which contain a file called final_v2, and some of which contradict each other because policy changed in 2023 and nobody retired the old version.
Retrieval quality is mostly document quality. When a pilot moves from the folder to the estate, answers get vaguer and occasionally wrong, and users conclude the thing is unreliable. They are right, but the fault is upstream. Nobody had decided which source wins when two sources disagree, or which documents are authoritative at all.
Permissions were left for later
In the demo everyone had access to everything, because the demo had one user. Then the first real question arrives: if an associate asks about a matter they are not staffed on, what happens?
If the answer is being worked out after the build, the pilot stops there, and it should. Retrofitting access control onto a system that assumed a single trusted user is not a configuration change. It reaches into how content is indexed, how requests are scoped, and what gets logged. In financial services, professional services and legal it is the first thing that is asked and the last thing that is planned for.
It sat outside the systems people actually use
The pilot has its own URL. Using it means leaving the thing you were doing, going somewhere else, retyping context that already exists in the file you had open, and coming back with a result you then paste in by hand.
People will do that during a pilot because they are being watched and because it is novel. They will not do it in week six. Adoption failures usually get diagnosed as change management, and change management gets diagnosed as a training problem, and none of that is the issue. The tool asked for effort it did not repay.
No one owned it in production
The pilot was built by an innovation group, a consultant, or one capable engineer with spare capacity. Production requires someone who is accountable when it is slow, when a model provider changes behaviour, when the document store drifts, when spend triples in a month because a workflow got popular.
When that owner does not exist, IT will not accept the handover, and they are right to refuse it. An unowned system in production is a liability with a login page. So the pilot stays a pilot, which is the only state in which nobody has to sign for it.
What shipping actually requires
None of these gaps is exotic and none of them is a research problem. They are the ordinary requirements of putting software into an organisation, and they get skipped because AI pilots are usually run as experiments rather than as builds.
The correction is to design for the second object from the start. Decide what good means before you build, in numbers, against real cases. Treat the document estate as part of the system and fix what is wrong with it. Model permissions on day one, not after the security review. Put the capability where the work already happens. Name the person who owns it when it breaks.
Done in that order, the demo stops being a separate event. It becomes an early view of something already headed for production, and the question of whether to roll it out never needs a special meeting.
Where we come in
Alppoint puts engineers on the problem and ships the system. Not a prototype for your team to productionise, and not extra hands on a plan somebody else drew. We take the whole function: the evaluation harness, the retrieval layer, the permission model, the integration into your environment, and the handover to people who can run it. It goes live in your environment and it stays your asset.
If you have a pilot that demoed well and stalled, the diagnosis usually takes one conversation. We are happy to have it.



