AI deployment: where projects stall
An AI deployment runs through five stages: establishing what is actually wrong, getting at the data, building, getting the system used, and keeping it working. Two of those five are where projects die, and neither is the building. They die at data access, where a credential sits in a queue for eleven weeks while a timeline quietly becomes fiction, and they die at adoption, where a system that works correctly is delivered on time and nobody uses it. Outright technical failure is rare enough to be worth ignoring when you plan. This has a practical consequence for how you scope a project: the questions that predict a timeline are not about model capability or integration complexity. They are about who can grant access to which system, and who is accountable for the outcome the system is supposed to change.
Stage one: what is actually wrong
Every organisation arrives with a described problem, and the description has usually been simplified twice before it reaches anyone technical. This is not a criticism of anybody; it is what happens when a problem travels through a business case.
The work here is to look at how the task is done today rather than at how it is supposed to be done. The reliable technique is to ask to see the last five real instances rather than asking for a description. Descriptions contain the standard case. The five instances contain the exceptions, and the exceptions are usually where the time goes and therefore where the value is.
This stage should also produce the first piece of bad news, and if it does not, it has not been done properly. There is nearly always something in the original scope that will not work, and finding it in week two costs a fraction of finding it in month four.
Stage two: getting at the data
The stage that consumes the most calendar time and the least attention in planning. The data is in systems owned by teams who were not consulted about this project, access requires approvals with their own queues, and the person who understood the schema has left.
There is a question worth asking before any timeline is agreed, and it is uncomfortable because the honest answer is often embarrassing: how long did it take, the last time, to get a read credential on the main system involved. Organisations that can answer in days and organisations that cannot are running different projects, and pretending otherwise produces a plan that is wrong from the first week.
The artefact this stage should produce is undervalued: a written account of what the fields actually contain as opposed to what they are named. Every organisation has data that is systematically wrong in a predictable way, corrected mentally by the people who use it, and written down nowhere. That document tends to outlive the project.
Stage three: building
The stage everyone plans for and the one least likely to be the problem. It is ordinary engineering against your constraints, with one discipline that distinguishes deployment work from product work: the system must fail visibly rather than confidently.
A model that is fluently wrong on the cases your experts care about will destroy trust faster than a system that declines to answer. The design decision that matters most is therefore not about accuracy in aggregate; it is about what happens on the inputs where the system is least reliable, and whether the person reading the output can tell which those were.
The evaluation set built during this stage is the project's political instrument as much as its technical one. Building it forces your own experts to agree in advance on what a correct answer looks like, and the arguments that produces are the same disagreements that would otherwise surface in month five as a claim that the system does not work.
Stage four: getting it used
The stage that decides whether the previous three counted, and the one most projects treat as training. It is not training. The people whose work changes did not ask for this, may be measured on the old process, and owe nobody their cooperation.
What works is watching them not use it, one at a time. The reasons people abandon a working system are rarely what they say when asked directly: it is three clicks too deep, it does not show the one number they are accountable for, or it was wrong once in front of their manager. All three are fixable and none appears in a requirements document.
The other half is identifying who inside the organisation wants this to work, and making the outcome their win rather than the vendor's. Deployments that make an internal champion look right propagate after the vendor leaves. Deployments that make the vendor look clever stop the day they go.
Stage five: keeping it working
Systems drift. The upstream export changes format, a model is updated, a process the system assumed is reorganised. Without someone watching, the decay is silent: the system keeps producing output and the output gets quietly worse.
This is why the evaluation set is the most durable deliverable. Run monthly against a fixed set of cases, it turns silent decay into a visible number, and it is the only mechanism that reliably catches a degradation before a user does. It also outlives the model: when you change provider in two years, the evaluation set is what tells you whether the change was an improvement.
The three questions that predict a timeline
Ask these before scoping anything, and treat hesitation as information rather than as evasion.
Who can grant read access to the systems involved, and how long did it take last time. Who is accountable today for the outcome this is meant to improve, by name. And what happens to the people currently doing this work, because if nobody has decided, they will decide for themselves that the project is a threat, and they will be neither wrong nor cooperative.
A project where all three have clear answers can be planned. A project where any of them is unresolved can still be worth doing, but the plan should say so rather than assume the answer will arrive.
Bring one problem, not a roadmap
The most useful first conversation is about a single specific thing that takes too long today. Thirty minutes is usually enough to tell whether it is worth building, and we will say when it is not.
A first conversation is thirty minutes and is not a sales call. If the answer is that you do not need us, that is a useful outcome and we will say so.