Somebody suggests trying it before committing. Everybody agrees, because trying something is obviously more sensible than buying it untested.
Eighteen months later the trial is still running. Four departments depend on it, a year of data sits inside it, and the question of whether to buy it has never been formally asked, because the answer became obvious somewhere around month five without anybody noticing.
You did not pilot that tool. You bought it, in instalments, without an approval.
Why pilots do not end
A decision needs three things: a question, a way to answer it, and somebody who decides. Most pilots have none of them.
The question is usually "let's see how it goes", which nothing can fail. The measure is an impression formed by whoever used it most. And the decider is unnamed, so the outcome is settled by inertia.
Under those conditions a pilot cannot conclude, so it continues. Meanwhile the cost of stopping rises every week: more data inside the tool, more people trained on it, more informal process built around it. By the time anybody asks the purchase question, the honest answer is that it was answered months ago by accumulation.
This is not a failure of discipline by the people running it. It is what happens by default when nobody set the exit conditions, and it is entirely preventable in about an hour of work before the pilot starts.
Nobody involved wants it to end
Worth naming the incentives, because they explain why this keeps happening to sensible people.
The supplier does not want the pilot to conclude, because a running pilot is a customer accumulating switching costs while a concluded pilot is a decision that might go against them. Every additional week of free access is cheap for them and expensive for you to unwind. This is not sharp practice, it is simply how the commercial logic works, and it operates whether or not anybody is being cynical.
The internal champion does not want it to conclude either, because they advocated for the tool and an inconclusive continuation is more comfortable than a verdict. Extending is never framed as avoidance. It is framed as being thorough, wanting more data, giving it a fair chance.
And the people using it have adapted. Six weeks in, they have built habits around it and the workarounds have become invisible. Ask them whether it is working and they will say yes, because the alternative is learning something else.
So the only person with an interest in the pilot ending is the person paying for what comes next, and that person is usually not in the room where the extension gets agreed.
That is why the end date and the named decider have to be set at the start by somebody outside the pilot. Left to the people running it, no pilot ever concludes, and none of them will have done anything wrong.
The five things to agree first
Public bodies run pilots constantly and have had to formalise what makes one useful. The US Government Accountability Office applies a consistent set of leading practices when it reviews federal pilot programmes: clear and measurable objectives, communication with stakeholders, an assessment methodology and data-gathering strategy, an evaluation plan that names who does what, and criteria for assessing whether the pilot would scale [1][2].
Those translate directly to a business trialling software.
What question does this answer? Specific enough that it can fail. "Can our operations team process a claim end to end without leaving the tool" is a question. "Evaluating whether the platform meets our needs" is not, because there is no result that would count as no.
How will we measure it? Decided before the pilot starts, not after. This is the one people skip, and it matters because whoever picks the measure after seeing the results will pick one that supports the conclusion they already hold.
How long does it run? Long enough to cover one full cycle of whatever you are testing. If the process has a monthly close, a two-week pilot proves nothing about the hardest part. If it is a daily workflow, six weeks is generous. Take the date from the cycle, not the calendar.
Who decides? One named person, agreed in advance, who can say no as well as yes. Ideally not the person who championed the pilot.
What happens in each outcome? All three of them, which brings us to the part that actually causes the problem.
Plan for the inconclusive result
Everybody plans for yes and no. Almost nobody plans for the outcome that actually occurs most often, which is that the results are ambiguous.
Faced with an ambiguous result and no prepared response, a team extends the pilot. It is the reasonable-sounding option, it postpones a difficult conversation, and it requires no approval. That single move is how a six-week trial becomes a two-year arrangement.
So decide in advance what inconclusive triggers. Usually it should be one short, specific extension with a new and narrower question, or a decision to stop and revisit later. What it should never be is "carry on and see".
Write down what the extension would test and what it would cost. If you cannot describe that in advance, the honest position is that an ambiguous result means no.
What a well-designed pilot looks like on one page
Concretely, because the list above can read as bureaucracy until you see how short it actually is.
Question: Can our three-person accounts team issue, send and reconcile a full month of invoices in this tool without falling back to the current spreadsheet?
Measure: Number of invoices issued in the tool versus outside it during October. Time to close the month, against a September baseline of six working days. Count of errors found at reconciliation.
Duration: One full monthly cycle, 1 to 31 October. Decision meeting 5 November.
Participants: All three accounts staff, including the one who did not want to change tools, plus the finance manager who signs off the close.
Decider: Finance director. Not attending the pilot, not the person who proposed it.
If it works: Adopt from January, budget for licences and two days of setup, retire the spreadsheet by end of Q1.
If it does not: Stop, revert to the spreadsheet, and revisit only if the close time exceeds eight days for two consecutive months.
If it is inconclusive: One further cycle in November testing only the reconciliation step, which is where we expect the difficulty. No further extension after that.
Data: Real invoices. Export confirmed available as CSV. If we do not proceed, data is exported and the account is closed within 14 days.
That is the entire design, and it takes twenty minutes to write. Every part of it is a decision that would otherwise be made later, under pressure, by whoever felt most strongly.
Notice particularly the last two lines. Deciding in advance that November tests only reconciliation prevents the vague second extension, and confirming the export before starting prevents the discovery, in month four, that leaving is harder than arriving.
Who takes part changes the answer
A pilot run by five enthusiastic volunteers proves that five enthusiastic volunteers can make the tool work. It tells you very little about two hundred people who did not choose it, and this is exactly the scalability question GAO treats as a separate design requirement [1].
Choose participants deliberately. Include people who are sceptical, people who are busy, and people who are less confident with technology, because those are the population you will roll out to. Their experience is what predicts adoption; the enthusiasts' experience predicts only their own.
A cross-section of roles matters more than a headcount. Five people covering three different jobs will tell you more than twenty people all doing the same one.
And keep the measurement on your side. A supplier can configure the tool, train your team and support the whole exercise, but a supplier who also owns the assessment and the conclusion has an obvious interest in the result. Objective, measure and decision stay with you even when all the technical work sits with them.
Free trials are not free
The cost of a free pilot is your own people's time plus the switching cost you accumulate while it runs.
Every week puts more of your data in somebody else's system, trains more of your staff on their interface, and builds more informal process around their way of doing things. That accumulation is not an accident. It is what a generous trial period is designed to produce, and it works.
This gives you the real design constraint for a pilot: it should be small enough that stopping is easy. If a pilot has cost enough that abandoning it would be embarrassing, the sunk cost is now making the decision for you and you have lost the main benefit of piloting at all. Keep it cheap specifically so that no remains a comfortable answer.
Two related questions worth asking before the first login. Where will our data sit, and how do we get it back? A trial that is easy to enter and hard to leave has made the exit part of its commercial design. Ask how you export everything, in what format, and how long you have after cancelling. The answer tells you whether the supplier expects to win on merit.
If a discount appears that expires before your agreed decision date, ask for it to be held until then. You are being asked to decide without the evidence the pilot was meant to produce. The response to that request is itself useful information.
Pilots and proofs of concept are different things
A proof of concept answers whether something is technically possible. It is fast, narrow, and the code is meant to be thrown away.
A pilot answers whether something works in your business, with your people and your data. Different question, different design, different participants.
Confusing the two is expensive in one specific direction: treating a successful proof of concept as though it settled the business question. It did not. It established that the integration is feasible, which is genuinely useful and tells you nothing about whether your team will use the result.
The other expensive move is putting proof of concept code into production. It was written to answer a question quickly, which means without the error handling, security hardening and testing production work requires. Once it is running and useful, replacing it never reaches the top of anybody's list, and you have acquired a permanent system built to temporary standards. Our guide on what fast-built software costs later covers where that leads.
The fix is to say so before it is written. If the code might realistically go into production, do not call it a proof of concept and do not price it as one. Build the small version properly instead.
The cost of a pilot that never ended
Worth being concrete about what this actually costs, because "it just carried on" sounds harmless.
You pay list price rather than a negotiated rate, because the negotiation happens at the point of purchase and you never reached one. You lose the bargaining position a genuine comparison would have given you, since the supplier knows perfectly well that your data and your habits are already inside their product.
You also carry a tool nobody assessed against the alternatives, which means any advantage a competing product had is now invisible and will stay that way. And you have trained an organisation to treat trials as adoptions, which makes the next pilot harder to run honestly, because everybody already knows how it ends.
The subtler cost is that the decision never got the scrutiny it deserved. A purchase goes through budget approval, a comparison, and somebody asking what it replaces. A drifting pilot skips all three. So the tool your business now depends on is the one item in your stack that nobody ever formally examined, and often it is a fairly significant one.
None of that is catastrophic on any single occasion. It is simply worse than the outcome you would have got from an hour of design at the start.
Reading the result honestly
Two failure modes at the end, and they pull in opposite directions.
The first is declaring success because the pilot completed. Finishing is not the same as passing. Go back to the written objective and the written measure and check against those specifically, not against the general feeling in the room.
The second is treating poor adoption as a verdict on the software. If the tool worked but people did not use it, that is a genuine finding about process and accountability rather than about the product, and buying the full licence will not close a gap the pilot has just identified for you. Our guide on whether it is actually a software problem covers that distinction, and a pilot that surfaces it has done its job even though it feels like a failure.
Write two pages at the end: the original objective, the measure, what happened against it, what surprised you, what adopting it properly would cost, and a recommendation. The test of that document is whether somebody uninvolved could read it in a year and understand the basis of the decision.
When not to bother
Pilots are not free and they are not always warranted.
For a widely used tool with quick setup and an easy exit, a pilot can cost more in attention than simply adopting it and reversing if it disappoints. The overhead of designing, measuring and reporting on a trial is real.
Pilots earn their keep where switching later would be expensive, where the fit is genuinely uncertain rather than merely unfamiliar, or where you need evidence to persuade somebody else. Outside those cases, consider just deciding.
If you do want help: defining the question, the measure and the decision points for a pilot, or reviewing one already running to establish what it can and cannot conclude, starts from around AED 1,500 with us. A fuller discovery producing a written specification you own outright starts from around AED 4,000. Final pricing depends on scope, and these are our own figures rather than a market survey.
The thing to do today costs nothing. List every pilot currently running in your business, and against each one write the end date and the name of the person who decides. Where either is missing, that arrangement stopped being a trial some time ago, and the hour it takes to work out when usually ends at least one of them.
References
- US Government Accountability Office, leading practices for pilot program design (GAO-24-106847)
- US Government Accountability Office, DATA Act: Section 5 pilot design issues (GAO-16-438)
- US Digital Services Playbook
- SKIMBOX, is this actually a software problem
- SKIMBOX, the real cost of fast-built software
- SKIMBOX, build versus buy for UAE businesses
GAO's pilot design practices are summarised above from how GAO applies them across its published federal pilot reviews rather than quoted from a single canonical list. They concern US public programmes and are cited for the underlying design principle, which applies broadly.



