Strategy

The Pilot That Never Ends: How to Run a Trial That Actually Decides Something

SKIMBOX Team

A pilot with no success criteria and no end date is not a trial, it is a slow purchase nobody has approved. Here is what has to be agreed before it starts, and how to tell a real proof of concept from a permanent one.

The Pilot That Never Ends: How to Run a Trial That Actually Decides Something

Somebody suggests trying it before committing. Everybody agrees, because trying something is obviously more sensible than buying it untested.

Eighteen months later the trial is still running. Four departments depend on it, a year of data sits inside it, and the question of whether to buy it has never been formally asked, because the answer became obvious somewhere around month five without anybody noticing.

You did not pilot that tool. You bought it, in instalments, without an approval.

Why pilots do not end

A decision needs three things: a question, a way to answer it, and somebody who decides. Most pilots have none of them.

The question is usually "let's see how it goes", which nothing can fail. The measure is an impression formed by whoever used it most. And the decider is unnamed, so the outcome is settled by inertia.

Under those conditions a pilot cannot conclude, so it continues. Meanwhile the cost of stopping rises every week: more data inside the tool, more people trained on it, more informal process built around it. By the time anybody asks the purchase question, the honest answer is that it was answered months ago by accumulation.

This is not a failure of discipline by the people running it. It is what happens by default when nobody set the exit conditions, and it is entirely preventable in about an hour of work before the pilot starts.

Nobody involved wants it to end

Worth naming the incentives, because they explain why this keeps happening to sensible people.

The supplier does not want the pilot to conclude, because a running pilot is a customer accumulating switching costs while a concluded pilot is a decision that might go against them. Every additional week of free access is cheap for them and expensive for you to unwind. This is not sharp practice, it is simply how the commercial logic works, and it operates whether or not anybody is being cynical.

The internal champion does not want it to conclude either, because they advocated for the tool and an inconclusive continuation is more comfortable than a verdict. Extending is never framed as avoidance. It is framed as being thorough, wanting more data, giving it a fair chance.

And the people using it have adapted. Six weeks in, they have built habits around it and the workarounds have become invisible. Ask them whether it is working and they will say yes, because the alternative is learning something else.

So the only person with an interest in the pilot ending is the person paying for what comes next, and that person is usually not in the room where the extension gets agreed.

That is why the end date and the named decider have to be set at the start by somebody outside the pilot. Left to the people running it, no pilot ever concludes, and none of them will have done anything wrong.

The five things to agree first

Public bodies run pilots constantly and have had to formalise what makes one useful. The US Government Accountability Office applies a consistent set of leading practices when it reviews federal pilot programmes: clear and measurable objectives, communication with stakeholders, an assessment methodology and data-gathering strategy, an evaluation plan that names who does what, and criteria for assessing whether the pilot would scale [1][2].

Those translate directly to a business trialling software.

What question does this answer? Specific enough that it can fail. "Can our operations team process a claim end to end without leaving the tool" is a question. "Evaluating whether the platform meets our needs" is not, because there is no result that would count as no.

How will we measure it? Decided before the pilot starts, not after. This is the one people skip, and it matters because whoever picks the measure after seeing the results will pick one that supports the conclusion they already hold.

How long does it run? Long enough to cover one full cycle of whatever you are testing. If the process has a monthly close, a two-week pilot proves nothing about the hardest part. If it is a daily workflow, six weeks is generous. Take the date from the cycle, not the calendar.

Who decides? One named person, agreed in advance, who can say no as well as yes. Ideally not the person who championed the pilot.

What happens in each outcome? All three of them, which brings us to the part that actually causes the problem.

Plan for the inconclusive result

Everybody plans for yes and no. Almost nobody plans for the outcome that actually occurs most often, which is that the results are ambiguous.

Faced with an ambiguous result and no prepared response, a team extends the pilot. It is the reasonable-sounding option, it postpones a difficult conversation, and it requires no approval. That single move is how a six-week trial becomes a two-year arrangement.

So decide in advance what inconclusive triggers. Usually it should be one short, specific extension with a new and narrower question, or a decision to stop and revisit later. What it should never be is "carry on and see".

Write down what the extension would test and what it would cost. If you cannot describe that in advance, the honest position is that an ambiguous result means no.

What a well-designed pilot looks like on one page

Concretely, because the list above can read as bureaucracy until you see how short it actually is.

Question: Can our three-person accounts team issue, send and reconcile a full month of invoices in this tool without falling back to the current spreadsheet?

Measure: Number of invoices issued in the tool versus outside it during October. Time to close the month, against a September baseline of six working days. Count of errors found at reconciliation.

Duration: One full monthly cycle, 1 to 31 October. Decision meeting 5 November.

Participants: All three accounts staff, including the one who did not want to change tools, plus the finance manager who signs off the close.

Decider: Finance director. Not attending the pilot, not the person who proposed it.

If it works: Adopt from January, budget for licences and two days of setup, retire the spreadsheet by end of Q1.

If it does not: Stop, revert to the spreadsheet, and revisit only if the close time exceeds eight days for two consecutive months.

If it is inconclusive: One further cycle in November testing only the reconciliation step, which is where we expect the difficulty. No further extension after that.

Data: Real invoices. Export confirmed available as CSV. If we do not proceed, data is exported and the account is closed within 14 days.

That is the entire design, and it takes twenty minutes to write. Every part of it is a decision that would otherwise be made later, under pressure, by whoever felt most strongly.

Notice particularly the last two lines. Deciding in advance that November tests only reconciliation prevents the vague second extension, and confirming the export before starting prevents the discovery, in month four, that leaving is harder than arriving.

Who takes part changes the answer

A pilot run by five enthusiastic volunteers proves that five enthusiastic volunteers can make the tool work. It tells you very little about two hundred people who did not choose it, and this is exactly the scalability question GAO treats as a separate design requirement [1].

Choose participants deliberately. Include people who are sceptical, people who are busy, and people who are less confident with technology, because those are the population you will roll out to. Their experience is what predicts adoption; the enthusiasts' experience predicts only their own.

A cross-section of roles matters more than a headcount. Five people covering three different jobs will tell you more than twenty people all doing the same one.

And keep the measurement on your side. A supplier can configure the tool, train your team and support the whole exercise, but a supplier who also owns the assessment and the conclusion has an obvious interest in the result. Objective, measure and decision stay with you even when all the technical work sits with them.

Free trials are not free

The cost of a free pilot is your own people's time plus the switching cost you accumulate while it runs.

Every week puts more of your data in somebody else's system, trains more of your staff on their interface, and builds more informal process around their way of doing things. That accumulation is not an accident. It is what a generous trial period is designed to produce, and it works.

This gives you the real design constraint for a pilot: it should be small enough that stopping is easy. If a pilot has cost enough that abandoning it would be embarrassing, the sunk cost is now making the decision for you and you have lost the main benefit of piloting at all. Keep it cheap specifically so that no remains a comfortable answer.

Two related questions worth asking before the first login. Where will our data sit, and how do we get it back? A trial that is easy to enter and hard to leave has made the exit part of its commercial design. Ask how you export everything, in what format, and how long you have after cancelling. The answer tells you whether the supplier expects to win on merit.

If a discount appears that expires before your agreed decision date, ask for it to be held until then. You are being asked to decide without the evidence the pilot was meant to produce. The response to that request is itself useful information.

Pilots and proofs of concept are different things

A proof of concept answers whether something is technically possible. It is fast, narrow, and the code is meant to be thrown away.

A pilot answers whether something works in your business, with your people and your data. Different question, different design, different participants.

Confusing the two is expensive in one specific direction: treating a successful proof of concept as though it settled the business question. It did not. It established that the integration is feasible, which is genuinely useful and tells you nothing about whether your team will use the result.

The other expensive move is putting proof of concept code into production. It was written to answer a question quickly, which means without the error handling, security hardening and testing production work requires. Once it is running and useful, replacing it never reaches the top of anybody's list, and you have acquired a permanent system built to temporary standards. Our guide on what fast-built software costs later covers where that leads.

The fix is to say so before it is written. If the code might realistically go into production, do not call it a proof of concept and do not price it as one. Build the small version properly instead.

The cost of a pilot that never ended

Worth being concrete about what this actually costs, because "it just carried on" sounds harmless.

You pay list price rather than a negotiated rate, because the negotiation happens at the point of purchase and you never reached one. You lose the bargaining position a genuine comparison would have given you, since the supplier knows perfectly well that your data and your habits are already inside their product.

You also carry a tool nobody assessed against the alternatives, which means any advantage a competing product had is now invisible and will stay that way. And you have trained an organisation to treat trials as adoptions, which makes the next pilot harder to run honestly, because everybody already knows how it ends.

The subtler cost is that the decision never got the scrutiny it deserved. A purchase goes through budget approval, a comparison, and somebody asking what it replaces. A drifting pilot skips all three. So the tool your business now depends on is the one item in your stack that nobody ever formally examined, and often it is a fairly significant one.

None of that is catastrophic on any single occasion. It is simply worse than the outcome you would have got from an hour of design at the start.

Reading the result honestly

Two failure modes at the end, and they pull in opposite directions.

The first is declaring success because the pilot completed. Finishing is not the same as passing. Go back to the written objective and the written measure and check against those specifically, not against the general feeling in the room.

The second is treating poor adoption as a verdict on the software. If the tool worked but people did not use it, that is a genuine finding about process and accountability rather than about the product, and buying the full licence will not close a gap the pilot has just identified for you. Our guide on whether it is actually a software problem covers that distinction, and a pilot that surfaces it has done its job even though it feels like a failure.

Write two pages at the end: the original objective, the measure, what happened against it, what surprised you, what adopting it properly would cost, and a recommendation. The test of that document is whether somebody uninvolved could read it in a year and understand the basis of the decision.

When not to bother

Pilots are not free and they are not always warranted.

For a widely used tool with quick setup and an easy exit, a pilot can cost more in attention than simply adopting it and reversing if it disappoints. The overhead of designing, measuring and reporting on a trial is real.

Pilots earn their keep where switching later would be expensive, where the fit is genuinely uncertain rather than merely unfamiliar, or where you need evidence to persuade somebody else. Outside those cases, consider just deciding.

If you do want help: defining the question, the measure and the decision points for a pilot, or reviewing one already running to establish what it can and cannot conclude, starts from around AED 1,500 with us. A fuller discovery producing a written specification you own outright starts from around AED 4,000. Final pricing depends on scope, and these are our own figures rather than a market survey.

The thing to do today costs nothing. List every pilot currently running in your business, and against each one write the end date and the name of the person who decides. Where either is missing, that arrangement stopped being a trial some time ago, and the hour it takes to work out when usually ends at least one of them.

References

  1. US Government Accountability Office, leading practices for pilot program design (GAO-24-106847)
  2. US Government Accountability Office, DATA Act: Section 5 pilot design issues (GAO-16-438)
  3. US Digital Services Playbook
  4. SKIMBOX, is this actually a software problem
  5. SKIMBOX, the real cost of fast-built software
  6. SKIMBOX, build versus buy for UAE businesses

GAO's pilot design practices are summarised above from how GAO applies them across its published federal pilot reviews rather than quoted from a single canonical list. They concern US public programmes and are cited for the underlying design principle, which applies broadly.

Frequently asked questions

  • What is wrong with most software pilots?

    They have no definition of success, no end date, and no agreement about who decides. Without those three, a pilot cannot conclude, so it continues. What began as a low-commitment trial becomes a system the business depends on, bought without anybody approving the purchase, setting a budget for it, or comparing it against the alternatives it was supposed to be tested against.

  • How do I know if my pilot has become permanent?

    Ask when it started and when it was meant to end. If the second date has passed, or was never set, you are no longer piloting. Another reliable test is whether anybody would notice if you switched it off tomorrow. If real work would stop, the trial concluded some time ago and nobody wrote the decision down, which means it was never actually weighed against anything.

  • What should be agreed before a pilot starts?

    What question the pilot answers, how you will measure the answer, how long it runs, who decides at the end, and what happens in each outcome. Five things, all agreed in writing before anybody logs in. Five things, none of which is technical, and all of which take about twenty minutes to write. Skipping them is the entire reason so many pilots drift quietly into becoming the default answer.

  • What does a good pilot objective look like?

    Specific and measurable rather than exploratory. Government guidance on designing pilots calls for well-defined, appropriate, clear and measurable objectives, which is a useful bar. Can our operations team process a claim end to end without leaving the tool is an objective. Evaluating whether the platform meets our needs is not an objective, because there is no result that would count as a failure against it.

  • How do I measure a pilot?

    Decide the method before you start rather than afterwards, and write it down where everybody can see it. Public sector guidance on pilot design treats an assessment methodology and a data-gathering strategy as core requirements, and the reason is straightforward: whoever chooses the measure after seeing the results will choose one that supports the conclusion they had already reached, and will do so entirely sincerely.

  • How long should a pilot run?

    Long enough to cover a full cycle of whatever you are testing, and no longer. If the process you are trialling has a monthly close, a two-week pilot proves nothing about the part that matters most. If it is a daily workflow, six weeks is usually generous. Take the end date from the cycle you are testing rather than from a round number on the calendar, because a pilot that stops before the hard part has proved nothing.

  • Should the pilot end date be firm?

    Yes, and the decision should be scheduled as a meeting rather than left to arise naturally. A pilot with a soft end date does not end. Put a date in the diary with the named decider present, and give the meeting an agenda that includes stopping as an explicit option. That single step converts a drift into a decision that somebody has to actually make.

  • Who should decide the outcome?

    One named person, agreed before it starts. Pilots frequently produce genuinely ambiguous results, and ambiguity plus no decider produces continuation by default. Name the person who can say no as well as yes, and where you can, make sure it is not the person who championed the pilot in the first place. They are not neutral, through no fault of their own.

  • What outcomes should I plan for?

    Three, written down in advance. It worked, so here is what adopting it involves and costs. It did not, so here is what we do instead. Or it was inconclusive, so here is the specific thing we would need to test next and what that costs. That third path is where most pilots actually land, and it is the one for which almost nobody prepares a response in advance, which is precisely why it produces an extension instead of a decision.

  • Why does the inconclusive outcome matter so much?

    Because it is the most common result and the one with no prepared response. Faced with an ambiguous outcome and no plan, teams extend the pilot rather than deciding, which is how a six-week trial becomes a two-year arrangement. Deciding in advance exactly what an inconclusive result triggers, including what the extension would test and what it would cost, removes that default entirely.

  • What is scalability and why should I test it?

    Whether what worked for a small group would still work for everybody. Government pilot guidance treats criteria for assessing scalability as a distinct design requirement, and it catches a specific failure: a pilot run by five enthusiastic volunteers proves that five enthusiastic volunteers can use the tool. It says very little about two hundred people who did not choose it, are busy, and have an existing way of working that already functions well enough for them.

  • How do I choose who takes part in a pilot?

    Deliberately, and not only from volunteers. Enthusiasts will make almost anything work and their success does not generalise. Include at least some people who are sceptical, some who are busy, and some who are less confident with technology, because between them they represent the population you will eventually roll out to. Their experience predicts adoption. The enthusiasts' experience predicts only their own.

  • Should suppliers run the pilot?

    They can support it, and they should not own the measurement or the conclusion. A supplier configuring the tool, training your team and also assessing whether it worked has an obvious interest in the answer. Keep the objective, the measure and the final decision firmly on your side, even in cases where every piece of the technical work sits with them and you could not do it yourself.

  • Is a free pilot actually free?

    Rarely, and the cost is mostly your own people's time plus the switching cost you accumulate. Every week of a free pilot puts more of your data in somebody's system, trains more of your staff on their interface, and builds more informal process around it. That accumulating commitment is precisely what a generous free trial period is designed to create, and it works reliably. It is not sharp practice, it is simply how the arrangement is built.

  • How does a pilot become a purchase nobody approved?

    Gradually and without any single decision. Data goes in, people learn it, a workflow forms around it, and by the time somebody asks whether to buy it the honest answer is that you already did. The commitment was made in small increments, none of which individually required an approval, which is exactly why the end date matters more than any other term you agree.

  • What is the difference between a pilot and a proof of concept?

    A proof of concept answers whether something is technically possible, usually quickly and with a small piece of throwaway work. A pilot answers whether it works in your business with real people and real data. They need different designs, different participants and different measures. Treating a successful proof of concept as though it had settled the business question is a common and expensive confusion.

  • Should proof of concept code go into production?

    Almost never, and it should be agreed up front that it will not. Code written to answer a question fast is built without the error handling, security or testing production requires. Once it is running and people find it useful, replacing it never reaches the top of anybody's priority list, and you have quietly acquired a permanent system built to deliberately temporary standards.

  • How do I stop a proof of concept becoming the product?

    Say so before it is written, in the same document that scopes it, and plan for the rebuild rather than hoping it will not be needed. If there is any realistic chance it goes into production, do not call it a proof of concept and do not price it as one. Build a small version properly instead, and accept that it costs more than a throwaway would have.

  • What if the pilot shows the tool is fine but adoption is poor?

    That is a genuine finding rather than a failed pilot, and it usually means the problem sits in the process or the accountability rather than the software. Our guide on whether something is actually a software problem covers that distinction in full. Buying the licence will not close a gap in accountability that the pilot has just helpfully identified for you at low cost.

  • Should I run two tools in parallel?

    Only for a short, defined period and with the same success measure applied to both. Comparative pilots are genuinely useful and they double the effort, so keep them short. What a comparative pilot must never become is a permanent arrangement in which different teams use different tools because nobody ever made the call. That is the most expensive outcome available and it is depressingly common.

  • How much should a pilot cost?

    Small enough that stopping is easy. That is the real design constraint. If the pilot has cost enough that abandoning it would be embarrassing, the sunk cost is now making your decision for you, and you have lost the main benefit of piloting at all. Keep it deliberately cheap so that no remains a comfortable answer right up until the decision date, because the moment stopping becomes embarrassing you have lost the main benefit of piloting.

  • What if the supplier offers a discount to convert now?

    Treat urgency during a trial as information about their process rather than about your decision. A discount that expires before your pilot ends is asking you to decide without the evidence the pilot was meant to produce. Ask for the offer to be held open until your agreed decision date. Most suppliers will agree, and the ones who refuse have told you something useful about how they expect to win the business.

  • Should I pilot with real data?

    Usually yes for realism, and with care about which data. Test data hides the problems real records cause: duplicates, missing fields, and the odd historical entries every business has. Use real data wherever you can do so safely, and establish before you start where it will be stored, who can see it, and what happens to it if you decide not to proceed.

  • What happens to my data if I stop the pilot?

    Ask before you start rather than after you decide. How do I export everything, in what format, and how long do I have to do it? A trial that is easy to enter and hard to leave has made the exit part of its commercial design. How readily a supplier answers that question is a fair test of whether they expect to win the business on merit.

  • How many people should be in a pilot?

    Enough to be representative rather than enough to be statistically rigorous, which for most businesses is unattainable and not really the point. A deliberate cross-section of roles matters far more than any particular headcount. Five people covering three genuinely different jobs will tell you considerably more than twenty people all doing the same one, because the failure modes cluster at the boundaries between roles.

  • What should the pilot report contain?

    The original objective, the measure, what actually happened against it, what surprised you, what it would cost to adopt properly, and a recommendation. Two pages. Two pages, no more. The test is whether somebody who was not involved could read it and understand the basis of the decision, including in a year's time when nobody remembers any of the detail.

  • Do I need a pilot at all?

    Not always. For a widely used tool with a short setup and easy exit, a pilot can cost more in attention than simply adopting it and reversing if necessary. Pilots earn their keep where switching later would be expensive, where the fit is genuinely uncertain rather than merely unfamiliar, or where you need documented evidence to persuade somebody else in the business.

  • Can a pilot be too long?

    Easily, and length is the main way pilots become permanent. Beyond the natural cycle of what you are testing, extra time adds commitment rather than information. If you are three months into what was agreed as a six-week pilot, the additional weeks are no longer producing evidence about the tool. They are producing dependency on it, which is a different thing entirely.

  • Can you help design or review a pilot?

    We can. Defining the question, the measure and the decision points, or reviewing a pilot already running to establish what it can and cannot conclude, starts from around AED 1,500 with us. A fuller discovery producing a written specification you own starts from around AED 4,000. Final pricing depends on scope, and these are our own figures rather than a market survey, since no official body publishes rates for this kind of work.

  • What is the single thing to fix today?

    Put an end date and a named decider on every pilot currently running, and write down what success would look like even retrospectively. If you cannot state what would make one of them a failure, that pilot is not a trial and has not been for some time. That exercise takes about an hour across a whole business, costs nothing, and in our experience usually ends at least one arrangement that everybody had stopped questioning.

SKIMBOX Team

Tech Consultancy

Get fresh writing in your inbox

One email a fortnight. No filler.

By subscribing, you agree to our privacy policy.

Want us to build something?

We work with teams across MENA, UK, USA, and India to build products, run programs, and grow.

Get in touch

Continue reading