How to Choose a First AI Pilot
A worksheet for defining an AI pilot’s task, data limits, owner, review process and stop condition, with a fictional help-desk example.

A useful first pilot tests one task that someone can review and an owner can stop. Write down the current process so you can compare the results and decide what to do next.
Use this worksheet to describe one candidate and decide whether it is ready to test. It is a planning aid, not legal, security, compliance, or procurement advice.
Choose the task before the technology
Start with a task or decision people repeat. Describe who does it, how it works now, what information it uses, and the failure you need to address. Use that description to define the AI’s role.
“Use a language model in customer support” is still too broad. “Draft a proposed reply for one internal support queue, using an approved knowledge base, with an agent reviewing every answer” is specific enough to question. That description lets you examine what access is needed, who reviews the answers, who owns the pilot, and how to evaluate it.
The first-pilot worksheet
Blank worksheet
- Task: What repeated task or decision are we examining?
- User: Who performs it, and who is affected by the result?
- Current process: What happens now, including reviews, delays, and known ways the process can fail?
- Proposed AI role: Is the system retrieving, classifying, extracting, drafting, summarising, or recommending? What remains human work?
- Allowed information: Which approved sources may the pilot use?
- Excluded information: Which personal, confidential, regulated, or otherwise restricted data must stay out?
- Owner: Who can change the pilot, pause it, and answer for the decision?
- Human review: What must a reviewer inspect before the output is used?
- Success evidence: How will we compare results on this task to decide whether to continue?
- Likely failure: What plausible error matters most, and how will the test expose it?
- Stop condition: What result, incident, or uncertainty ends or pauses the pilot?
- Next review: On what date will the owner decide to stop the pilot, revise it, extend it, or put it into regular use?
Use five checks to decide whether the pilot is ready
| Gate | Question | Not ready when… |
|---|---|---|
| Usefulness | Would a better result change real work? | No user owns the task or wants the output. |
| Boundaries | Can the task, users, sources, and exclusions be named? | The proposal depends on unrestricted access or undefined autonomy. |
| Review | Can a qualified person judge the output before it matters? | Errors are difficult to detect or are acted on automatically. |
| Evidence | Can the team compare the pilot with a current baseline on representative cases? | Success means only that the demo ran. |
| Reversibility | Can the owner pause the pilot and recover? | There is no stop mechanism, fallback, or accountable owner. |
Synthetic worked example
This example is fictional. It does not describe an employer, client, or measured deployment.
Internal help-desk reply drafts
- Task: Draft a response to common internal software-access questions.
- User: Help-desk agents; employees receive the final reply.
- Current process: An agent searches approved guidance, writes a response, and checks account-specific details.
- AI role: Retrieve passages from an approved knowledge base and draft a proposed response. It does not send messages or change access.
- Allowed information: Published internal guidance and a synthetic ticket set created for evaluation.
- Excluded information: Employee records, credentials, live ticket histories, and confidential attachments.
- Owner and review: The service-desk lead owns the test. An agent verifies the cited guidance, requested action, tone, and missing context before any use.
- Evidence: On a pre-agreed representative set, compare factual support, citation quality, missed questions, unsafe suggestions, reviewer effort, and task completion with the current process.
- Stop condition: Pause if the system exposes excluded data, fabricates access steps, or cannot reliably ground an answer in the approved source.
- Next review: The owner reviews the evidence after the fixed test set and decides whether to revise, stop, or design a controlled next phase.
One convincing example doesn’t establish reliability
Choose your evaluation set before tuning against it. Cover the range of cases the pilot needs to handle, including those where the correct behaviour is to abstain or ask for more information. Compare results with the current process and keep failures as evidence.
Passing a pilot gives the owner evidence for the next decision. Production monitoring, vendor risk, accessibility, security, legal obligations, cost at scale and organisational change still need to be addressed before rollout.
