Back to insights

How to choose your first AI pilot.

A first pilot is not a demonstration. It is a test designed to produce a decision you are willing to act on, in a workflow you can measure, with a named person accountable for the output.

AI adoption, practical guideBrainforest

A useful first pilot starts with a problem narrow enough to measure and important enough to matter. Agree in advance what evidence would support expansion, revision or stopping.

First, check that AI is the right instrument

Three different interventions get requested under the same label, and they have very different costs and failure modes.

ApproachFits whenMain risk
Rules-based automationThe steps are known, the inputs are structured, and the same input should always produce the same output.Brittleness when the process has undocumented exceptions.
AI assistanceInputs are unstructured or inconsistent, judgment is involved, and a person can verify the output against a source.Plausible but wrong output that survives a shallow review.
Process changeThe work is slow because of handoffs, missing fields, unclear ownership or waiting, not because of effort.Changing handoffs or ownership without involving the people who do the work.

Compare all three before committing. If a required field, a shared template or a named owner would address the problem, test that change before adding AI. A documented process also gives any later AI evaluation a clearer starting point.

Define one workflow, one owner, one baseline

Starting with one workflow can make outcomes easier to isolate. A multiworkflow pilot is not inherently uninterpretable, but it needs separate baselines, owners and decision rules for each workflow. Follow each tested workflow from trigger to completed output.

Name a single accountable owner. Not a sponsor and not a committee: the person who will review outputs, decide when something is unacceptable, and report the result. If no one will take that role, the pilot is not ready.

Then measure the current state before anything changes. Choose a sample size based on the variation in the work, the cost of errors and the decision the pilot must support. Include difficult cases, current error and rework, current cost, and relevant exceptions. Do not describe a small convenience sample as statistically representative.

Settle data permissions before the trial

Decide, in writing and with whoever owns information governance, which data the pilot may use, where it may be processed, how long it is retained, and whether it may be used to improve a vendor model. The data owner should confirm purpose, access, retention and the permissions required for this specific use. An existing policy alone does not establish authorization. Confirm that access can be revoked cleanly when the pilot ends.

Resolve these questions before the trial. If authorization is unclear, pause the pilot until the accountable data owner and governance team have answered it.

Set review, reversal and error cost rules

For every output the pilot produces, answer three questions. Who reviews it. How the reviewer verifies it against a source rather than judging whether it reads well. And what it costs if a wrong output reaches a customer, a payment, a regulator or a commitment.

Prefer workflows where the error is caught before it propagates and the action can be reversed. A drafted summary that a person approves is a good first candidate. An automatically sent message or an executed transaction is not. Exclude high-stakes human resources, legal, medical and safety decisions from this illustrative first-pilot pattern.

An illustrative example: internal knowledge retrieval

Illustrative example only. This is a hypothetical scenario used to show the method. It contains no client work and no measured results.

A hypothetical operations team answers internal questions using approved, low-risk policy documents and guidance. The proposed pilot is a retrieval assistant that returns a drafted answer with links to the source passages it used, while enforcing each user's existing source access.

This can be tested against agreed criteria. The approved evidence is available internally. A reviewer checks whether each citation supports the statement and whether the source is current; a citation does not ensure correctness. The answer goes to an eligible colleague who already has access to the source and is expected to confirm it. Nothing irreversible happens automatically.

The same team asking for an assistant that approves expenses or decides eligibility is a different proposition. Those decisions commit the organization and require separate controls and governance. High-stakes HR, legal, medical and safety decisions are outside this example.

Run a representative trial with pass and stop conditions

Agree the conditions before the trial starts so the final decision can refer to the same criteria the team accepted at the beginning.

  1. Sample. A fixed set of real cases, chosen to include the awkward ones, not a curated set that flatters the tool.
  2. Duration. Long enough for the team to stop performing for the pilot and simply work.
  3. Pass condition. A stated quality bar plus a stated improvement over baseline, both measured after review effort is included.
  4. Stop condition. What result ends the pilot immediately. A rate of unsupported statements, a governance breach, or reviewers quietly abandoning the tool are all legitimate stop conditions.
  5. Decision date and decision maker. Written down at the start.

Include an adoption measure agreed for the eligible pilot users. Results from two enthusiastic users should not be generalized to the whole workforce, but a bounded pilot also need not measure people who were never included in its agreed scope.

Measure honestly, including the work the pilot creates

Do not compare the old task time only against the new generation time and call the difference a saving. The assisted time must include reading, verifying, correcting and occasionally discarding the output, plus the time spent on cases the tool handled badly.

  • Time per completed, accepted task, before and after.
  • Review and rework effort, counted separately so it cannot be hidden.
  • Quality of the completed work against the agreed bar.
  • Exception frequency and what happened to those cases.
  • Actual operating cost: licences, usage, integration, administration and training.
  • Effective use among the agreed eligible pilot users, not only the strongest individual result.

Then distinguish released capacity from financial value. Time released is capacity. Record any actual incremental contribution separately, and record avoidable cash cost only when a cost truly changes. Do not count the same hours twice as capacity and payroll saving. A pilot can still provide a useful stop or redesign decision when benefits do not exceed review effort and operating cost.

Decide, then write the next brief

A pilot should end in one of three explicit outcomes: expand, refine and retest with a named change, or stop. Record which, and why, with the evidence attached. That learning record informs the next decision; it does not guarantee that a later pilot will be faster or cheaper.