Executive reading · ~60 seconds

An AI pilot does not exist to prove that a model produces artifacts, but to reduce the uncertainty of a decision. The unit that brings technology and business together is the accepted outcome: delivered result, appropriate to the quality criteria, within the cost and risk limits. The minimum panel combines outcome, acceptance, flow, savings and risk; tokens continue as an operational measure of consumption.

An AI pilot can produce a convincing demonstration in just a few days. The response comes quickly, the prototype looks smart, and the team finally sees tasks that could be automated. Then easy-to-celebrate numbers emerge: active users, prompts sent, tokens consumed, documents generated and hours apparently saved.

None of this alone answers the question that decides an investment: Did the work generate a change accepted by those who needed to receive it?

Tokens measure consumption. Artifacts measure production. Value appears when an outcome crosses the pipeline, reaches the recipient, meets an explicit quality criterion, and costs less—in money, time, risk, and attention—than the relevant alternative.

This paper proposes a simple contract to take AI pilots out of the demonstration theater and into a defensible operational decision.

Executive summary

A pilot does not exist to prove that a model can perform a task. It exists to reduce the uncertainty of a decision: expand, redesign, restrict or terminate a change in work.

Field studies show that gains can be real and materially different depending on the person, task and environment. In customer support, a survey of 5.172 agents found an average increase of 15% in problems resolved per hour, with a greater effect among professionals in the lowest skill quintile; the result belongs to a specific company and occupation. In three experiments with 4.867 developers, the data set showed 26.08% more tasks completed for those who had access to the wizard, but the experiments are noisy and do not measure broad software quality or net employment. [CAREER-A18-C1] [CAREER-A18-C2]

The responsible conclusion is not “AI increases productivity by X%”. The measurement design needs to capture the task, the population, the system, and the quality of the outcome. DORA 2025 reinforces this frontier by describing AI as an amplifier of the strengths and weaknesses of the organizational system, in research focused on software development. [CAREER-A18-C5]

Therefore, a pilot's minimum panel must combine outcome, acceptance, flow, savings and risk. The decision unit is the cost per accepted outcome, not the cost per token nor the raw quantity of artifacts.

Pilot is not a small version of the deployment

A demonstration answers “is it possible to produce something?”. A serious pilot answers “what do we need to learn before changing the system?”

This difference changes the entire design. Instead of just choosing cases that look good on a presentation, the pilot includes representative work, exceptions, people with different experience levels, and the full path to acceptance. Instead of hiding human review, it measures its cost. Instead of declaring success when the output appears, wait for the effect to reach the customer, user or team that needed it.

Before you begin, write down the decision the experiment is intended to inform:

If the flow produces more accepted outcomes, without exceeding the limits of quality, cost and risk, we will expand to this scope. If not, we will redesign or end the pilot.

Without this phrase, it's easy to change the definition of success after seeing the data.

The unit that avoids self-deception

Imagine a copilot who generates one hundred commercial proposals per day. If twenty reach the customer, eight respect the pricing policy and two move forward, “one hundred proposals” is a measure of production, not value. The same goes for lines of code, answered tickets, created campaigns or summary reports.

One outcome accepted must satisfy four conditions:

  1. reached the recipient or defined end state;
  2. passed known quality criteria before the experiment;
  3. preserved risk, security and liability restrictions;
  4. had total observable cost, including review, context, integration and operation.

The formula does not need to feign financial precision where it does not yet exist. You need to prevent an infrastructure unit from being presented as a business result:

custo por outcome aceito = custo total do fluxo / outcomes aceitos

Tokens, calls, seats and GPU time remain important. They explain consumption and help operate capacity. A FinOps Foundation distinguishes resource efficiency units from business-related units. A token can compose the numerator; it should not take the place of the outcome in the denominator.

Five dimensions, one decision

DimensionOperational questionMinimum signal
OutcomeHas anyone received a relevant change?outcomes delivered and accepted
AcceptanceHow much has passed without material correction?acceptance and rework rate
FlowIs the full path better?end-to-end and queue time
EconomyWhat was the actual cost per result?total cost per accepted outcome
Riskwhat failed and how reversible was it?incidents, severity and detection

The five dimensions form a set. Speed ​​without acceptance can only produce inventory. Quality without operable cost may not scale. Risk-free savings can transfer a larger bill into the future. A pilot advances when the set improves within the agreed boundaries.

Outcome

Define change from the recipient's point of view. “Generate summary” is an activity. “Reduce the time for an analyst to make a correctly documented decision” is an outcome hypothesis.

Choose a drive that will survive a model change. This prevents the company from confusing supplier success with flow success.

Acceptance

Accepted does not mean perfect; means adequate to the stated criteria. Criterion can combine automatic testing, policy, sampling, and human decision. Record rejection and material correction, not just final approval.

If AI creates something in two minutes that requires forty minutes to correct, the generation time tells an incomplete story.

Flow

Measure from the relevant input to the outcome, not just the automated portion. A faster generator may increase service wait times. A classifier can move the queue for exceptions. An automatic response can reduce service and increase later rework.

Local time is diagnostic. End-to-end time is a decision.

Economy

Include licensing, consumption, context preparation, integration, review, observability, incidents, and operation. In the early cycles, a range estimate is more honest than an ROI with many decimal places.

Compare with the relevant alternative: current process, change without AI, or do nothing. “Zero cost” rarely exists; the work just appears in another cost center.

Risk

Define what can't get worse: data exposure, undue promise, unauthorized decision, regulatory failure, customer harm or irreversible action. Tell incidents and near-incidents with context.

Absence of observed failure in a small sample does not prove safety. It only shows what that experiment managed to observe.

Design the baseline before turning on the AI

Without a baseline, the pilot compares enthusiasm with memory. Register the current process using the same unit and criteria that will be applied to the new flow.

A useful baseline doesn't need months of instrumentation. For a delimited stream, you can start with:

  • volume and composition of tasks;
  • accepted outcomes;
  • end-to-end and waiting time;
  • rework and escalations;
  • approximate cost of people and systems;
  • failures, exceptions and severity;
  • differences by experience, task type or channel.

Segmentation matters. The evidence cited in this article found different effects by experience level. An average can hide who received value and who shouldered the cost of change.

A pilot contract on one page

Before the first run, record:

FieldContent
Decisionexpand, redesign, restrict or terminate
Populationpeople, tasks, channels, and exceptions included
Baselinecurrent flow period, volume and metrics
Outcomeobservable final unit and recipient
Acceptanceautomatic and human criteria
Limitsquality, cost, risk and authority
Evidencedata origin, version and responsible
Windowduration and moment of decision

The contract also states what the pilot no proof. An internal test can inform operational capability without proving customer preference. A volunteer group can show membership without representing the entire company. A stable month may not cover seasonality.

Three mature decisions

Expand

The outcome improved, the acceptance criteria were preserved and cost and risk are within limits. Expansion takes place by stage, maintaining observability and reversion capacity.

Redesign

There is a sign of value, but the bottleneck has changed, the review has become expensive or a population has been harmed. The next cycle changes a specific hypothesis—process, context, tool, verification, or training—and preserves the comparison.

Close

The case does not improve the outcome, costs more than the alternative or creates an incompatible risk. Closing is not a failure of the program. It is a return on investment in learning, as long as the decision and its evidence are recorded.

The fourth state, “continue driving without deciding”, tends to be the most expensive.

What leaders should ask in the review

  • What decision should this pilot inform?
  • What was the baseline and was it measured with the same criteria?
  • Who received the outcome and who declared acceptance?
  • What rework was left out of automation?
  • The gain is concentrated in which task or population?
  • What is the total cost per accepted outcome?
  • What cannot be concluded from this sample?
  • What mechanism will change if we expand?

These questions do not require the board to choose models. They require the company to maintain the link between technology, work and results.

Conclusion

An AI pilot must buy information for a decision, not applause for a demonstration.

Producing more can be useful, but only if the flow transforms this capacity into accepted outcomes. The mature measure preserves the complete path: result, quality, time, cost and risk. It recognizes heterogeneity, explains limits and allows us to say “no” when technology does not improve that system.

Tokens remain on the operational panel. Value begins when the customer, team, or business receives a verifiable change — and someone can make the case, with evidence, why it deserves to be scaled.

Previous reading: AI amplifies the business you already have. Next reading: *Spec-driven without self-deception*, in preparation at Trustyu Forge.

Editorial and responsibility note

Research cutoff
Last review
Recorded corrections
No corrections recorded.

This article combines cited sources, analysis, and the author's professional experience. Verifiable data and factual statements are linked to their respective sources. Interpretations, hypotheses, projections, recommendations, and opinions represent the author's professional point of view at the time of publication; they do not constitute proven facts, a promise of results, or legal, financial, or technical advice applicable to a specific case. Consult the original sources and qualified professionals before making decisions.

Claims and sources

CAREER-A18-C1

In the field study with 5.172 customer support agents, access to the assistant increased the average number of problems resolved per hour by 15%; the lowest skill quintile had a gain of 36%, and the gains were concentrated among less skilled and less experienced professionals.

Limit: The result comes from customer support in a company and does not prove equal effect in every occupation, nor disappearance of input functions; 36% does not represent the entire group of beginners or less experienced. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.

CAREER-A18-C2

In three field experiments with 4.867 developers, the data set showed 26.08% more tasks completed among those who had access to the assistant, with greater adoption and gains among less experienced professionals.

Limit: Individual experiments are noisy, use a specific tool and company, and do not measure broad software quality or net employment. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.

CAREER-A18-C5

DORA 2025 describes AI primarily as amplifying existing strengths and weaknesses and associates the greatest returns with the organizational system, not just the tools.

Limit: The research focuses on software development; the article uses this formulation as a design principle, not as a universal labor law. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.