Executive reading · ~60 seconds

AI accelerates preparation, variation and analysis. Validation continues depending on observed behavior, commitment, transaction, or retention in the world the hypothesis aims to explain.

A team uses AI to create personas, mock interviews, surveys, landing pages, prototypes, copies, campaigns and analyses. In just a few days, it produces a volume that would previously have consumed weeks. The speed seems to confirm that the idea is working.

But almost everything still happened within the team's own system. The hypothesis concerned the world; the evidence came from a simulation of that world.

Executive summary

AI is an excellent experiment accelerator. It reduces preparation costs, expands variations, helps to clarify hypotheses and shortens the cycle between question and test. It does not automatically transform a produced artifact into market evidence.

Use an evidence ladder:

LevelExampleWhat really informs
Simulationpersona or interview generatedimprove questions and explore scenarios
Declarationperson says they would uselanguage, objection and perception
Behaviorperson tries to complete taskusability and contextual adherence
Commitmentprovides data, time, access or indicationrelative trust and priority
Transactionpay or sign an economic commitmentwillingness to exchange value
Retentioncome back and join the workrecurring utility in the observed context

Levels are not medals. Each hypothesis requires a type of proof. The mistake is to attribute more authority to a sign than it possesses.

The difference between producing and learning

An experiment exists to change a decision. If the result could not lead to stopping, reducing, reordering or redesigning, it is probably a demo.

Before building, write:

  1. hypothesis — what we believe about the customer, problem, channel, value or operation;
  2. greater risk — which uncertainty makes the others secondary;
  3. observation — what behavior would support or weaken the belief;
  4. decision threshold — what we will do in each result;
  5. validity — until when and for which segment the learning is valid.

This contract reduces a frequent bias: reinterpreting any result as success after the experiment ends.

Where AI enhances the quality of the experiment

Make assumptions explicit

Ask the AI to break down the thesis into customer, problem, alternative, channel, trust, price and operation. The gain is not in accepting the answer, but in making assumptions debatable.

Create cheap variations

Value propositions, scripts, screens and messages can be explored in parallel. Variation helps to avoid attachment to the first form, as long as the test compares the same hypothesis.

Prepare counterpoints

AI can act as a red team: list existing alternatives, reasons not to change, adoption costs, dependencies and failures. This improves the drawing before involving real people.

Reduce mechanical work

Transcribing, grouping observations and generating first syntheses saves time. Preserve the original records and review the classification: a convincing summary can hide important contradictions.

Where simulation turns into self-deception

Generated persona treated as customer

A model reproduces patterns from the data and the prompt. It does not bear the cost, politics, legacy, or reputation of a person in the actual situation.

Hypothetical interview treated as demand

Asking “would you buy it?” produces courtesy and imagination. Observing a decision with a trade-off produces stronger evidence.

Click treated as value

A campaign can validate message or channel without validating delivery, retention or savings. The signal is useful; the conclusion needs to remain narrow.

Impressive prototype treated as a product

Visual quality can increase interest and yet hide the integration, data, support, risk and cost that define real adoption.

A seven-day protocol

Day 1 — choose a hypothesis

Don’t test “the product”. Choose the assumption with the greatest combination of uncertainty and impact.

Day 2 — define the test

Decide which behavior, commitment, or transaction might contradict the thesis. Also record what will not be completed.

Days 3 and 4 — use AI to prepare

Create script, variations, minimum prototype and recording instruments. Remove personal data and do not expose confidential information to models without a proper contract.

Days 5 and 6 — observe the world

Conduct conversations, tasks, or offers with eligible participants. Record facts separate from interpretations.

Day 7 — decide

Compare the result to the threshold written before. Continue, reframe, restrict or stop. An “inconclusive” conclusion is legitimate when the sample or instrument does not support anything else.

Experiment cost needs denominator

Comparing just tokens or hours can reward volume. The cost must be linked to a unit of value: material hypothesis tested, decision changed, risk removed or outcome accepted. Useful observability connects telemetry and spend to this denominator, without turning an estimate into an invoice. [R3-C2]

The same discipline applies to learning. “We generated 30 variations” describes production. “We discarded segment X because none of the five eligible companies accepted commitment Y” describes decision, sample and limit.

What this article proves — and what it doesn't prove

A Harvard Business School research presents AI as an accelerator of an organization driven by hypotheses and experiments. This source offers a working model; does not validate your specific idea. FORGE uses P1 to maintain simulation, discovery, and market evidence in distinct classes.

The article does not claim that payment is the only proof, that interviews are worthless or that all experiments require inferential statistics. It states that the conclusion needs to respect the authority of the observed signal.

Conclusion

AI made the lab faster. The market remains outside of it.

Use models to prepare better, explore more options, and reduce mechanical work. Then cross the border: put the hypothesis in front of real people, constraints and choices. The advantage is not to produce more experiments. It's abandoning wrong beliefs before they become an expensive product.

Previous reading: From idea to first revenue. Next suggested reading: From pilot to value.

Editorial and responsibility note

Research cutoff
Last review
Recorded corrections
No corrections recorded.

This article combines cited sources, analysis, and the author's professional experience. Verifiable data and factual statements are linked to their respective sources. Interpretations, hypotheses, projections, recommendations, and opinions represent the author's professional point of view at the time of publication; they do not constitute proven facts, a promise of results, or legal, financial, or technical advice applicable to a specific case. Consult the original sources and qualified professionals before making decisions.

Claims and sources

R3-C2

Useful AI observability must connect technical telemetry to a unit of value and its cost, instead of optimizing aggregate spend without an outcome denominator.

Limit: A unit must be product-specific; this framework-only run measures intake cost but not customer outcome. This is a research-synthesis design input; it does not prove product adoption, operational maturity, independent attestation or outcome.

  • OpenTelemetry — OpenTelemetry, CC-BY-4.0 documentation; analysis-only excerpts
  • FinOps Foundation — FinOps Foundation, CC-BY-4.0 framework; analysis-only excerpts