Executive reading · ~60 seconds
In 90 historical production bugs from Google, an agent who wrote contracts before generating tests detected 63.2% of defects in five attempts, compared to 53.4% in the baseline. The 9.8 percentage point gain came with 38% more tokens. The study supports contracts as a useful stage of reasoning and control, not as a substitute for coverage, review or human decision.
An agent can generate many tests and still not touch the property that distinguishes correct behavior from a defect. The problem isn't just writing test code. It is to make explicit the contract that the system should preserve before choosing inputs and oracles.
Status: Evidence Brief approved for publication. This text analyzes a study with internal historical bugs of Google. It does not claim that Trustyu Forge reproduced the experiments nor that the results automatically transfer to another stack, model or organization.
The experiment
The paper *Grounding AI Agents in Contracts* evaluated 90 historical production bugs, all with fixes in a single file, in C++, Java, Python and Go. The same model — Gemini 3 Flash — received up to five attempts per bug, totaling 450 executions per approach. At baseline, the agent generated tests directly. In the spec-driven condition, I first wrote preconditions, postconditions and test suggestions in natural language; then used this contract to produce the suite.
The optional human curation proposed by the architecture was deliberately removed from the experiment. This allows you to measure the contribution of the specification stage, but does not measure a complete engineering flow with experts reviewing the contract.
The gain, with denominator
In five attempts, the contract-driven approach detected 63.2% of the 90 bugs, compared to 53.4% in the baseline — a difference of 9.8 percentage points, reported with p=0.0352. There were 57 bugs found by the spec-driven approach and 48 by the baseline: 45 in common, 12 exclusive to the specification and three exclusive to the baseline.
Branch coverage increased from 46.4% to 48.9%, a difference of 2.5 points; the variation in line coverage was -0.4 points and was not statistically significant. This matters because more detection didn't simply come from running many more lines. The contract appears to have guided what behaviors to look for.
When the suite covered the contract produced, detection occurred in 54.9% of attempts, 151 of 275. Without contract coverage, it occurred in 19.4%, 34 of 175. It is a strong association within the experiment, not universal causal proof.
The cost that should not be hidden
The baseline consumed 243.9 million tokens. The spec-driven approach consumed 336.7 million: 38% more. Entry grew by 36.2% and exit by 59.1%. Due to the exclusive bug found, the estimated cost increased from 5.1 million to 5.9 million tokens, an increase of 16.2%.
This number changes the operational decision. Better detection may justify more inference on critical software, but not on any change. The policy needs to define when the additional contract is mandatory, when a simpler template is sufficient and when the cost is not payable.
How to turn the find into a gate
- Declare ownership: record preconditions, postconditions, invariants, and prohibited behaviors before ordering tests.
- Separate implementation contract: the oracle must observe behavior, not copy the known fix.
- Measure contract coverage: line and branch coverage remains useful, but does not alone inform whether the critical property has been exercised.
- Preserve divergence: If direct testing and contract-driven testing find different sets of bugs, use diversity rather than choosing a single agent.
- Budget the inference: record tokens, latency and cost per additional defect detected.
- Link to artifact: contract, suite, result, model, prompt and commit need to form a reproducible track.
- Maintain human decision: a failed test is evidence to investigate; a passing test does not certify the entire system.
What the study doesn't prove
The sample comes from a single monorepo and an internal Google process. Bugs are historical, the fix touches one file, and the experiment uses a family of models. Distributed systems, configuration failures, identity, concurrency, dependencies, and emerging incidents may respond differently.
The suite quality assessment also used Gemini 3.1 Pro as a judge. The study anonymized and randomized the order and repeated the trial five times, but an LLM rater may carry noise, style preference, or affinity with the same technology family. The preference of 77.8% over the baseline and 56.7% over human tests, among 83 successfully generated tests, does not equate to general superiority over experts.
All authors are linked to Google. This does not invalidate the work, but it makes affiliation, access to the internal dataset and lack of independent reproduction relevant to the reading. Until the cutoff of 10/04/2026, we did not find an external replication of the complete experiment.
Architecture rule
Specification is not documentation written later. It is a verifiable intermediary between intent and execution. When an agent needs to declare the contract before testing, the team gains an object that they can review, version, compare and use to explain why that suite should detect a certain class of error. [HARNESS31-C2]
The useful result is not ‘the AI tested’. It is: what ownership was declared, what evidence exercised it, what cost was consumed, what remained out of scope, and who authorized the residual risk.
Direct sources
- Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation — arXiv. Preprint, method, sample, tables and limitations.
- DOI ACM of the article. Persistent record of the publication.
Editorial and responsibility note
This text combines facts attributed to the sources with the author's analysis and technical proposal. Personal and professional views are not proven facts; data, denominators, limits and conflicts are indicated when available. The content is informative and does not replace technical, legal, financial or security assessment. Tech Human and Trustyu work commercially on related topics. Research, structure and writing were assisted by AI; Factual review, authorial approval, and publication were confirmed by Fernando Parreiras on 10/4/2026, without additional independent human review.
Research cutoff: 10/04/2026. Editorial status: Special Evidence Brief authorized for 10/6/2026 at 9am BRT.
Editorial and responsibility note
- Research cutoff
- Last review
- Recorded corrections
- No corrections recorded.
The cutoff above applies to the canonical claims. Additional sources and their access dates are identified in the article body.
This article combines cited sources, analysis, and the author's professional experience. Verifiable data and factual statements are linked to their respective sources. Interpretations, hypotheses, projections, recommendations, and opinions represent the author's professional point of view at the time of publication; they do not constitute proven facts, a promise of results, or legal, financial, or technical advice applicable to a specific case. Consult the original sources and qualified professionals before making decisions.
Claims and sources
HARNESS31-C2
A durable harness should maintain session history, orchestration policy and execution isolation as explicit boundaries, while adapters and plugins prevent model or channel choice from becoming the security boundary.
Limit: The sources show two implementations, not a neutral interoperability standard. Actual permissions must be independently enforced and tested outside model choices. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.
- Anthropic — Anthropic, proprietary-site-terms
- YC Software —YC Software, MIT
- YC Software —YC Software, MIT