Mechanism coverage
7.9 / 1038 / 48
FORGE 3.2 · September 6, 2026 review
After Brownfield, evolution extends to durable specs, multi-agent work, verifiable governance and pipeline integrity. The next step is proving these controls reach every consumer and protect delivery.
Reconcile capabilities, data and risks before deciding what to preserve, consolidate or rebuild.
Persistent specs, per-machine seats, organization audits, bot PRs and stricter verifiers.
Measure the executed body, close per-consumer gates and exercise deployment, recovery and telemetry within data boundaries.
Baseline review that informed FORGE 3.3; it is not its qualification. Integrated Brownfield recommendation remains suspended; completed review does not mean approval.
Editorial self-assessment: 12 equally weighted dimensions and two independent scales. Ratings are not certification, market percentiles or completion percentages.
38 / 48
20 / 48
Coverage: 0 no material located; 1 intent/proposal; 2 documented contract; 3 identified implementation; 4 implementation and tests identified in scope. A high rating does not prove universal deployment.
Evidence: 0 no eligible proof located; 1 local, synthetic or partial; 2 bounded remote execution; 3 bounded chain with recorded cross-component validation; 4 longitudinal measurement and independent effectiveness verification. Uncertainty or lack of access limits the rating; it does not prove absence.
Each index = rating sum ÷ 48 × 10, rounded to one decimal. The 12 assessment dimensions reorganize the method; they do not replace the 23 editorial axes in the August inventory.
Five core repositories, 88 workflows and 215 external Actions references: 215 pinned to full SHA values; another 25 references are local. This is a syntax inspection of pinned files, not an organization census or proof of execution.
Internal sources were checked through commits, ADRs, issues and runs. Full Claude Code conversations were not inspected. The internal basis has restricted access; this public summary alone cannot reproduce every conclusion. Market sources are primary and accessible.
Maps, generated instructions, counts and link checks reduce documentation drift.
Verifiers and bounded remote runs exist; context coverage and quality impact lack a longitudinal series.
Restricted internal basis: Docs · ADR-0074, canon-check, PRs 698/699
Comparison reference: OpenAI · Harness engineering · DORA · AI Capabilities Report 2025
Durable specs, traceability and empty-corpus rejection strengthen change evaluation.
One SDD run was observed as ineligible-shadow. Guard propagation and real agent outcomes still need per-consumer evidence.
Restricted internal basis: Docs · ADR-0074; Infra · PR 625; SCORE · SDD119
Comparison reference: Anthropic · Demystifying evals for AI agents · OpenAI · Harness engineering · NIST · SP 800-218A
Canonical worktrees, per-machine seats and fleet-doctor address collisions and environment drift.
Directory isolation is not an operating-system sandbox. Longitudinal collision and recovery metrics are missing.
Restricted internal basis: Docs · ADR-0076/0079; AI Env · PR 187
Comparison reference: OpenAI · Harness engineering · NIST · SSDF 1.1
Organization invariants, repository matrices and cached audits distinguish drift, exceptions and denied reads.
372 local tests passed. The remote observation still contains warnings and partial reads: a green workflow is not complete conformance.
Restricted internal basis: Docs · ADR-0081; run 34055858909; local suite Sep 6
Comparison reference: NIST · SSDF 1.1 · OpenSSF · Scorecard checks
Cancellations, stale verdicts and scanners with no input received explicit failure handling.
The invoked body, required checks and delivered SHA must converge. Complete propagation was not demonstrated in this review.
Restricted internal basis: Infra · PRs 597/599/600/617; Docs · proposed ADR-0082
Comparison reference: GitHub · Secure use of Actions · OpenSSF · Scorecard checks
Bot-authored PRs and release PRs separate authorship, review and publication. Skipped review no longer counts as review.
Review activity does not prove independent substantive approval. Exceptions and gate coverage still limit assurance.
Restricted internal basis: Infra · PRs 620/631; Template · PR 322; proposed ADR-0080
Comparison reference: OpenSSF · Scorecard checks · NIST · SSDF 1.1
Baseline 0.20.1 has recorded verification; 0.21.0 was published with native immutability metadata. External Actions references are pinned in the inspected scope.
The evidence rating includes the bounded historical chain, not a new 0.21.0 signature verification. Pinning and signatures do not award a SLSA level or prove fleet adoption.
Restricted internal basis: Release 0.20.1 · public record; Infra · release 0.21.0
Comparison reference: SLSA · Specification 1.2 · GitHub · Secure use of Actions
Characterization, reconciliation, target decisions, reversible waves and retirement have contracts and adversarial tests.
Independent review took place, but integrated approval remains withheld. No real pilot or customer migration is claimed.
Restricted internal basis: Docs · ADR-0075; qualification v3; issues 563/584
Comparison reference: NIST · SSDF 1.1 · DORA · AI Capabilities Report 2025
Tool policies, provenance and permission boundaries are part of the method; hardened images now have a dedicated standard.
Auto mode is not a sandbox. End-to-end adversarial enforcement and image rollout were not demonstrated globally.
Restricted internal basis: Docs · ADR-0047/0049/0078; hardened-base-images
Comparison reference: OWASP · Agent Control Standard · MCP · Authorization security considerations · NIST · SP 800-218A
The historical baseline retains traces and receipts. The new telemetry proposal requires data classification and a perimeter decision before installation.
ADR-0083 is still proposed. This review has no proof of integrated longitudinal telemetry or automatic authorization to collect prompts.
Restricted internal basis: Docs · ADR-0054; proposed ADR-0083
Comparison reference: Anthropic · Demystifying evals for AI agents · NIST · SP 800-218A · MCP · Authorization security considerations
Budgets, fewer calls, conditional caching and reusable workflows address infrastructure waste.
Fewer calls are not business ROI. Cost per accepted change, rework and quality need a shared denominator.
Restricted internal basis: Docs · ADR-0027; cached audit; Infra · reusable workflows
Comparison reference: DORA · AI Capabilities Report 2025 · DORA · Software delivery performance metrics
Readiness, deployment evidence and rollback have contracts; recent review exposed gaps between documentation and effective automation.
Merge, artifact and deployment must be linked, and recovery exercised. The five DORA metrics were not measured per service here.
Restricted internal basis: Docs · CI/CD audit 705/707; documentation correction 725 merged, without automatic rollback
Comparison reference: DORA · Software delivery performance metrics · NIST · SSDF 1.1
The method aligns with harness engineering, secure development, provenance and outcome evals. The main gap is demonstrated effectiveness: genuinely required controls, exercised recovery, and quality and cost per accepted change. Without a comparable sample of other organizations, there is no basis for claiming world-leading performance.
DORA uses five delivery metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. Compare per application or service; agent activity and commit volume do not replace these outcomes. DORA
Bind source, artifact, checks and deployment; test that a relevant failure blocks delivery. Rehearse rollback and record recovery time and state.
Measure capabilities by executed content, not pin labels. Define owned, expiring exceptions and demonstrate independent review by risk class.
Resolve Brownfield's remaining human decisions, revalidate the convergent chain and record the decision. Then run only the authorized pilot and measure all five DORA metrics, cost and quality within an approved data perimeter.
References inform the criteria; they do not certify FORGE. No SLSA level was assessed, OpenSSF Scorecard was not run and the new OWASP ACS was not deployed in this review.