Trustyu

FORGE 3.2 · September 6, 2026 review

From a documented method to verifying what actually runs.

After Brownfield, evolution extends to durable specs, multi-agent work, verifiable governance and pipeline integrity. The next step is proving these controls reach every consumer and protect delivery.

  1. 3.2 baseline · August

    Understand and transform legacy systems

    Reconcile capabilities, data and risks before deciding what to preserve, consolidate or rebuild.

  2. Integrated advances · September

    Govern how work happens

    Persistent specs, per-machine seats, organization audits, bot PRs and stricter verifiers.

  3. Open frontier · not complete

    Prove adoption and outcomes

    Measure the executed body, close per-consumer gates and exercise deployment, recovery and telemetry within data boundaries.

Baseline review that informed FORGE 3.3; it is not its qualification. Integrated Brownfield recommendation remains suspended; completed review does not mean approval.

Ratings with explicit criteria

Editorial self-assessment: 12 equally weighted dimensions and two independent scales. Ratings are not certification, market percentiles or completion percentages.

Mechanism coverage

7.9 / 10

38 / 48

Adoption evidence

4.2 / 10

20 / 48

Coverage: 0 no material located; 1 intent/proposal; 2 documented contract; 3 identified implementation; 4 implementation and tests identified in scope. A high rating does not prove universal deployment.

Evidence: 0 no eligible proof located; 1 local, synthetic or partial; 2 bounded remote execution; 3 bounded chain with recorded cross-component validation; 4 longitudinal measurement and independent effectiveness verification. Uncertainty or lack of access limits the rating; it does not prove absence.

Each index = rating sum ÷ 48 × 10, rounded to one decimal. The 12 assessment dimensions reorganize the method; they do not replace the 23 editorial axes in the August inventory.

What was inspected

Five core repositories, 88 workflows and 215 external Actions references: 215 pinned to full SHA values; another 25 references are local. This is a syntax inspection of pinned files, not an organization census or proof of execution.

Internal sources were checked through commits, ADRs, issues and runs. Full Claude Code conversations were not inspected. The internal basis has restricted access; this public summary alone cannot reproduce every conclusion. Market sources are primary and accessible.

12 dimensions, including what is missing

Context and canon

Mechanism coverage
4 / 4
Adoption evidence
2 / 4

What evolved

Maps, generated instructions, counts and link checks reduce documentation drift.

What the evidence does not yet support

Verifiers and bounded remote runs exist; context coverage and quality impact lack a longitudinal series.

Restricted internal basis: Docs · ADR-0074, canon-check, PRs 698/699

Comparison reference: OpenAI · Harness engineering · DORA · AI Capabilities Report 2025

Specifications and evals

Mechanism coverage
3 / 4
Adoption evidence
2 / 4

What evolved

Durable specs, traceability and empty-corpus rejection strengthen change evaluation.

What the evidence does not yet support

One SDD run was observed as ineligible-shadow. Guard propagation and real agent outcomes still need per-consumer evidence.

Restricted internal basis: Docs · ADR-0074; Infra · PR 625; SCORE · SDD119

Comparison reference: Anthropic · Demystifying evals for AI agents · OpenAI · Harness engineering · NIST · SP 800-218A

Multi-agent work

Mechanism coverage
4 / 4
Adoption evidence
2 / 4

What evolved

Canonical worktrees, per-machine seats and fleet-doctor address collisions and environment drift.

What the evidence does not yet support

Directory isolation is not an operating-system sandbox. Longitudinal collision and recovery metrics are missing.

Restricted internal basis: Docs · ADR-0076/0079; AI Env · PR 187

Comparison reference: OpenAI · Harness engineering · NIST · SSDF 1.1

Verifiable governance

Mechanism coverage
4 / 4
Adoption evidence
2 / 4

What evolved

Organization invariants, repository matrices and cached audits distinguish drift, exceptions and denied reads.

What the evidence does not yet support

372 local tests passed. The remote observation still contains warnings and partial reads: a green workflow is not complete conformance.

Restricted internal basis: Docs · ADR-0081; run 34055858909; local suite Sep 6

Comparison reference: NIST · SSDF 1.1 · OpenSSF · Scorecard checks

CI integrity

Mechanism coverage
3 / 4
Adoption evidence
2 / 4

What evolved

Cancellations, stale verdicts and scanners with no input received explicit failure handling.

What the evidence does not yet support

The invoked body, required checks and delivered SHA must converge. Complete propagation was not demonstrated in this review.

Restricted internal basis: Infra · PRs 597/599/600/617; Docs · proposed ADR-0082

Comparison reference: GitHub · Secure use of Actions · OpenSSF · Scorecard checks

Review and authority

Mechanism coverage
3 / 4
Adoption evidence
1 / 4

What evolved

Bot-authored PRs and release PRs separate authorship, review and publication. Skipped review no longer counts as review.

What the evidence does not yet support

Review activity does not prove independent substantive approval. Exceptions and gate coverage still limit assurance.

Restricted internal basis: Infra · PRs 620/631; Template · PR 322; proposed ADR-0080

Comparison reference: OpenSSF · Scorecard checks · NIST · SSDF 1.1

Software supply chain

Mechanism coverage
3 / 4
Adoption evidence
3 / 4

What evolved

Baseline 0.20.1 has recorded verification; 0.21.0 was published with native immutability metadata. External Actions references are pinned in the inspected scope.

What the evidence does not yet support

The evidence rating includes the bounded historical chain, not a new 0.21.0 signature verification. Pinning and signatures do not award a SLSA level or prove fleet adoption.

Restricted internal basis: Release 0.20.1 · public record; Infra · release 0.21.0

Comparison reference: SLSA · Specification 1.2 · GitHub · Secure use of Actions

Brownfield and consolidation

Mechanism coverage
4 / 4
Adoption evidence
1 / 4

What evolved

Characterization, reconciliation, target decisions, reversible waves and retirement have contracts and adversarial tests.

What the evidence does not yet support

Independent review took place, but integrated approval remains withheld. No real pilot or customer migration is claimed.

Restricted internal basis: Docs · ADR-0075; qualification v3; issues 563/584

Comparison reference: NIST · SSDF 1.1 · DORA · AI Capabilities Report 2025

Agent security

Mechanism coverage
3 / 4
Adoption evidence
1 / 4

What evolved

Tool policies, provenance and permission boundaries are part of the method; hardened images now have a dedicated standard.

What the evidence does not yet support

Auto mode is not a sandbox. End-to-end adversarial enforcement and image rollout were not demonstrated globally.

Restricted internal basis: Docs · ADR-0047/0049/0078; hardened-base-images

Comparison reference: OWASP · Agent Control Standard · MCP · Authorization security considerations · NIST · SP 800-218A

Observability and data

Mechanism coverage
2 / 4
Adoption evidence
1 / 4

What evolved

The historical baseline retains traces and receipts. The new telemetry proposal requires data classification and a perimeter decision before installation.

What the evidence does not yet support

ADR-0083 is still proposed. This review has no proof of integrated longitudinal telemetry or automatic authorization to collect prompts.

Restricted internal basis: Docs · ADR-0054; proposed ADR-0083

Comparison reference: Anthropic · Demystifying evals for AI agents · NIST · SP 800-218A · MCP · Authorization security considerations

FinOps and efficiency

Mechanism coverage
3 / 4
Adoption evidence
2 / 4

What evolved

Budgets, fewer calls, conditional caching and reusable workflows address infrastructure waste.

What the evidence does not yet support

Fewer calls are not business ROI. Cost per accepted change, rework and quality need a shared denominator.

Restricted internal basis: Docs · ADR-0027; cached audit; Infra · reusable workflows

Comparison reference: DORA · AI Capabilities Report 2025 · DORA · Software delivery performance metrics

Delivery and recovery

Mechanism coverage
2 / 4
Adoption evidence
1 / 4

What evolved

Readiness, deployment evidence and rollback have contracts; recent review exposed gaps between documentation and effective automation.

What the evidence does not yet support

Merge, artifact and deployment must be linked, and recovery exercised. The five DORA metrics were not measured per service here.

Restricted internal basis: Docs · CI/CD audit 705/707; documentation correction 725 merged, without automatic rollback

Comparison reference: DORA · Software delivery performance metrics · NIST · SSDF 1.1

Market alignment, not a leadership badge

The method aligns with harness engineering, secure development, provenance and outcome evals. The main gap is demonstrated effectiveness: genuinely required controls, exercised recovery, and quality and cost per accepted change. Without a comparable sample of other organizations, there is no basis for claiming world-leading performance.

DORA uses five delivery metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. Compare per application or service; agent activity and commit volume do not replace these outcomes. DORA

Next steps, with exit evidence

  1. Priority 1 · close the delivery chain

    Bind source, artifact, checks and deployment; test that a relevant failure blocks delivery. Rehearse rollback and record recovery time and state.

  2. Priority 2 · verify the fleet and review

    Measure capabilities by executed content, not pin labels. Define owned, expiring exceptions and demonstrate independent review by risk class.

  3. Priority 3 · complete qualification and measure outcomes

    Resolve Brownfield's remaining human decisions, revalidate the convergent chain and record the decision. Then run only the authorized pilot and measure all five DORA metrics, cost and quality within an approved data perimeter.