Trustyu Forge Technical area
Technical benchmark v1.9 · source-bound cut Aug 13, 2026

Current status
of Forge 3.1.

This is the engineering layer behind the executive narrative. Seventeen capabilities are assessed by the lowest sustainable level of the system, with evidence linked to release, repository and SHA. The historical 2.0–3.0 cutoff remains below without being rewritten.

Specified Shadow 1 shadow-operational profile Fleet Enforced · not eligible No claim Operational / Attested
0.18signed and verifiable runtime
0.16Template default in shadow
21 / 21structured profile campaign
34 / 60executed CRM contracts
49 / 49classified Hub surfaces
NoFleet Enforced/GA declared
Core 3.1 Verifiable

Contracts, SDD, Graph/Loop, coverage, API/events, receipts, supply chain and rollback are distributed across the runtime and the Template.

Profile supported Shadow operational

OpenAI Responses readonly in Loop proven broker, token envelopes, causal negatives, rollback, re-entry and teardown.

terminal release Not closed

Independent Observer, CRM/Hub compliance, release gates, and operational window still block Fleet Enforced.

Public contract cutting: machine-readable manifest, in which each source declares whether the reader can open it: private repository references are marked as restricted, because GitHub responds 404 to those who don't have access and that is indistinguishable from non-existent. The shadow-operational profile is a supported and constrained surface; does not automatically promote runtime, products or fleet to Operational, Attested or Fleet Enforced.
Reading contract

Six states.
No shortcuts.

Status follows provenance. A demonstrated mechanism does not automatically qualify the Template; a Template default does not qualify a product; a pilot's observation does not become universal property.

Mechanismartifact or control demonstrated in the cited section; does not prove adoption.
Template defaultdefault issued by the scaffold; The product still requires its own evidence.
Pilot observationmeasurement with repo, SHA and window; It is not a universal property.
Targetdesired objective or threshold; is not a result achieved.
Pendingno current eligible evidence; claim delivered is prohibited.
Historicalfrozen previous cut; does not represent current qualification.
Current cut · 17 capacities

Forge 3.1 now.
Evidence and residual.

This matrix does not score the presence of files. It uses the lowest demonstrated maturity between contract, execution, independent verification and adoption. V means verified in the mentioned scope; does not mean Fleet Enforced.

Maturity Ddocumented Ccontrolled Vchecked Aattested · not declared Ooperational · not declared
AxleForge 3.1 · 13 AugCurrent evidenceResidual for the terminal state
Context engineering
VMechanized context
Canon-check, progressive disclosure, profiles and source-bound traceability.Reconcile freshness between program, issues and live status.
ADR and alive status
CControlled contract
ADRs 0060, 0062 and 0063 accepted; specified shadow topology.ADR-0063 was accepted on 12/08: the G6 criterion exists, but the ≥30 day window has not yet started.
Executable controls
VVerifiable runtime
runtime 0.18 validates manifests, receipts, bindings, Graph/Loop and causal envelopes.Enforcement remains off until fleet qualification.
Spec-driven development
VProportional trace
docs#418 incorporated SPEC → change → test → evidence into gates.Source closeout completed; products and operational profile still need to satisfy the terminal gates.
Writing isolation
VSingle writer
Worktree, D5 lock, overlap guards, ownership and teardown executable.Effectiveness depends on continued adherence by each consumer.
Sandbox and least privilege
CRestricted profile
Deny-by-default, pre-dispatch broker, pinned network and transport; adversarial suite of 29 tests covers the four vectors of the CONNECT-only proxy.Unpublished OCI image; the current E2E runs in SHA previous at the border merge and needs to be repeated. Pending independent observer.
Tool/MCP governance
CReadonly qualified
Grant, allowlist, broker and tool loop have been proven for OpenAI Responses readonly.Hosted built-ins, broad MCP, nested agents, and Claude fall outside the supported 3.1 profile.
AI evals
VTrace-native
hub#765 closed trace-native quality and context integrity.hub#791: execution, limits and checkpoint still insufficient.
Evidence and attestation
CSigned, not attested
Assets, provenance, receipts and verifiers external to the producer in runtime 0.18; publication of the image by digest with SBOM and provenance delivered.Observer OCI and successor ARR are still missing; Attested is not declared.
supply chain
VPin and rollback
SHA256SUMS, Ed25519, provenance and immutable pins on Template 0.16.1, with approved rollback to 0.15.The release no longer depended on human credentials: it started to be published by machine identity trustyu-forge-bot, with ephemeral token. It remains to observe durability in a longitudinal window.
Readiness, SLO and rollback
CShadow operational
21/21 campaign, causal negatives, rollback, re-entry and profile teardown supported.A ADR-0063 accepts fixed criteria; window, SLI/error budget, game-day and human promotion remain mandatory.
Closed-loop learning
VDurable regressions
Findings became invariants, schemas, negatives and tests for the 0.18 runtime.Measure avoided recurrence and effectiveness, not record volume.
FinOps and unit economics
CLimited Tokens
Causal envelopes and token hard stops per turn; fleet CI Consolidation Unlocked ~3,200 min/month when removing jobs billed by rounding (M16 measurement).Monetary hard stop absent and post-wave confirmation not yet measured — projected savings are not observed savings.
Senior roles and human control
CExplicit authority
Profiles, Graph/Loop, promotions and approvals are not decided by the model; separating agent and human identity unlocked the first positive of the review gate.Measured: the automated reviewer did not review 57,3% 699 PRs in 30 days; actual gate coverage = 37,3%. Check turns green without reviewing. infra#521 · #522 · #523
Agentic observability
CMetadata-only
Use, sequence, digests and receipts without persisting prompt, output or tenant.ai-env#139: independent observation remains pending.
Issue provenance/closeout
CRebuildable
Feature specs, PRs, source SHAs, artifacts and residuals form an auditable chain.EPIC 3.1 remains open until G1–G6: release truth and half of the review already proven, but the gate remains in shadow in the four repositories.
Adoption and assurance
CShadow fleet
Template 0.16.1, CRM 1.141.0 and Hub 0.72.1 keep the 0.18 runtime in shadow.CRM 34/60 · Partial API Hub; Fleet Enforced ineligible.
Definition of Done terminal

The final point
from 3.1.

The program's nine technical flows — excluding Articles, conducted separately — reached engine or shadow status. This does not end the release. To honor the name GA — Fleet Enforced, six qualification gates remain mandatory.

  1. G1
    Fail-closed release truth

    Release proven end-to-end and published by machine identity; the review half unlocked on 08/13 with the first gate positive — Human Approve enabled after separating the identity of the agent from that of the author. The gate continues in shadow in the four repositories and promotion to required remains blocked.

    infra#486 · #508 · #521 · #522
    Partial · proven
  2. G2
    Boundary OCI and independent observer

    Merged border, CONNECT-only proxy adversarial suite delivered (29 tests, four vectors) and digest publishing with SBOM/provenance ready. The published image, the successor receipt/ARR and the external observer are missing — and the current protected E2E runs on pre-merge SHA, so it needs to be repeated.

    ai-env#139 · PR #140 · #147 · #157
    Ongoing
  3. G3
    Supported profile hard stops

    Tokens, deadline and cost verified outside the producer; out-of-profile modes denied.

    ai-env#72 · ai-env#81
    Pending
  4. G4
    CRM with full contract

    60/60: authorized fixtures and two-tenant boundary in ephemeral PostgreSQL, without real data.

    crm#1840
    34 / 60
  5. G5
    Hub with sufficient API/events and Graph/Loop

    OAS lint/diff/conformance, frame schemas, denominators and observed execution/limits/checkpoint.

    hub#764 · hub#791
    Partial
  6. G6
    Operational qualification and human promotion

    Protected E2E, campaign, rollback/game-day, minimum 30-day window, SLI/error budget, and recorded decision. The ADR-0063 was accepted on 08/12, then the criterion exists; the window has not yet started because it depends on the frozen candidate.

    ADR-0063 · EPIC #366
    Pending
Don't block 3.1 with new scope: Claude full parity, hosted built-ins, wide MCP, nested agents, independent external adopter and level Attested are explicit candidates for 3.2. In 3.1, these surfaces must remain outside the supported profile and fail closed.
Historical reading

From the documentary foundation to
verifiable system.

Evolution was not a name change. Each milestone increased the ability to transform intent into reliable execution. The canonical study froze quantitative comparisons as of 2.0; the initial phase is preserved as historical origin, without manufacturing a 1.0 column that never had a comparable formal baseline.

Origin · Forge 1.x · Apr–Jun 2026

Method and foundation

P1–P6 Templates, ADRs, architectural principles, squads and first real products formed the basis of what would become the framework.

Qualitative historical record · not scored in frozen benchmark.
Forge 2.0 · Jun 21, 2026

Squad AI-first

Skills, worktrees, spec-loop, review-remediate, hooks, MCP, security and FinOps made the process repeatable.

Formal baseline used in this matrix.
Forge 2.1 · transition Jul 11, 2026

Context and controls

Context budget, control catalog, eval baseline, readiness/SLO and squad activation by risk and phase.

Analytical intermediate state; there was no separate public release.
Forge 3.0 · release 11 Jul 2026

engineering harness

Policy, sandbox, verifier, evidence, observer and closed-loop learning connect written rules to provable control.

Historic decision: framework-only DONE on Jul 19; without operational qualification.
Historical cut decision: the reconciliation of the 24 items recorded framework-only as DONE and FORGE 3.0 as GO technical on July 19, 2026. The runtime forge-v0.8.0 was published in source SHA eae88cfe39c49e829b4ec4b624887403ff15bbb4, with Template 0.8 and rollback 0.7. A synthetic reference reproduced Evidence Pack 6/6 and conformance Verified offline; it proves the mechanism, not current state, product, adopter, or operation. Operational and Attested are not declared.
17 capabilities

The market and what existed
up to cut 3.0.

The unit of analysis is engineering capability — not files, rules, PRs, or lines of code. This v1.7 frame remains frozen to preserve the 2.0–3.0 comparison. It does not represent the current state; cut 3.1 is in the previous matrix.

Maturity absent Ddocumented Ccontrolled Vchecked Acertificate Ooperating
AxleMarket consensus 2025–2026Forge 2.0Forge 2.1Forge 3.0 cut · Jul 17Historic verdict
Context engineeringShort map, versioned knowledge, progressive disclosure and mechanical freshness.
CStrong but extensive ADRs, skills and instructions.
VBudgets, fragments, history and capability register.
VCanon-check and progressive disclosure mechanics.
StrongWhat remains is semantic drift between anchors and alive state.
ADR and alive statusPersistent/superseded decision; separate current operating state.
DVersioned decisions, status with residual drift.
CGovernance and living state catalog.
C0039/0040 reconciled and submitted for ratification.
AlignedSemantic freshness still requires human review.
Executable controlsOwner, applicability, evidence, enforcement, exception and deadline.
DRules distributed between ADRs and CI.
CDeterministic catalog and profiles.
V14 controls, diff facts and explicit evidence.
AboveIt remains to prove coverage in every applicable real change.
Spec-driven developmentSpec as source, out-of-scope, contracts before code and trace until acceptance.
CLoop SPEC→PR and Contract-First.
CStandardized feature spec and gate.
VTrace and independent verifiers.
StrongMeasure defects and rework avoided, not using the template.
Writing isolationA task/workspace/branch; parallel reading; single-writer per artifact.
CWorktree by feature and D5/D6 protocol.
CExplicit seats and roster.
VOverlap, ownership and teardown guards.
StrongRecord teardown on every real task.
Sandbox and least privilegeFS, network, secrets and tools per task; deny-by-default; short credential; HITL.
DGuardrails and worktree insulation.
CRisk profiles and capabilities.
CHub #683 and CRM #1621 passed provenance without passing the aggregate; #688 failed.
PartialLack of approved global gate, approvals/capabilities and longitudinal coverage.
Tool/MCP governanceMinimal tools, clear schemas, scopes, allowlist and auditability.
CCanonical MCP and inventory per repo.
CActivation by context and phase.
VCapability routing and validated contracts.
AlignedTest tool profile per task and control context cost.
AI evalsModel + harness + environment; isolated trials; calibrated graders; cost and latency.
DEval-harness in part of the skills.
CAI change baseline and gates.
V12/12 technical corpus with source/profile/SHA binding.
StrongScale up to 20–50 real failures and measure cost per task.
Evidence and attestationExecutor does not self-attest; source/SHA/profile binding; verifiable artifact.
DHeterogeneous evidence on CI and PR.
CEvidence contracts and defined provenance.
CSHA256SUMS manifest for runtime 0.8 with Ed25519 signature; synthetic fixture 6/6 Verified; Observer EVD=2/5 and VER=2/9.
PartialReal pack, coverage ≥95%, approved global gate and longitudinal window are missing.
supply chainImmutable lock/pin, SBOM, provenance, trusted builder and verification.
CLocks, scans and actions partially pinned.
CBaseline and evidence collectors.
VSHA256SUMS 0.8 manifest with Ed25519 signature and 9 inventoried assets; Template pinned with rollback 0.7.
StrongDo not declare SLSA level without all requirements met.
Readiness, SLO and rollbackSLI/SLO, error budget, tested rollback, runbook and stop condition.
DMostly manual readiness.
CSLO, scorecard and gates defined.
CObserver signed: 11 changes in 2 days; open window.
PartialRequires ≥20 changes, 30 days, stability and pilot drills.
Closed-loop learningReview or incident becomes test, eval, control or durable improvement.
CReview-remediate and versioned history.
CLearning records and ownership.
VFeedback produces tested controls, evals and regressions.
AlignedMeasure avoided recurrence and effectiveness, not number of records.
FinOps and unit economicsCost per unit of value/outcome, human attention and denial-of-wallet.
CDeployment skip and cost per product.
CBudgets by branch, agent and gate.
CM15: Actions −47.20%/day; minutes per change +28.27%.
Up in governanceMissed target by 2.80 p.p.; unit efficiency regressed; missing models and human time.
Senior roles and human controlSenior defines intention, risk and taste; proportional autonomy and clear approvals.
DPersonas and squad roles.
CActivation by risk/phase and human gates.
CExplicit owner and decision; single technical approver.
Aligned with caveatFounder/CTO is today the only real technical approver.
Agentic observabilityModel/tool ​​spans, outcome, tokens, cost, policy and privacy.
DTelemetry broken down by product.
CSchema of receipts and harness metrics.
CSubset OTel GenAI privacy-safe in runtime 0.8, covered by the signed SHA256SUMS manifest.
BackReal product telemetry linked to tokens/cost/outcome/receipt is missing.
Issue provenance/closeoutAuthorship, decision, implementation, validation and reconstructable residual.
DInconsistent closeouts and unproven records.
DProblem recognized, still no common gate.
CADR-0051 and closeout gate in canonical.
Above in the contractPartial in the fleet; Spread will occur by normal touch.
External adoption and assuranceIndependent user, package/license, support, upgrade/rollback and evidence pack.
No external framework product.
DDocumented extraction strategy.
DConformance/EEP inventoried by SHA256SUMS manifest signed in 0.8; no external adopters.
BackNo independent external adopters until cut off.

Where Forge 3.0 is strong

  • Integrated context engineering, executable catalog and spec-driven development.
  • Task isolation, evals, supply chain and source-bound verification.
  • FinOps of CI/agents and more rigorous issue closeout in the design.

What prevents the next level

  • Approved global gate and high-risk/≥95% coverage; today EVD=2/5 and VER=2/9.
  • ≥20 changes, 30 days, 2.80 p.p. residual FinOps and cost of models/human attention per outcome.
  • True adoption of the OTel GenAI profile, assurance and first external adopter.
Source-bound research

The market was researched,
not just cited.

The public gate was performed against peer-reviewed snapshots from official, primary, and regulatory sources. Each acquisition generated a receipt linked to the URL, snapshot, content hash, policy, retention and measured cost. External content was treated as untrusted data and was not persisted.

20 / 20approved sources and receipts
11independent groups
3 / 3multi-source studies
9 / 9claims mapped and reviewed
0model calls and tokens
0persisted raw content
R1 Reliable Engineering Harness

Reliability is a property of the system

Workflow, workspace cycle, isolation, small changes, feedback and durable decisions need to be explicit. A better prompt does not replace harness.

R2 · Secure and auditable delivery

Controls complement each other

Reproducible testing, least privilege, sandbox, provenance and independent verification are different gates. None of them alone prove the system.

R3 · Evals, observability and FinOps

Quality and cost need an outcome

Evals require criteria and traces; cost requires a unit of value. Non-zero spend must come from official interfaces, never an estimate treated as an invoice.

Reproducible public record

The sanitized mirror publishes the catalogue, research pack, 20 manifests and 20 receipts. 423,265 bytes were received in 21 HTTP calls and 37,675 ms combined, without a model call and without retaining the retrieved content.

Inspect evidence
Official billing: the OpenAI, Anthropic, and Google engines have been documented, but remain documented-not-configured. This framework-only run did not call models, so the measured direct model cost was $0. Future product cost remains unknown until there is authorized export and reconciliation by project/SKU.

Fail-closed: Pages that returned 403, triggered DLP/prompt-injection, or changed hash were not forced or counted. They were replaced by official immutable snapshots. The result proves the intake and research mechanism in this cutoff; it does not prove operation, attestation, product adoption, or global leadership.

How to read

Evidence before
adjective.

“Market consensus” is the intersection of public practices and standards — not a certification or a statistical average. Unknown, unavailable, and zero states are treated as different values.

Hierarchy of evidence

  1. E1: real event attested, linked to the source, SHA, policy and verifier.
  2. E2: test, CI, signed artifact or reproducible contract.
  3. E3: implementation present in the audited SHA.
  4. E4: decision accepted and linked to implementation.
  5. E5: statement without independent proof.

Qualification rule

No status A or O it is supported only by implementation or documentation. Valid artifact or provenance alone does not make the system compliant; operation requires SLO, cost, rollback, ownership and sustained sample.

Release 0.8: sanitized public proof

The tag and source SHA identify the cut; the names and digests of the 9 assets are in SHA256SUMS, whose Ed25519 signature can be verified with the public key. The signature does not cover the Git tag.

View proof record

Public border: the full canonical report lives in Trustyu's private documentation repository; This page and the evidence pack v1.7 are the sanitized public mirror of the approved cut. The organizations cited are sources of research and practice — there is no claim of partnership, certification or endorsement.