Target: reduce lead time with control
Repeatable work can be automated; Business, architectural and risk decisions remain under senior management.
Human knowledge + artificial intelligence.
Judgment to decide. AI to scale up.
We organize market knowledge, AI-assisted work, and senior leadership into a verifiable process. Outcomes, security, and scale are qualified by each product's evidence.
These four categories guide product discovery and design. Suitability, adoption and business outcome need to be validated in each context.
We treat oversight, architecture, and evidence as parts of the system. Source-bound research supports this design; it does not promise a universal win rate or replace product validation.
Strategic judgment, domain expertise, architectural decisions, accountability.
Security by design, multi-tenancy and compliance as versioned decisions for durable evolution.
Human in the loop at critical points. Tests before code. Evidence before declaring ready.
Agents can perform repeatable work within a spec, capability limits, and verifiable feedback.
Standards, tests, and receipts make the result inspectable; consistency continues to be measured, not assumed.
The Template offers defaults; each product still needs to prove its adoption and its results in the SHA itself.
Model capability does not replace isolation, least privilege, testing, provenance, and review of irreversible actions.
Secure adoption depends on explicit workflow, reviewable changes, and durable decisions; a better prompt does not replace harness.
Quality requires eval criteria and observable execution; cost is only useful when linked to a unit of product value.
Spec-driven loop defined as Template default. Since ADR-0068, branch convention, source footer and session seat are no longer warnings and are now blocking within the always-on gate. The coverage of each product continues to be declared by evidence linked to the SHA itself.
Why formal DoD changes the game: AI without a correction oracle invents a way out. With SPEC→RED→GREEN, AI has a verifiable target — not code that looks right, but tests that prove it is right.
Out-of-scope protection: each feature declares what it doesn't do. Kills scope creep and prevents AI from expanding scope on its own.
Architecture, security and delivery practices are translated into defaults and controls. Adoption and results remain dependent on the evidence for each product.
Six phases form the cycle default. The product only advances when its evidence meets the applicable criteria; the human maintains the decision.
Why this cycle matters: domain, architecture and infrastructure are explained before scaling. Reduction in rework and maintenance costs are hypotheses to be measured per product.
The protocol models agents by role, allowed tools, loop and seat SQ. The diagram below describes the mechanism; adoption and outcome depend on observed execution.
The role defines responsibility and capability; the model is selected by task. Model names, prices and status are in the radar dated models.
The catalog associates role, expected behavior, guardrails, DoD and eval-harness. Inventory and adoption are verified at the applicable cutoff.
Catalog · Eval-drivenRisk, complexity, eval, cost and latency decide the tier. The concrete model can change without rewriting squad roles, tools or gates.
Stable policy · live catalogMandatory pre-work, named branches, identified PR, isolated worktrees and session seating. D1–D5 coordinate · D6 isolate · D7 identify.
D1-D7 · Multi-sessionRuns SPEC→SELF-VERIFY autonomously — but scales to human in ambiguous spec, structural decision without ADR or access to PRD.
HITL · Explicit GatesThe scaffold provides a reusable base. Each product must demonstrate adoption, configuration and enforcement in its own repository and SHA.
Pipeline, supply chain, runtime and data receive reference controls. The consumer must demonstrate presence, configuration and enforcement.
Default · 5 layersSchema, roles and identity in JWT are part of the default. Isolation is a property to test in the context of the product.
Isolation defaultSemVer and VersionBadge are defaults. Since ADR-0064, the branch staging is justified by have a staging environment deployed, not by type of repository: those who publish by tag or SHA go from feature → main. Availability and FinOps require measurement in the consumer environment.
Spec, tests and risk coverage are included as defaults; enforcement is only declared with evidence from the consumer.
Template defaultProduct is born internationalized. pt-BR default, en-US secondary. Override per client via bank — no rebuild.
Pt-BR · en-USOTel GenAI privacy-safe profile and source-bound checker integrate into runtime 0.18; the SHA256SUMS manifest has verifiable signature Ed25519. Real product telemetry requires rollout and evidence.
OTel GenAI · ShadowThe design uses defense in depth: each layer has its own control and boundary. ADR and Template define the default; each product needs to prove that control is present and effective.
A historical cross-environment caching incident prompted three benchmark invariants: environment isolation, tenant isolation, and defense in depth. Learning was codified in ADR; each product must still demonstrate its adoption.
Models and providers change quickly. The architecture proposes to separate product and supplier behavior; real routes require evals, telemetry and product evidence.
LLMFactory is the proposed mechanism to separate behavior and supplier; portability must be tested.
Schema and telemetry are targets. Official billing, reconciliation and unit of value still require product evidence.
Gates before irreversible actions are part of the design; execution is checked per task.
Embedding isolation and knowledge inheritance are defaults subject to tenant testing.
The approved public cut does not support TAM, CAGR or universal commercial productivity. That's why these numbers leave the page until they go through intake and license review.
What this cut allows us to say: reliability, security, provenance, evals and cost need to be properties of the system. Commercial results and product adoption remain outside this evidence.
The question is not just “which AI to use?” AND how to accelerate without turning customers, data and investment into an experiment. FORGE organizes people, decisions, controls and evidence around AI so that the product evolves without losing its foundation.
Repeatable work can be automated; Business, architectural and risk decisions remain under senior management.
The separation between domain and provider is an architectural mechanism; Portability requires product testing.
Requirements, tests, approvals and releases are only rebuildable when linked to the source, SHA and responsible person.
The closed loop foresees that lessons return to the method as a test, control or rule; effectiveness needs to be measured.
The public synthesis is derived only from the approved projection. External sources undergo intake, hashing, usage review and receipt; a discovery link alone does not support a claim.
Research performed, not just cited. At the RC cutoff, 20 approved sources from 12 independent groups supported 20 verifiable receipts, three studies, and nine traceable claims — without persisting raw external content. Inspect the public record →
Framework-only readiness was recorded in that cut. It demonstrates the mechanism and availability of the then basis, not coverage, contracts or operation of every current product.
Rollout 3.1: Template defaults, pilot observations and operational qualification remain separate states. Operational and Attested are not declared.
The organizations cited are public sources of research and market practices; there is no claim of partnership, certification or endorsement. Current source-bound cutoff: Aug 4, 2026 · 03:27 UTC. Readiness framework-only historical reconciled on Jul 19, 2026.
This is the technical table originally published on the homepage and preserved in the benchmark. It compares capabilities, processes and controls with market consensus. The 17 verdicts belong to the historic July 17 cutoff; FORGE 3.1 appears separately so as not to transform specification or shadow into operational proof.
Primary and normative sources define what to observe in context, security, evals, evidence, operation and cost. They do not represent partnership or certification.
P1–P6 Templates, ADRs, architectural principles, squads and real products formed the basis. The period is qualitative: it did not receive a column or retroactive score.
Skills, worktrees, SPEC→PR, review-remediate, hooks, MCP, security and FinOps made the process comparable; the 2.1 transition added controls and readiness.
Coverage, contracts and harnesses are defined as evidence by risk and SHA. Fleet Enforced, Operational and Attested remain outside the published claim.
How to read: the scale —/D/C/V/A/O means absent, documented, controlled, verified, attested and operating. The current state of 3.1 does not silently recalculate a frozen benchmark.
Open method, sources and evidence →| Axle | Market consensus 2025–2026 | Forge 2.0 | Forge 2.1 | Forge 3.0 cut · Jul 17 | Historic verdict |
|---|---|---|---|---|---|
| Context engineering | Short map, versioned knowledge, progressive disclosure and mechanical freshness. | CStrong but extensive ADRs, skills and instructions. |
VBudgets, fragments, history and capability register. |
VCanon-check and progressive disclosure mechanics. |
StrongWhat remains is semantic drift between anchors and alive state. |
| ADR and alive status | Persistent/superseded decision; separate current operating state. | DVersioned decisions, status with residual drift. |
CGovernance and living state catalog. |
C0039/0040 reconciled and submitted for ratification. |
AlignedSemantic freshness still requires human review. |
| Executable controls | Owner, applicability, evidence, enforcement, exception and deadline. | DRules distributed between ADRs and CI. |
CDeterministic catalog and profiles. |
V14 controls, diff facts and explicit evidence. |
AboveIt remains to prove coverage in every applicable real change. |
| Spec-driven development | Spec as source, out-of-scope, contracts before code and trace until acceptance. | CLoop SPEC→PR and Contract-First. |
CStandardized feature spec and gate. |
VTrace and independent verifiers. |
StrongMeasure defects and rework avoided, not using the template. |
| Writing isolation | A task/workspace/branch; parallel reading; single-writer per artifact. | CWorktree by feature and D5/D6 protocol. |
CExplicit seats and roster. |
VOverlap, ownership and teardown guards. |
StrongRecord teardown on every real task. |
| Sandbox and least privilege | FS, network, secrets and tools per task; deny-by-default; short credential; HITL. | DGuardrails and worktree insulation. |
CRisk profiles and capabilities. |
CHub #683 and CRM #1621 passed provenance without passing the aggregate; #688 failed. |
PartialLack of approved global gate, approvals/capabilities and longitudinal coverage. |
| Tool/MCP governance | Minimal tools, clear schemas, scopes, allowlist and auditability. | CCanonical MCP and inventory per repo. |
CActivation by context and phase. |
VCapability routing and validated contracts. |
AlignedTest tool profile per task and control context cost. |
| AI evals | Model + harness + environment; isolated trials; calibrated graders; cost and latency. | DEval-harness in part of the skills. |
CAI change baseline and gates. |
V12/12 technical corpus with source/profile/SHA binding. |
StrongScale up to 20–50 real failures and measure cost per task. |
| Evidence and attestation | Executor does not self-attest; source/SHA/profile binding; verifiable artifact. | DHeterogeneous evidence on CI and PR. |
CEvidence contracts and defined provenance. |
CSHA256SUMS manifest for runtime 0.8 with Ed25519 signature; synthetic fixture 6/6 Verified; Observer EVD=2/5 and VER=2/9. |
PartialReal pack, coverage ≥95%, approved global gate and longitudinal window are missing. |
| supply chain | Immutable lock/pin, SBOM, provenance, trusted builder and verification. | CLocks, scans and actions partially pinned. |
CBaseline and evidence collectors. |
VSHA256SUMS 0.8 manifest with Ed25519 signature and 9 inventoried assets; Template pinned with rollback 0.7. |
StrongDo not declare SLSA level without all requirements met. |
| Readiness, SLO and rollback | SLI/SLO, error budget, tested rollback, runbook and stop condition. | DMostly manual readiness. |
CSLO, scorecard and gates defined. |
CObserver signed: 11 changes in 2 days; open window. |
PartialRequires ≥20 changes, 30 days, stability and pilot drills. |
| Closed-loop learning | Review or incident becomes test, eval, control or durable improvement. | CReview-remediate and versioned history. |
CLearning records and ownership. |
VFeedback produces tested controls, evals and regressions. |
AlignedMeasure avoided recurrence and effectiveness, not number of records. |
| FinOps and unit economics | Cost per unit of value/outcome, human attention and denial-of-wallet. | CDeployment skip and cost per product. |
CBudgets by branch, agent and gate. |
CM15: Actions −47.20%/day; minutes per change +28.27%. |
Up in governanceMissed target by 2.80 p.p.; unit efficiency regressed; missing models and human time. |
| Senior roles and human control | Senior defines intention, risk and taste; proportional autonomy and clear approvals. | DPersonas and squad roles. |
CActivation by risk/phase and human gates. |
CExplicit owner and decision; single technical approver. |
Aligned with caveatFounder/CTO is today the only real technical approver. |
| Agentic observability | Model/tool spans, outcome, tokens, cost, policy and privacy. | DTelemetry broken down by product. |
CSchema of receipts and harness metrics. |
CSubset OTel GenAI privacy-safe in runtime 0.8, covered by the signed SHA256SUMS manifest. |
BackReal product telemetry linked to tokens/cost/outcome/receipt is missing. |
| Issue provenance/closeout | Authorship, decision, implementation, validation and reconstructable residual. | DInconsistent closeouts and unproven records. |
DProblem recognized, still no common gate. |
CADR-0051 and closeout gate in canonical. |
Above in the contractPartial in the fleet; Spread will occur by normal touch. |
| External adoption and assurance | Independent user, package/license, support, upgrade/rollback and evidence pack. | —No external framework product. |
DDocumented extraction strategy. |
DConformance/EEP inventoried by SHA256SUMS manifest signed in 0.8; no external adopters. |
BackNo independent external adopters until cut off. |
This site only publishes what the corresponding artifact allows you to conclude. Product qualification 3.1 remains pending until evidence is linked to the repository, profile and SHA.
The verticals below are application hypotheses and roadmap. None represents adoption, operation, or proven results in this public cut.
We do not believe in human replacement. The objective is to combine senior leadership, AI, formal DoD, and specialized roles; gains and results need to be measured in the context of each product.
We bring together domain and engineering knowledge to specify, build and measure a product in the context of your market.