FORGE 3.1 · Specified/Shadow — by trustyu.ai

Serious architecture.
Real AI.
Lasting legacy.

We organize market knowledge, AI-assisted work, and senior leadership into a verifiable process. Outcomes, security, and scale are qualified by each product's evidence.

20
receipted sources
12
independent groups
3
research studies
9
projected claims
dod-loop — feature/auth-flow — PRD
# Loop spec-driven autônomo — ADR-0030
# SPEC → EXPLORE → PLAN → RED → IMPLEMENT → GREEN → SELF-VERIFY

01·SPEC feature-spec ··· ✓ out-of-scope declarado
02·EXPLORE blast-radius ·· ✓ 3 arquivos afetados
03·PLAN blueprint ···· ✓ revisão arquitetural
04·RED contract-first · ↯ 7 testes falham
05·IMPL codegen ····· → mypy: clean, ruff: clean
06·GREEN suite ······ ✓ 7/7 — coverage measured
07·VERIFY smoke ····· ✓ DoD satisfeito

feature-spec-gate ✓ CI gate green — PR ready
agent: claude/auth-flow origem: Claude Code

$
What We Build

Beyond SaaS.
Result-oriented software.

These four categories guide product discovery and design. Suitability, adoption and business outcome need to be validated in each context.

🏭
SaaS vertical
Deep software for a specific segment. Business domain coded, not configured. Neither too horizontal to be useful, nor too custom to scale.
🤖
Agentic Software
Target: agents execute delimited flows, with human supervision in critical actions and evidence of execution.
🧠
Vertical AI
Target: domain knowledge, persona, guardrails and trackable outputs are evaluated in the product.
📊
Outcome-as-a-Service
Commercial hypothesis: linking price to a unit of value; metering and billing remain pending.
Our Central Thesis

AI does not replace professionals.
Senior professionals lead AI.

We treat oversight, architecture, and evidence as parts of the system. Source-bound research supports this design; it does not promise a universal win rate or replace product validation.

👤 Senior Professional

Strategic judgment, domain expertise, architectural decisions, accountability.

🏗️ Architecture & Governance

Security by design, multi-tenancy and compliance as versioned decisions for durable evolution.

🔎 Empirical Validation

Human in the loop at critical points. Tests before code. Evidence before declaring ready.

+
> sum
of its parts
=

⚡ Assisted Execution

Agents can perform repeatable work within a spec, capability limits, and verifiable feedback.

🔁 Evaluable Consistency

Standards, tests, and receipts make the result inspectable; consistency continues to be measured, not assumed.

📐 Reuse with Qualification

The Template offers defaults; each product still needs to prove its adoption and its results in the SHA itself.

Risk 01

AI without human governance increases risk

Model capability does not replace isolation, least privilege, testing, provenance, and review of irreversible actions.

Risk 02

Adoption without a harness creates fragility

Secure adoption depends on explicit workflow, reviewable changes, and durable decisions; a better prompt does not replace harness.

Opportunity

Human + AI demands measurable gains

Quality requires eval criteria and observable execution; cost is only useful when linked to a unit of product value.

Definition of Done

Nothing is ready
without going through the 7 steps.

Spec-driven loop defined as Template default. Since ADR-0068, branch convention, source footer and session seat are no longer warnings and are now blocking within the always-on gate. The coverage of each product continues to be declared by evidence linked to the SHA itself.

01 · SPEC
Specification
Human defines
02 · EXPLORE
Exploration
AI performs
03 · PLAN
Blueprint
Revision
04 · RED
Tests First
Mandatory gate
05 · IMPLEMENT
Implementation
AI performs
06 · GREEN
Green Suite
Gate CI
07 · SELF-VERIFY
Self-check
AI iterates

Why formal DoD changes the game: AI without a correction oracle invents a way out. With SPEC→RED→GREEN, AI has a verifiable target — not code that looks right, but tests that prove it is right.

✓ spec present · ✓ tests in diff · ✓ coverage by profile/risk · ✓ lint clean
feature-spec-gate — CI check
# ADR-0030 — default verificável por produto

spec-present ······ ok
tests-in-diff ····· ok
coverage evidence · profile + SHA
out-of-scope ····· declared
ruff ············ clean
mypy ··········· clean

gate: EXAMPLE — template default
produto requer evidence própria

$

Out-of-scope protection: each feature declares what it doesn't do. Kills scope creep and prevents AI from expanding scope on its own.

Evidence-driven engineering

Public standards.
Verifiable controls.

Architecture, security and delivery practices are translated into defaults and controls. Adoption and results remain dependent on the evidence for each product.

🗺️
Domain-Driven Design
Bounded Contexts, Aggregate Roots and Domain Events are part of the reference design; each product validates the model in its own domain.
DDD · ADR-0012
🧱
Modular Monolith
Modular monolith is the architectural default. Service extraction, auth, billing and AI depend on the decision and product evidence.
Clean Architecture
🔒
Security by Design
Threat-model STRIDE, scanners, supply-chain hardening, hooks and RLS are reference controls; presence and enforcement are verified by consumer.
STRIDE · Supply-chain · RLS
📝
Conventional Commits + SemVer
Commitlint, versioning, CHANGELOG and VersionBadge are workflow defaults; execution must be demonstrated in the consumer repository.
semantic-release · ADR-0023
🧪
Risk Testing
Tests and coverage are measured by profile and risk. The ADR-0069 fixes the limit of this proportionality: depth governs the battery of quality and never the axes of security and semantics — a scanner runs because the code has changed, not because the risk has been classified as high.
Proportional by axis · ADR-0069
🌍
Hierarchical Multi-tenancy
Schema, RLS, RBAC and identity in JWT are defaults; Effective isolation requires testing linked to the environment and the SHA of the product.
RLS · RBAC · 5 roles
💰
Build & Deploy FinOps
Rules to avoid unnecessary deployments and measure usage are part of the design. Complete cost and unit economics follow in qualification per product.
Budgets · ADR-0027
🔄
Isolation Worktrees
One feature per worktree is the coordination protocol. Collisionlessness remains an observed result, not a universal promise.
1 feature = 1 worktree
📋
ADRs — Versioned Decisions
ADRs record decisions, status and evolution; Current status and compliance remain subject to verification by the applicable cutoff.
Decision · Explicit status
Product Cycle

From the problem to
qualifiable product.

Six phases form the cycle default. The product only advances when its evidence meets the applicable criteria; the human maintains the decision.

P1
Discovery & Domain
Domain, Bounded Contexts, personas, real pain validated.
Explain in 2 minutes who pays
P2
Architecture & Scope
Product ADRs, legacy stack, security before commit.
Documented decisions
P3
Bootstrap & Infrastructure
Repository receives CI/CD and environment defaults.
Health check as a target
P4
MVP Development
DoD loop by feature and gates proportional to risk.
Pilot criterion
P5
Validation & Polish
E2E, incidents and cost are measured in the defined window.
Product evidence
P6
Launch & Iterate
Structured feedback and learning returns to the method.
Outcome validated

Why this cycle matters: domain, architecture and infrastructure are explained before scaling. Reduction in rework and maintenance costs are hypotheses to be measured per product.

How AI Operates Correctly

Specialized squad.
Roles, tools, loops.

The protocol models agents by role, allowed tools, loop and seat SQ. The diagram below describes the mechanism; adoption and outcome depend on observed execution.

The role defines responsibility and capability; the model is selected by task. Model names, prices and status are in the radar dated models.

🏛️
Architect
ADRs, DDD, structural decisions
Frontier reasoningReview · read-only
🔐
Security
Threat modeling, SAST, RLS
Critical reviewIsolation read-only
💻
Backend
FastAPI, SQLAlchemy, Alembic
Balanced executionComplete tools · gates
🖥️
Frontend
Next.js, TypeScript, shadcn/ui
Balanced executionComplete tools · gates
🤖
AI Engineer
LangGraph, LLMFactory, RAG
Agentic executionDelimited Tools · eval
🧪
QA
Contract-First, Playwright E2E
VerificationTests · reading
⚙️
SRE/DevOps
CI/CD, deployment, monitoring
Supervised operationApproval gates · rollback
✍️
Tech Writer
ADRs, READMEs, specs
Technical summaryReading · document writing
routing-policy — reviewer:security
# política estável; seleção concreta em runtime

routing:
strategy: task-aware
criteria: [risk, complexity, eval, cost, latency]
tier: frontier-reasoning

tools:
Read ·········· allowed
Grep ········· allowed
Glob ········· allowed
Write ······· DENIED
Bash ······· DENIED

reviewer cannot write code. by design.
squad roster — sessões ativas (SQ)
# ADR-0034 — assento SQ por sessão de IA (D7)

claude·SQ7 · claude/auth-flow ··· :3070 RED→GREEN
codex·SQ3 · codex/webhook-retry :3030 REVIEW
claude·SQ11 · claude/rls-tenant ·· :3110 SPEC

seat alloc: wt new → menor SQ livre (sem colisão)
EXEMPLO · 3 sessões · 1 worktree por sessão

Specialized skills

The catalog associates role, expected behavior, guardrails, DoD and eval-harness. Inventory and adoption are verified at the applicable cutoff.

Catalog · Eval-driven

Routing by task, not by role

Risk, complexity, eval, cost and latency decide the tier. The concrete model can change without rewriting squad roles, tools or gates.

Stable policy · live catalog

Auditable D1–D7 Protocol

Mandatory pre-work, named branches, identified PR, isolated worktrees and session seating. D1–D5 coordinate · D6 isolate · D7 identify.

D1-D7 · Multi-session

Declared autonomy boundary

Runs SPEC→SELF-VERIFY autonomously — but scales to human in ambiguous spec, structural decision without ADR or access to PRD.

HITL · Explicit Gates
Foundation — Template Defaults

Product is born with
verifiable defaults.

The scaffold provides a reusable base. Each product must demonstrate adoption, configuration and enforcement in its own repository and SHA.

🔐 Security in Depth

Pipeline, supply chain, runtime and data receive reference controls. The consumer must demonstrate presence, configuration and enforcement.

Default · 5 layers

🏢 Multi-tenancy & RBAC

Schema, roles and identity in JWT are part of the default. Isolation is a property to test in the context of the product.

Isolation default

🚀 CI/CD with Double Gate

SemVer and VersionBadge are defaults. Since ADR-0064, the branch staging is justified by have a staging environment deployed, not by type of repository: those who publish by tag or SHA go from feature → main. Availability and FinOps require measurement in the consumer environment.

Branch by environment · ADR-0064

🧪 Quality as a Trail

Spec, tests and risk coverage are included as defaults; enforcement is only declared with evidence from the consumer.

Template default

🌍 i18n Native

Product is born internationalized. pt-BR default, en-US secondary. Override per client via bank — no rebuild.

Pt-BR · en-US

📡Governed Observability

OTel GenAI privacy-safe profile and source-bound checker integrate into runtime 0.18; the SHA256SUMS manifest has verifiable signature Ed25519. Real product telemetry requires rollout and evidence.

OTel GenAI · Shadow
Security in Depth

Six layers.
None trusts the previous one.

The design uses defense in depth: each layer has its own control and boundary. ADR and Template define the default; each product needs to prove that control is present and effective.

01
Pipeline
Gitleaks, CodeQL, Trivy and dependency alerts make up the reference profile; the consumer check demonstrates enforcement.
ADR-0028
02
Supply chain
Cooldown, dependency review and behavioral analysis are complementary controls provided for by default.
ADR-0039
03
Agent runtime
Hooks and per-role allowlists limit operations. The applied policy and its result must appear in the task evidence.
ADR-0037 0024
04
Data — Tenant Isolation
RLS, access extension and auditable escape hatch form the fail-closed design; Product tests prove insulation.
ADR-0040
05
Environment — STG ⊥ PRD
Cache namespace, separate instances, and infrastructure auditing are reference invariants, checked by environment.
ADR-0040 · STRIDE
06
Organization — who achieves the code
Access is granted to team, never the person: a direct grant to person→repository is treated as a governance defect, because disconnecting someone would depend on hunting for scattered grants. The default is zero exposure by default.
ADR-0067
Real case · incident P0 → platform invariant

A historical cross-environment caching incident prompted three benchmark invariants: environment isolation, tenant isolation, and defense in depth. Learning was codified in ADR; each product must still demonstrate its adoption.

AI Orchestration

AI as executor,
not as a copilot.

Models and providers change quickly. The architecture proposes to separate product and supplier behavior; real routes require evals, telemetry and product evidence.

Anthropic Claude evaluated route
Model family available to engineering partner
Activation by product requires eval and policy
OpenAI GPT & Codex evaluated route
Agentic engineering and runtime alternatives
Activation by product requires eval and policy
Google Gemini prepared
Long contexts and fast, lowest-cost routes
Declared routes; activation requires product eval
xAI Grok governed radar
New frontier option for reasoning and tool use
Evaluate by ADR + cost, quality, security and latency before adopting
Rule of Honesty: The engineering process roster does not prove runtime availability of any product. New provider should only enter with an adapter, tests and comparable evaluation.

Orchestration, not dependence

LLMFactory is the proposed mechanism to separate behavior and supplier; portability must be tested.

Cost linked to outcome

Schema and telemetry are targets. Official billing, reconciliation and unit of value still require product evidence.

Human at critical points

Gates before irreversible actions are part of the design; execution is checked per task.

RAG per domain

Embedding isolation and knowledge inheritance are defaults subject to tenant testing.

Source-bound research

Source received.
Bounded claim.

The approved public cut does not support TAM, CAGR or universal commercial productivity. That's why these numbers leave the page until they go through intake and license review.

20
Approved sources
Each source ID is linked to URL, snapshot, hash, usage basis and sanitized receipt.
Mechanism · cut Aug 4, 2026
12
Independent groups
The research avoids confusing the volume of links from the same publisher with corroboration.
Mechanism · source catalog
3
Multi-source research
APIs/coverage, multiplayer harness and source-bound publishing were evaluated separately.
Mechanism research pack
9
Projected claims
Owner, reviewer, validity, source bindings and inference boundary are machine readable.
Mechanism · projection claim
0
Raw external text
The mirror publishes provenance and receipts; prompt, transcript, PII, secret and full external corpus are prohibited.
Privacy-safe · zero-copy
6
Public states
Mechanism, default template, pilot observation, target, pending and historical are not to be confused.

What this cut allows us to say: reliability, security, provenance, evals and cost need to be properties of the system. Commercial results and product adoption remain outside this evidence.

Confidence to decide · evidence to audit

The model changes.
Your business continues.

The question is not just “which AI to use?” AND how to accelerate without turning customers, data and investment into an experiment. FORGE organizes people, decisions, controls and evidence around AI so that the product evolves without losing its foundation.

01 · SPEED WITH CONTROL

Target: reduce lead time with control

Repeatable work can be automated; Business, architectural and risk decisions remain under senior management.

02 · PROTECTED INVESTMENT

Target: reduce model switching cost

The separation between domain and provider is an architectural mechanism; Portability requires product testing.

03 · DEMONSTRABLE TRUST

Decisions that must leave evidence

Requirements, tests, approvals and releases are only rebuildable when linked to the source, SHA and responsible person.

04 · LEARNING THAT ACCUMULATES

A mistake doesn't have to become routine

The closed loop foresees that lessons return to the method as a test, control or rule; effectiveness needs to be measured.

Summary of approved sources.
Translated into method.

The public synthesis is derived only from the approved projection. External sources undergo intake, hashing, usage review and receipt; a discovery link alone does not support a claim.

Research performed, not just cited. At the RC cutoff, 20 approved sources from 12 independent groups supported 20 verifiable receipts, three studies, and nine traceable claims — without persisting raw external content. Inspect the public record →

0.8Historical SHA256SUMS with verifiable Ed25519 signature
17 Julcutoff date; not current status
20 / 20approved sources with verifiable receipt
17capabilities compared to the market
Historical · cut Jul 19, 2026

Framework-only readiness was recorded in that cut. It demonstrates the mechanism and availability of the then basis, not coverage, contracts or operation of every current product.

Rollout 3.1: Template defaults, pilot observations and operational qualification remain separate states. Operational and Attested are not declared.

The organizations cited are public sources of research and market practices; there is no claim of partnership, certification or endorsement. Current source-bound cutoff: Aug 4, 2026 · 03:27 UTC. Readiness framework-only historical reconciled on Jul 19, 2026.

Evolution board · 17 capabilities

From JARVIS to FORGE.
What really changed.

This is the technical table originally published on the homepage and preserved in the benchmark. It compares capabilities, processes and controls with market consensus. The 17 verdicts belong to the historic July 17 cutoff; FORGE 3.1 appears separately so as not to transform specification or shadow into operational proof.

Bench market 2025–2026

Reference standards

Primary and normative sources define what to observe in context, security, evals, evidence, operation and cost. They do not represent partnership or certification.

Source · JARVIS / Forge 1.x

Method and foundation

P1–P6 Templates, ADRs, architectural principles, squads and real products formed the basis. The period is qualitative: it did not receive a column or retroactive score.

FORGE 2.0 formal baseline

Repeatable process

Skills, worktrees, SPEC→PR, review-remediate, hooks, MCP, security and FinOps made the process comparable; the 2.1 transition added controls and readiness.

Current published status · FORGE 3.1

Specified/Shadow

Coverage, contracts and harnesses are defined as evidence by risk and SHA. Fleet Enforced, Operational and Attested remain outside the published claim.

How to read: the scale —/D/C/V/A/O means absent, documented, controlled, verified, attested and operating. The current state of 3.1 does not silently recalculate a frozen benchmark.

Open method, sources and evidence →
Maturity absent Ddocumented Ccontrolled Vchecked Acertificate Ooperating
Seventeen capabilities compared between the market consensus, FORGE 2.0, FORGE 2.1 and the historical cutoff of FORGE 3.0.
AxleMarket consensus 2025–2026Forge 2.0Forge 2.1Forge 3.0 cut · Jul 17Historic verdict
Context engineeringShort map, versioned knowledge, progressive disclosure and mechanical freshness.
CStrong but extensive ADRs, skills and instructions.
VBudgets, fragments, history and capability register.
VCanon-check and progressive disclosure mechanics.
StrongWhat remains is semantic drift between anchors and alive state.
ADR and alive statusPersistent/superseded decision; separate current operating state.
DVersioned decisions, status with residual drift.
CGovernance and living state catalog.
C0039/0040 reconciled and submitted for ratification.
AlignedSemantic freshness still requires human review.
Executable controlsOwner, applicability, evidence, enforcement, exception and deadline.
DRules distributed between ADRs and CI.
CDeterministic catalog and profiles.
V14 controls, diff facts and explicit evidence.
AboveIt remains to prove coverage in every applicable real change.
Spec-driven developmentSpec as source, out-of-scope, contracts before code and trace until acceptance.
CLoop SPEC→PR and Contract-First.
CStandardized feature spec and gate.
VTrace and independent verifiers.
StrongMeasure defects and rework avoided, not using the template.
Writing isolationA task/workspace/branch; parallel reading; single-writer per artifact.
CWorktree by feature and D5/D6 protocol.
CExplicit seats and roster.
VOverlap, ownership and teardown guards.
StrongRecord teardown on every real task.
Sandbox and least privilegeFS, network, secrets and tools per task; deny-by-default; short credential; HITL.
DGuardrails and worktree insulation.
CRisk profiles and capabilities.
CHub #683 and CRM #1621 passed provenance without passing the aggregate; #688 failed.
PartialLack of approved global gate, approvals/capabilities and longitudinal coverage.
Tool/MCP governanceMinimal tools, clear schemas, scopes, allowlist and auditability.
CCanonical MCP and inventory per repo.
CActivation by context and phase.
VCapability routing and validated contracts.
AlignedTest tool profile per task and control context cost.
AI evalsModel + harness + environment; isolated trials; calibrated graders; cost and latency.
DEval-harness in part of the skills.
CAI change baseline and gates.
V12/12 technical corpus with source/profile/SHA binding.
StrongScale up to 20–50 real failures and measure cost per task.
Evidence and attestationExecutor does not self-attest; source/SHA/profile binding; verifiable artifact.
DHeterogeneous evidence on CI and PR.
CEvidence contracts and defined provenance.
CSHA256SUMS manifest for runtime 0.8 with Ed25519 signature; synthetic fixture 6/6 Verified; Observer EVD=2/5 and VER=2/9.
PartialReal pack, coverage ≥95%, approved global gate and longitudinal window are missing.
supply chainImmutable lock/pin, SBOM, provenance, trusted builder and verification.
CLocks, scans and actions partially pinned.
CBaseline and evidence collectors.
VSHA256SUMS 0.8 manifest with Ed25519 signature and 9 inventoried assets; Template pinned with rollback 0.7.
StrongDo not declare SLSA level without all requirements met.
Readiness, SLO and rollbackSLI/SLO, error budget, tested rollback, runbook and stop condition.
DMostly manual readiness.
CSLO, scorecard and gates defined.
CObserver signed: 11 changes in 2 days; open window.
PartialRequires ≥20 changes, 30 days, stability and pilot drills.
Closed-loop learningReview or incident becomes test, eval, control or durable improvement.
CReview-remediate and versioned history.
CLearning records and ownership.
VFeedback produces tested controls, evals and regressions.
AlignedMeasure avoided recurrence and effectiveness, not number of records.
FinOps and unit economicsCost per unit of value/outcome, human attention and denial-of-wallet.
CDeployment skip and cost per product.
CBudgets by branch, agent and gate.
CM15: Actions −47.20%/day; minutes per change +28.27%.
Up in governanceMissed target by 2.80 p.p.; unit efficiency regressed; missing models and human time.
Senior roles and human controlSenior defines intention, risk and taste; proportional autonomy and clear approvals.
DPersonas and squad roles.
CActivation by risk/phase and human gates.
CExplicit owner and decision; single technical approver.
Aligned with caveatFounder/CTO is today the only real technical approver.
Agentic observabilityModel/tool ​​spans, outcome, tokens, cost, policy and privacy.
DTelemetry broken down by product.
CSchema of receipts and harness metrics.
CSubset OTel GenAI privacy-safe in runtime 0.8, covered by the signed SHA256SUMS manifest.
BackReal product telemetry linked to tokens/cost/outcome/receipt is missing.
Issue provenance/closeoutAuthorship, decision, implementation, validation and reconstructable residual.
DInconsistent closeouts and unproven records.
DProblem recognized, still no common gate.
CADR-0051 and closeout gate in canonical.
Above in the contractPartial in the fleet; Spread will occur by normal touch.
External adoption and assuranceIndependent user, package/license, support, upgrade/rollback and evidence pack.
No external framework product.
DDocumented extraction strategy.
DConformance/EEP inventoried by SHA256SUMS manifest signed in 0.8; no external adopters.
BackNo independent external adopters until cut off.
Evidence by state

Don't confuse mechanism
with operation.

This site only publishes what the corresponding artifact allows you to conclude. Product qualification 3.1 remains pending until evidence is linked to the repository, profile and SHA.

20 / 20
Mechanism
Intake and research exercised with receipts approved at the July 17th cutoff.
Default
Template default
Spec, tests and controls can be created in the scaffold; the consumer still needs to prove adoption.
2 days
Pilot observation
The Historical Observer recorded 11 shadow changes; This does not qualify the current fleet.
Target
Coverage and contracts
Risk Thresholds, OAS/events and ARR v2 are verifiable objectives, not delivered claims.
Pending
Product 3.1
No published eligible evidence to declare universal coverage, operation or attestation.
v1.7
Historical
Benchmark, runtime 0.8 and Observer remain frozen for traceability.
Sectors

Verticals with
real pain.

The verticals below are application hypotheses and roadmap. None represents adoption, operation, or proven results in this public cut.

⚖️
Legal Services
Lead qualification, case management and automation with agents specialized in the legal domain.
● Evidence 3.1 pending
🏥
Health
Clinical management, scheduling, medical records with AI. High regulation demands auditable architecture.
● Next
🏗️
Engineering & Construction
Hypothesis: project management, assisted reports and compliance.
● Pipeline
🎓
Education
Hypothesis: adaptive platforms, tutors and academic management.
● Roadmap
🏦
Financial Services
Hypothesis: compliance, risk analysis and auditable automation.
● Roadmap
🚀
Your Vertical
The product can come from the Template defaults and needs to be qualified in the context itself. Talk to us.
● Contact Us
Published source-bound cut

Human- and machine-readable evidence.

20
sources
20
receipts
12
groups
3
research studies
9
claims
21
HTTP calls
0
model calls
0
raw external text
Serious architecture. Real AI. Lasting legacy.

We do not believe in human replacement. The objective is to combine senior leadership, AI, formal DoD, and specialized roles; gains and results need to be measured in the context of each product.

Chat about your project Read our thesis

Shall we co-build AI for your market?

We bring together domain and engineering knowledge to specify, build and measure a product in the context of your market.

contato@trustyu.ai