Executive reading · ~60 seconds

The portfolio has changed, but the durable decision remains based on task, risk, authority, cost per outcome accepted, evals and evidence — never on a fixed role per model.

Within weeks, model names, prices, and capability narratives changed again. The easy reaction is to update a ranking and replace the model assigned to each squad role. The responsible reaction is to verify what changed, separate provider claims from evidence, and recalculate the execution route for each task.

This is the Radar Vivo FORGE 2026/T3 — edition 01. It does not choose a universal winner. It records a time-bounded snapshot and turns portfolio changes into testable decisions.

Executive summary

The August 21, 2026 cutoff shows four relevant movements:

  1. OpenAI now presents Sol, Terra and Luna as a frontier, balance and volume portfolio, not as a single model for everything;
  2. Anthropic positions Fable for long, ambitious work, with additional safety controls and fallback;
  3. the price difference between routes became large enough to make routing and evals part of the architecture, not a late optimization;
  4. no release eliminates context, tool, sandbox, policy, verification, or human judgment.

The useful question is not ‘which model is best?’. It is:

Which combination produces the lowest cost per accepted outcome, within the risk and latency that this task allows?

The cut: facts stated by suppliers

The numbers below were re-read from primary sources in the editorial cut. This is data from the suppliers themselves, not independent validation from Trustyu.

RoutePublished positioningInput / 1M tokensOutput / 1M tokensInitial reading
GPT-5.6 Solcomplex and frontier professional workUS$ 5,00US$ 30,00candidate for difficult tasks where marginal quality can pay the cost
GPT-5.6 Terrabalance between intelligence, speed and costUS$ 2,00US$ 12,00Default route candidate only after winning product baseline
GPT-5.6 Lunacost-sensitive, high-volume work$0.20$1.20candidate for screening, transformation and volume when the acceptance criteria remains preserved
Claude Fable 5long and ambitious work, with safeguards and fallbackUS$ 10,00US$ 50,00candidate for long-horizon tasks that justify additional cost, latency, and policy

Sources: GPT-5.6 portfolio guide, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Claude Fable and Anthropic Transparency Hub.

List price helps formulate hypotheses. It excludes cache, tools, retries, human review, platform costs, incidents, and rejection rate. It also does not establish performance on your case distribution.

What really changed

Portfolio became an architectural decision

When the price gap between Luna and Sol spans orders of magnitude, ‘use the best model’ is no longer a sufficient policy. An application may use a low-cost route for classification, a balanced route for frequent execution, and a frontier route for exceptions — provided the router is observable, testable, and unable to expand authority silently.

Long horizon became a product, not a generic promise

Fable reinforces competition for tasks that span longer stages, tools and periods. This matters for research, engineering and operations. But duration is not progress. Without external state, handoff, budget, checkpoints, and stopping criteria, a long run can only err for longer.

Security entered the model’s positioning

Provider safeguards, fallback, and controls are part of the assessment. They must not become the company's security boundary. Identity, capability, data, network, sandbox, and approval remain under the control of the system using the model.

The cost of coordination became more visible

Adding a planner, executor, evaluator, or multiple models may improve a use case while materially increasing cost and latency. FORGE requires task-specific evals and ablation before making any layer a default. [HARNESS31-C3]

Why not pin “Architect=Fable” or “Backend=Terra”

Roles describe responsibility. Models provide capabilities that change with releases, prices, limits, and policies. Pinning a model name to a role creates three forms of debt:

  • obsolescence: the map ages as soon as the portfolio changes;
  • false security: model choice appears to replace capability and review;
  • invisible cost: the route is repeated without comparing acceptance, retries, and intervention.

The most durable contract is orchestrated:

tarefa + risco + dado + authority + orçamento + latência
                         ↓
                  eval da rota candidata
                         ↓
             execução isolada + verificação
                         ↓
             outcome aceito + evidência + custo

The role remains stable. The model route is versioned and can change without rewriting the organization.

A decision matrix that survives the next release

Evaluate each candidate route against the same set of cases:

DimensionGate questionMinimal evidence
taskwhat bounded result needs to be produced?spec and representative examples
riskWhat harm can a wrong response or action cause?risk class and escalation rule
dataWhat information can enter and remain with the provider?policy, retention and classification
authorityDoes the model only propose or also execute?capabilities deny-by-default
qualitywhat material difference is there against the baseline?Blind eval with acceptance criteria
costhow much does the accepted unit cost, not just the call?tokens, tools, retries and human time
latencyDoes the route fit into the actual flow?p50, p95 and timeout by class
verifiabilityCan we explain and reproduce acceptance?logs, artifacts, digests and reviewer

No external source defines the correct threshold of the FORGE or a specific product. Baselines and thresholds need to be calibrated per use case. [R3-C1]

Route by cost per outcome accepted

An honest comparison uses the same denominator:

modelo + contexto + ferramentas + plataforma + verificação
+ retries + retrabalho + intervenção + incidentes
------------------------------------------------------
outcomes aceitos

The framework offers the method; the unit of value remains specific to the product. [R3-C2]

In practice, the cheapest route may generate more correction. The most expensive one may not produce material gain. The result only appears when cost, quality and operation are measured together.

A seven-step adoption protocol

  1. freeze the case set and the current baseline;
  2. record model, version, parameters, policy and test date;
  3. execute the candidate routes with the same authority and the same tools;
  4. evaluate quality without revealing the supplier when this is feasible;
  5. measure acceptance rate, retries, latency, intervention and total cost;
  6. ablate layers — planner, evaluator, search and memory — before keeping them;
  7. promote a versioned routing policy, with fallback, budget, and rollback.

The objective is not to immortalize the winner of the cut. It's about making the next change small, observable and reversible.

What this Radar proves — and what it doesn't prove

This Radar establishes that the official pages published the portfolio, prices, and positioning described at the stated cutoff. OpenAI sources are not independent from one another; neither are Anthropic sources. System cards and transparency reports improve provider visibility, but they do not replace independent evaluation or evidence from your product.

We did not run our own comparative benchmark between Sol, Terra, Luna and Fable in this edition. Therefore, Radar does not claim overall superiority, productivity gains, operational safety or real savings for any Trustyu product.

Editorial validity

FieldCommitment of this edition
factual cutAugust 21, 2026
prices and availabilityreread by September 20, 2026 at the latest
architectural recommendationreview by November 19, 2026 at the latest
early reviewnew price, access, system card, data policy or material independent result
next editionreplaces this reading; the previous edition remains archived and dated

Conclusion

Fable, Sol, Terra and Luna expand the options. They do not remove the responsibility to choose, limit, measure and verify.

The premium system is not the one that displays the newest name on each card. It is what changes models without losing memory, policy, security, evidence or cost clarity.

Further reading: The model is not the system and FORGE references.

Editorial and responsibility note

Research cutoff
Last review
Recorded corrections
No corrections recorded.

This article combines cited sources, analysis, and the author's professional experience. Verifiable data and factual statements are linked to their respective sources. Interpretations, hypotheses, projections, recommendations, and opinions represent the author's professional point of view at the time of publication; they do not constitute proven facts, a promise of results, or legal, financial, or technical advice applicable to a specific case. Consult the original sources and qualified professionals before making decisions.

Claims and sources

HARNESS31-C3

Anthropic reports that planner, generator and evaluator layers improved a long-running application result while materially increasing cost and latency; FORGE should require task-specific evals and ablation before making such layers mandatory.

Limit: This is a provider experiment tied to a selected model, benchmark and application task; it cannot prove universal superiority of multi-agent or evaluator layers. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.

  • Anthropic — Anthropic, proprietary-site-terms

R3-C1

An AI quality claim requires explicit eval criteria and observable execution evidence; traces alone and subjective inspection alone are incomplete.

Limit: Neither source defines FORGE-specific pass thresholds; baselines must be calibrated per use case. This is a research-synthesis design input; it does not prove product adoption, operational maturity, independent attestation or outcome.

  • Anthropic — Anthropic, MIT repository; analysis-only excerpts
  • OpenAI — OpenAI, MIT repository; analysis-only excerpts

R3-C2

Useful AI observability must connect technical telemetry to a unit of value and its cost, instead of optimizing aggregate spend without an outcome denominator.

Limit: A unit must be product-specific; this framework-only run measures intake cost but not customer outcome. This is a research-synthesis design input; it does not prove product adoption, operational maturity, independent attestation or outcome.

  • OpenTelemetry — OpenTelemetry, CC-BY-4.0 documentation; analysis-only excerpts
  • FinOps Foundation — FinOps Foundation, CC-BY-4.0 framework; analysis-only excerpts