Executive reading · ~60 seconds
The portfolio has changed, but the durable decision remains based on task, risk, authority, cost per outcome accepted, evals and evidence — never on a fixed role per model.
Within weeks, model names, prices, and capability narratives changed again. The easy reaction is to update a ranking and replace the model assigned to each squad role. The responsible reaction is to verify what changed, separate provider claims from evidence, and recalculate the execution route for each task.
This is the Radar Vivo FORGE 2026/T3 — edition 01. It does not choose a universal winner. It records a time-bounded snapshot and turns portfolio changes into testable decisions.
Executive summary
The August 21, 2026 cutoff shows four relevant movements:
- OpenAI now presents Sol, Terra and Luna as a frontier, balance and volume portfolio, not as a single model for everything;
- Anthropic positions Fable for long, ambitious work, with additional safety controls and fallback;
- the price difference between routes became large enough to make routing and evals part of the architecture, not a late optimization;
- no release eliminates context, tool, sandbox, policy, verification, or human judgment.
The useful question is not ‘which model is best?’. It is:
Which combination produces the lowest cost per accepted outcome, within the risk and latency that this task allows?
The cut: facts stated by suppliers
The numbers below were re-read from primary sources in the editorial cut. This is data from the suppliers themselves, not independent validation from Trustyu.
| Route | Published positioning | Input / 1M tokens | Output / 1M tokens | Initial reading |
|---|---|---|---|---|
| GPT-5.6 Sol | complex and frontier professional work | US$ 5,00 | US$ 30,00 | candidate for difficult tasks where marginal quality can pay the cost |
| GPT-5.6 Terra | balance between intelligence, speed and cost | US$ 2,00 | US$ 12,00 | Default route candidate only after winning product baseline |
| GPT-5.6 Luna | cost-sensitive, high-volume work | $0.20 | $1.20 | candidate for screening, transformation and volume when the acceptance criteria remains preserved |
| Claude Fable 5 | long and ambitious work, with safeguards and fallback | US$ 10,00 | US$ 50,00 | candidate for long-horizon tasks that justify additional cost, latency, and policy |
Sources: GPT-5.6 portfolio guide, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Claude Fable and Anthropic Transparency Hub.
List price helps formulate hypotheses. It excludes cache, tools, retries, human review, platform costs, incidents, and rejection rate. It also does not establish performance on your case distribution.
What really changed
Portfolio became an architectural decision
When the price gap between Luna and Sol spans orders of magnitude, ‘use the best model’ is no longer a sufficient policy. An application may use a low-cost route for classification, a balanced route for frequent execution, and a frontier route for exceptions — provided the router is observable, testable, and unable to expand authority silently.
Long horizon became a product, not a generic promise
Fable reinforces competition for tasks that span longer stages, tools and periods. This matters for research, engineering and operations. But duration is not progress. Without external state, handoff, budget, checkpoints, and stopping criteria, a long run can only err for longer.
Security entered the model’s positioning
Provider safeguards, fallback, and controls are part of the assessment. They must not become the company's security boundary. Identity, capability, data, network, sandbox, and approval remain under the control of the system using the model.
The cost of coordination became more visible
Adding a planner, executor, evaluator, or multiple models may improve a use case while materially increasing cost and latency. FORGE requires task-specific evals and ablation before making any layer a default. [HARNESS31-C3]
Why not pin “Architect=Fable” or “Backend=Terra”
Roles describe responsibility. Models provide capabilities that change with releases, prices, limits, and policies. Pinning a model name to a role creates three forms of debt:
- obsolescence: the map ages as soon as the portfolio changes;
- false security: model choice appears to replace capability and review;
- invisible cost: the route is repeated without comparing acceptance, retries, and intervention.
The most durable contract is orchestrated:
tarefa + risco + dado + authority + orçamento + latência
↓
eval da rota candidata
↓
execução isolada + verificação
↓
outcome aceito + evidência + custo
The role remains stable. The model route is versioned and can change without rewriting the organization.
A decision matrix that survives the next release
Evaluate each candidate route against the same set of cases:
| Dimension | Gate question | Minimal evidence |
|---|---|---|
| task | what bounded result needs to be produced? | spec and representative examples |
| risk | What harm can a wrong response or action cause? | risk class and escalation rule |
| data | What information can enter and remain with the provider? | policy, retention and classification |
| authority | Does the model only propose or also execute? | capabilities deny-by-default |
| quality | what material difference is there against the baseline? | Blind eval with acceptance criteria |
| cost | how much does the accepted unit cost, not just the call? | tokens, tools, retries and human time |
| latency | Does the route fit into the actual flow? | p50, p95 and timeout by class |
| verifiability | Can we explain and reproduce acceptance? | logs, artifacts, digests and reviewer |
No external source defines the correct threshold of the FORGE or a specific product. Baselines and thresholds need to be calibrated per use case. [R3-C1]
Route by cost per outcome accepted
An honest comparison uses the same denominator:
modelo + contexto + ferramentas + plataforma + verificação
+ retries + retrabalho + intervenção + incidentes
------------------------------------------------------
outcomes aceitos
The framework offers the method; the unit of value remains specific to the product. [R3-C2]
In practice, the cheapest route may generate more correction. The most expensive one may not produce material gain. The result only appears when cost, quality and operation are measured together.
A seven-step adoption protocol
- freeze the case set and the current baseline;
- record model, version, parameters, policy and test date;
- execute the candidate routes with the same authority and the same tools;
- evaluate quality without revealing the supplier when this is feasible;
- measure acceptance rate, retries, latency, intervention and total cost;
- ablate layers — planner, evaluator, search and memory — before keeping them;
- promote a versioned routing policy, with fallback, budget, and rollback.
The objective is not to immortalize the winner of the cut. It's about making the next change small, observable and reversible.
What this Radar proves — and what it doesn't prove
This Radar establishes that the official pages published the portfolio, prices, and positioning described at the stated cutoff. OpenAI sources are not independent from one another; neither are Anthropic sources. System cards and transparency reports improve provider visibility, but they do not replace independent evaluation or evidence from your product.
We did not run our own comparative benchmark between Sol, Terra, Luna and Fable in this edition. Therefore, Radar does not claim overall superiority, productivity gains, operational safety or real savings for any Trustyu product.
Editorial validity
| Field | Commitment of this edition |
|---|---|
| factual cut | August 21, 2026 |
| prices and availability | reread by September 20, 2026 at the latest |
| architectural recommendation | review by November 19, 2026 at the latest |
| early review | new price, access, system card, data policy or material independent result |
| next edition | replaces this reading; the previous edition remains archived and dated |
Conclusion
Fable, Sol, Terra and Luna expand the options. They do not remove the responsibility to choose, limit, measure and verify.
The premium system is not the one that displays the newest name on each card. It is what changes models without losing memory, policy, security, evidence or cost clarity.
Further reading: The model is not the system and FORGE references.
Editorial and responsibility note
- Research cutoff
- Last review
- Recorded corrections
- No corrections recorded.
This article combines cited sources, analysis, and the author's professional experience. Verifiable data and factual statements are linked to their respective sources. Interpretations, hypotheses, projections, recommendations, and opinions represent the author's professional point of view at the time of publication; they do not constitute proven facts, a promise of results, or legal, financial, or technical advice applicable to a specific case. Consult the original sources and qualified professionals before making decisions.
Claims and sources
HARNESS31-C3
Anthropic reports that planner, generator and evaluator layers improved a long-running application result while materially increasing cost and latency; FORGE should require task-specific evals and ablation before making such layers mandatory.
Limit: This is a provider experiment tied to a selected model, benchmark and application task; it cannot prove universal superiority of multi-agent or evaluator layers. This is a source-bound design input; it does not prove product adoption, operational maturity, independent attestation, search ranking, AI citation or outcome.
- Anthropic — Anthropic, proprietary-site-terms
R3-C1
An AI quality claim requires explicit eval criteria and observable execution evidence; traces alone and subjective inspection alone are incomplete.
Limit: Neither source defines FORGE-specific pass thresholds; baselines must be calibrated per use case. This is a research-synthesis design input; it does not prove product adoption, operational maturity, independent attestation or outcome.
R3-C2
Useful AI observability must connect technical telemetry to a unit of value and its cost, instead of optimizing aggregate spend without an outcome denominator.
Limit: A unit must be product-specific; this framework-only run measures intake cost but not customer outcome. This is a research-synthesis design input; it does not prove product adoption, operational maturity, independent attestation or outcome.
- OpenTelemetry — OpenTelemetry, CC-BY-4.0 documentation; analysis-only excerpts
- FinOps Foundation — FinOps Foundation, CC-BY-4.0 framework; analysis-only excerpts