🧠 AI Model Comparison

July 2026 field guide for your stack: Claude Max 5x (Fable 5 · Opus 4.8) · ChatGPT Plus (GPT-5.6) · Grok (Grok 4.5 High) · Ollama Cloud. Pick model then effort. Hover chips top-right to highlight. Default: plan Fable/Opus/Sol high → build Terra/Grok high → review another session.

How to read this tab: Charts compare published peaks — not always daily effort. Fable leads SWE-Pro (80.3%). Opus 4.8 = dual Claude Max 5x (SWE-Pro 69.2 · TB 74.6 Terminus-2 · AA 61.4 launch · $5/$25). Sol leads many agentic/terminal scores. Terra = cost-aware implementer. Grok 4.5 high = your Grok plan (CA~76, TB 83.3%, cheap dual). Ollama Cloud = Qwen plan · GLM code.

🧪 SWE-Bench Pro % (repo-level coding)

Real multi-file GitHub tasks. Fable leads; Sol/Terra trail on this bench even when they win agentic scores.

🧠 General Intelligence Index

Artificial Analysis composite (reasoning, agents, science). Good proxy for planning / hard thinking.

⚡ Cost Efficiency (Score per $)

Ordinal guide (not a lab metric). Higher = more capability per dollar. Free Ollama scores high on value; flagships trade money for ceiling. Opus $5/$25 ≈ better $/M than Fable $10/$50 at near-frontier peaks.

🔭 Long-Context (1M token quality)

Peer-scale 1M quality (not AA-LCR %). Critical for big repos & RAG. Opus AA-LCR is roughly flat vs 4.7 (~68% on public AA-LCR boards — different scale).

Model Provider Type Params Context SWE-Bench Pro Terminal-Bench 2.1 Intel Index Input $/M Output $/M
Claude Fable 5 high Anthropic ☁️ Cloud 1M 80.3 ~84–86 60–65 $10 $50
Claude Fable 5 xhigh Anthropic ☁️ Cloud 1M 80.3+ ~86 ~60 max $10 $50
Claude Opus 4.8 high → max Anthropic ☁️ Cloud 1M 69.2 74.6† 61.4‡ $5 $25
GPT-5.6 Sol medium OpenAI ☁️ Cloud 1M ~65 ~86 ~56 $5 $30
GPT-5.6 Sol high OpenAI ☁️ Cloud 1M ~66 ~88 ~58 $5 $30
GPT-5.6 Sol xhigh / max OpenAI ☁️ Cloud 1M 64.6 Pro 88.8 59 max $5 $30
GPT-5.6 Terra medium OpenAI ☁️ Cloud 1M ~62 ~85 ~52 $2.50 $15
GPT-5.6 Terra high OpenAI ☁️ Cloud 1M ~63 ~86 ~54 $2.50 $15
GPT-5.6 Terra xhigh / max OpenAI ☁️ Cloud 1M ~64 87.4 55 max $2.50 $15
Grok 4.5 high xAI / SpaceXAI ☁️ Cloud MoE 500K–1M* 64.7 83.3 54 (high) $2 $6
Kimi K3 Moonshot AI ☁️ Cloud 2.8T (MoE) 1M ~70 $3 $15
Qwen3.5 Alibaba 🦙 Ollama 397B (17B act) 1M 76.4* 52.5 Free Free
GLM-5.2 Zhipu AI 🦙 Ollama — (MoE) 1M 62.1 81.0 Free Free
Kimi K2.7 Code Moonshot AI 🦙 Ollama 1T (32B act) 256K ~55 Free Free
MiniMax M3 MiniMax 🦙 Ollama — (MoE) 1M 59.0 66.0 Free Free
Gemma 4 31B Google 🦙 Ollama 31B 128K ~35 Free Free
DeepSeek V4 Pro DeepSeek 🦙 Ollama 1.6T (49B act) 1M ~50 Free Free

* Qwen3.5 uses SWE-bench Verified where SWE-Pro unpublished. Claude Opus 4.8 (sources, May 2026): SWE-Bench Pro 69.2% · SWE-Verified 88.6% · Terminal-Bench 2.1 74.6% on Terminus-2 harness († harness-sensitive — GPT-5.5 Codex CLI can read higher on the same suite) · AA Intelligence Index 61.4 at launch max effort (‡ AA May 28 analysis; later AA index revisions may re-rank) · GDPval-AA Elo 1890 · OSWorld-Verified 83.4% · HLE no-tools 49.8% · list $5/$25 per M (fast mode $10/$50). No separate published SWE-Pro for xhigh — raise effort only when quality fails. Grok 4.5 (high): AA Intel 54 · Coding Agent ~76 · Terminal-Bench 2.1 83.3% · SWE-Pro 64.7% · ~$2/$6. Sol Coding Agent max ≈ 80; Fable SWE-Pro 80.3%. Context windows vary by product surface.

⚡ Reasoning Effort Ladders — Deep Research (July 2026)

Same model family, different intelligence / latency / cost trade-offs. Effort does not change the $/M token rate — it changes how many tokens the model spends (thinking, tool calls, verification). Rule of thumb: start medium/high; only raise to xhigh/max when quality fails after 2–3 attempts.

Key finding: pick tier first, then dial effort

Anthropic Fable 5 (Claude Max 5x): API default is high. Use xhigh only for long-horizon work (30+ min agents, multi-million token budgets). Dual-capable: plan and implement.

Claude Opus 4.8 (Claude Max 5x): Prior Claude flagship — plan, implement, adversarial review. Published peaks: SWE-Pro 69.2%, TB 2.1 74.6% (Terminus-2), AA Intel 61.4 (launch), $5/$25. Prefer high daily; raise for hard multi-file. Best second Claude when Fable authored the work.

OpenAI GPT-5.6 (ChatGPT Plus / API): Tiers Sol ($5/$30) · Terra ($2.50/$15) · Luna ($1/$6). Effort: none → low → medium (default) → high → xhigh → max. Ultra = multi-agent, not an effort level. Terra high = default implementer after a clear plan.

Grok 4.5 High (Grok / xAI): Your practical Grok setting. AA Intelligence ≈ 54 (high) · Coding Agent ≈ 76 in Grok Build · Terminal-Bench 2.1 83.3% · SWE-Pro 64.7% · SWE Marathon 29% (leads Fable/Opus on that long-horizon harness). List ~$2/$6 per M and very token-efficient (~⅓–½ the tokens of peers on AA tasks). Dual-capable for fast plan+code; verify factual claims (higher hallucination risk than Claude/GPT on pure QA).

Pareto note: Grok 4.5 often sits on the cost–capability frontier (~$0.31/AA task, ~$2.5/Coding Agent task in public AA writeups). Sol (max) Intel ≈ 59 vs Fable ≈ 60; Terra ≈ 55. Use Grok high for cheap near-frontier agents; escalate to Fable/Sol when SWE-Pro-class repo quality is mandatory.

1. Claude Fable 5 — Effort API

high Claude Fable 5 High (default)

$10 in / $50 out · 1M context · adaptive thinking always on
  • API default — same as omitting effort
  • Anthropic guidance: start here for most tasks
  • Almost always thinks deeply; strong coding + knowledge work
  • SWE-Bench Pro 80.3% · SWE-Verified 95%

xhigh Claude Fable 5 Extra High

Extended capability · long-horizon · multi-million token budgets
  • Always thinks deeply with extended exploration
  • Target: agentic/coding runs over 30 minutes
  • More tool calls, plans, self-verification, code comments
  • Set max_tokens high (start ~64k) — hard ceiling on think+text
  • Claude Code ultracode = xhigh + multi-agent permission (not a separate API level)
  • Cost: same $/token, often 2–4× more output tokens vs high

max Claude Fable 5 Max (reference)

No token constraints · absolute ceiling
  • AA Intelligence Index leader (~60); often overthinks structured tasks
  • Reserve for genuinely frontier problems only
  • Caveat: safeguard may fallback ~5–8% cyber/bio to Opus 4.8

2. GPT-5.6 Sol — Medium / High / Extra High

medium GPT-5.6 Sol Medium

$5 / $30 · API default · gpt-5.6-sol
  • OpenAI recommended balanced starting point
  • Agents' Last Exam: even at medium, beats Fable by ~11.4 pts @ ~¼ cost
  • Best for: interactive coding, design/UI, multi-turn product work
  • Latency-friendly vs high/xhigh; quality still frontier-class

high GPT-5.6 Sol High

$5 / $30 · hard reasoning · complex debugging
  • Use when quality > latency: deep plans, multi-file refactors, security reviews
  • Terminal / agentic coding climbs toward published Sol peaks
  • Eval both medium and high on your harness before locking defaults

xhigh GPT-5.6 Sol Extra High

Deep research · async agent runs · long tool loops
  • UI label "Extra High" = API reasoning.effort: "xhigh"
  • Target: long-running agentic workflows, deep research, verification-heavy tasks
  • Terminal-Bench 2.1 single-agent peak 88.8% (Ultra 4-agent: 91.9%)
  • AA Coding Agent Index (max) 80 🥇 · AA Intel (max) 59
  • SWE-Bench Pro 64.6% (Fable still leads repo-level Pro)
  • BrowseComp 92.2% · OSWorld 2.0 62.6% · ALE 53.6
  • Compare xhigh vs max on hard evals — max explores even longer

3. GPT-5.6 Terra — Medium / High / Extra High

medium GPT-5.6 Terra Medium

$2.50 / $15 · half Sol price · gpt-5.6-terra
  • Practical default for most teams after GPT-5.5
  • Agents' Last Exam ~50.4 family peak; medium stays close for everyday workflows
  • Strong enough for coding, reasoning, agents when cost matters
  • MRCR long-context ~89–90% (not Luna's 41% cliff)

high GPT-5.6 Terra High

$2.50 / $15 · quality step-up without Sol bill
  • Escalate when medium misses edge cases but Sol is overkill
  • Hard debugging, multi-step planning, higher-stakes knowledge work
  • Often within 2–3 pts of Sol on Terminal / Coding Agent at ~½ price

xhigh GPT-5.6 Terra Extra High

Long agentic Terra · sweet-spot escalations
  • AA Coding Agent Index (max) 77.4 ≈ Fable 77.2
  • DeepSWE 69.6% ≈ Fable 69.7%
  • Terminal-Bench 2.1 87.4% · AA Intel (max) 55
  • ~$0.55 / AA Intelligence task (~50% less than Sol max)
  • Best "budget frontier" for long-horizon engineering without Sol cost

AA Intelligence Index (max-effort published)

AA Coding Agent Index (max-effort)

Terminal-Bench 2.1 (%)

Cost per AA Intelligence Task ($)

Config API effort $/M in $/M out AA Intel Coding Agent Terminal 2.1 SWE-Pro ALE When to use
Claude Fable 5 xhigh xhigh $10 $50 ~60 77.2 ~84–86 80.3 40.5 Repo migrations, multi-day agents, deep research
GPT-5.6 Sol medium medium $5 $30 ~56 ~76 ~86 ~63 ~51+ Default Sol interactive work
GPT-5.6 Sol high high $5 $30 ~58 ~78 ~88 ~64 ~52 Hard debug, high-stakes plans
GPT-5.6 Sol Extra High xhigh→max $5 $30 59 80 88.8 64.6 53.6 Cyber, terminal SOTA, multi-agent Ultra
GPT-5.6 Terra medium medium $2.50 $15 ~52 ~74 ~85 ~61 ~49 Daily driver after GPT-5.5
GPT-5.6 Terra high high $2.50 $15 ~54 ~76 ~86 ~63 ~50 Escalate without Sol bill
GPT-5.6 Terra Extra High xhigh→max $2.50 $15 55 77.4 87.4 ~64 50.4 Budget long-horizon frontier

Sources: Anthropic effort docs + Fable 5 launch; OpenAI GPT-5.6 GA (Jul 9, 2026); Artificial Analysis Intelligence/Coding Agent Index & cost-per-task (Jul 2026); Vellum / independent writeups for Terminal-Bench & ALE. Medium/high midpoints are interpolated where labs only publish max-effort curves — treat as decision guide, re-eval on your harness.

Routing cheat-sheet

Default stack: GPT-5.6 Terra Medium → GPT-5.6 Terra High if quality slips → GPT-5.6 Sol High/xhigh for agentic/cyber/UI ceiling → Claude Fable 5 High/xhigh for SWE-Pro-class repo migrations & multi-day coding.

Cost trap: Fable xhigh and Sol max can burn 2–7× tokens vs medium/high on the same prompt. Measure thinking_tokens / reasoning tokens; raise max_tokens before blaming quality.

Split brain: Fable orchestrator + cheaper workers (Sonnet/Haiku or Terra) can hit ~96% of all-Fable BrowseComp quality at ~46% cost (Anthropic multi-agent guidance).

📖 Effort Guide — pick the right dial for the job

Effort does not change the $/million token price. It changes how much the model thinks, how many tools it uses, and how long it runs. Higher effort = more tokens, latency, and cost. Start low enough to finish fast; raise only when quality fails.

🧠 Plan / Brainstorm Use Fable high, Opus high, or Sol high. Not xhigh by default — interactive strategy rarely needs multi-hour exploration. These models are dual-capable (can implement later).
🛠️ Implement (normal) Default Terra high or Grok 4.5 high (your Grok plan — cheap dual). Or keep Fable/Opus/Sol high in a fresh session. Clear ticket, one PR.
🛠️ Implement (long / hard) Raise to xhigh (Fable, Opus, Sol, or Terra). Multi-file, multi-hour, migrations, agent loops that keep failing at high.
🔎 Review Plan review → Fable/Opus/Sol high. Code review → same families at high. Prefer a different model/session than the author (e.g. Fable wrote → Opus or Sol reviews).

❌ Why not Fable xhigh / Opus 4.8 xhigh / Sol Extra High for everything? → jump to answer · 🧭 How to choose

🧭 How to choose the right model

Stop guessing. Answer 3 questions (or use the cheat sheet). Two knobs matter: which model (Fable / Sol / Terra / Ollama) and how hard it thinks (medium · high · xhigh). Effort does not change $/M token price — it burns more tokens.

1. Model (who) Brain + price class: Fable · Sol · Terra · Ollama
2. Effort (how hard) medium · high (default work) · xhigh · max
Order Job type → size → effort. Never start with “strongest model.”

Four default roles

🧠 Planner
Sol high · Fable high
Brainstorm, architecture, “how should we build X?”
🛠️ Builder
Terra high (default)
Implement after the plan is clear
🔎 Reviewer
Fable/Sol high · other session
Review plan or PR — not the same chat that wrote it
🚀 Long agent / epic
Fable/Sol xhigh · budget Terra xhigh
Multi-hour, multi-file, migration, keep failing at high

Dual workflow: Sol/Fable plans → Terra builds → Sol/Fable reviews

❌ Why not always use the “best”? (Fable 5 xhigh · Opus 4.8 xhigh · Sol Extra High)

“Best model + max effort on everything” feels safe. It is usually slower, more expensive, and sometimes worse for small jobs. Max capability is a tool for hard, long jobs — not a default lifestyle. Direct link: #why-not-always-max

Claude Fable 5 xhigh for everything?

  • Same $/M, more tokens: price per million stays ~$10/$50, but xhigh often burns 2–4× tokens (thinking + tools). A rename can cost like a small feature.
  • Wrong job shape: Anthropic positions xhigh for long-horizon agent coding (often 30+ min, many tools). Most plans, reviews, and PRs are short interactive turns — high is the default and enough.
  • Over-scope: max agents often rewrite neighbors, invent “improvements,” and expand tickets you wanted small.
  • Latency: minutes of thinking for a 5-line CSS change wastes your day, not just money.

Claude Opus 4.8 xhigh for everything?

  • Not the daily driver in this stack: your comparison defaults to Fable 5 (newer coding line) + Sol/Terra. Opus 4.8 is still strong, but “always Opus xhigh” is the old “always strongest Claude” habit.
  • Same cost trap: top Claude + xhigh = highest token volume on top tier rates. Tiny work does not get smarter past a point — it just gets slower.
  • Quality plateau: for clear tickets (migration already planned, one PR, known stack), Terra high / Fable high often match or beat “max everything” because the task is execution, not deep research.
  • Special cases only: keep Opus/Fable max for hard reasoning, high-stakes design, or when Fable/Sol high already failed — not for every chat.

GPT-5.6 Sol Extra High (xhigh) for everything?

  • Flagship price × max effort: Sol is already ~2× Terra ($5/$30 vs $2.50/$15). Extra High multiplies tokens on top. 20 small tickets/day can torch quota.
  • Ceiling ≠ better on clear work: for “implement this approved plan,” Terra high is the designed default. Sol Extra High shines on long agents, hard debug, multi-angle verify — not renames.
  • OpenAI’s own ladder: raise effort when evals show a gain — don’t jump to Extra High by default. Try medium/high first; xhigh when high fails or the job is multi-hour.
  • Not always best at repo migrations: Fable often leads pure SWE-Pro-style long coding. Sol xhigh is not automatically #1 at everything.
What “always max” actually costs you
  • Money — 2–7× tokens vs medium/high on the same prompt class
  • Time — you wait; interactive flow dies
  • Quality on small jobs — over-edit, scope creep, harder reviews
  • Quota — no budget left for the migration that actually needs xhigh
  • Blind spots — same max model plans, codes, and self-reviews
Better habit
  • Default high (or Terra high for implement), not xhigh
  • Use xhigh only when: multi-file · multi-hour · migration · high already failed 2–3×
  • Split roles: plan model ≠ implement model ≠ review model
  • Save Fable/Opus/Sol Extra High for the hard 10% of work
  • Tiny work → Terra medium or local Ollama ($0)
Bottom line: Fable xhigh, Opus 4.8 xhigh, and Sol Extra High are peak tools, not daily coffee. Always-max does not mean always-best — it means always-expensive, often-slow, sometimes-worse. Match the model to the job; escalate only when quality fails.

Interactive: answer 3 questions

1 Is the problem still unclear? (no plan, open design, “how should we…?”)
Your pick

Cheat sheet (print this)

Situation Use Why
Unclear / design / “how should we…?” Sol high or Fable high Interactive strategy; high is enough
Plan / architecture Sol high · Fable high Not xhigh by default
Review plan or PR Fable/Sol high · other session Author ≠ reviewer
Clear implement (normal PR) Terra high Default builder; ~½ Sol price
XS typo / rename / 1-line Terra medium Don’t burn high/xhigh
Long multi-file / migration Fable/Sol xhigh · budget Terra xhigh Long-horizon agent coding
Docs / release notes Terra medium/high Known facts → cheap work
Local / privacy / offline GLM-5.2 code · Qwen3.5 plan Free; escalate cloud if hard
Stuck after 2–3 fails Raise one step only Terra→Sol high→Fable/Sol xhigh
Effort: high vs xhigh
  • medium — small obvious edits, volume work
  • high — default for plan, review, normal PRs
  • xhigh — long agent: many files, 30+ min tools, migrations
  • max — only if xhigh still fails

Plan = high is not “plan is less important.” xhigh is for long coding agents, not every design chat.

Price intuition
  • Terra — everyday implementer · ~½ Sol
  • Sol — flagship plan / hard agent · ~2× Terra
  • Fable — deep plan + long SWE; xhigh burns 2–4× tokens vs high
  • Ollama — free drafts; below cloud frontier
Escalate ladder (when stuck)
  • After 2–3 fails, raise one step only
  • Terra high → Sol high → Fable/Sol xhigh
  • Also rewrite the prompt (repro, constraints, acceptance)
  • Don’t jump medium → max
Don’t do this
  • Fable xhigh “because ERP is serious”
  • Sol high for every rename
  • Plan + implement + review in one session/same model
  • Code while architecture is still open (e.g. DMS vs SAP)
One-line: Unclear/design → Sol or Fable high · Clear implement → Terra high · Review → Sol/Fable high (other session) · Tiny fix → Terra medium · Long multi-file → Fable/Sol xhigh · Stuck 2–3× → raise one level only.

↑ Back to: Why not always Fable/Opus xhigh / Sol Extra High?

Why Plan = Fable high, but long Implement = Fable xhigh?

Not because plan is less important. Anthropic’s xhigh is built for long-horizon agentic coding (often 30+ min, many tools, self-verify). Planning is usually short interactive turns — high is the API default and already strong. Using xhigh on every plan wastes 2–4× tokens. Using only high on a day-long migration under-explores. Rule: Fable high for plan & normal code; Fable xhigh when the coding job is multi-hour / multi-file / migration-class.

🔄 Software development lifecycle → which model

Follow the flow left → right. Author ≠ reviewer when quality matters. Effort: start medium/high; raise to xhigh only on long hard jobs.

1
Discover & Brainstorm
Claude Fable 5 high
GPT-5.6 Sol high
Ollama: Qwen3.5
2
Plan & Architecture
Claude Fable 5 high
GPT-5.6 Sol high
not xhigh by default
3
Review the Plan
Claude Fable 5 high*
GPT-5.6 Sol high*
Ollama Qwen3.5
*other session
4
Implement Code
GPT-5.6 Terra high
Claude Fable 5 high
Ollama GLM-5.2
xhigh if long job
5
Review Code / PR
Claude Fable 5 high*
GPT-5.6 Sol high*
GPT-5.6 Terra high
Ollama GLM-5.2
6
Test & Debug
GPT-5.6 Terra high
GPT-5.6 Sol high
Claude Fable 5 high
Ollama GLM-5.2
7
Docs & Release
GPT-5.6 Terra medium
GPT-5.6 Sol medium
Ollama Qwen3.5
8
Maintain & Hotfix
GPT-5.6 Terra high
GPT-5.6 Sol high
Claude Fable 5 xhigh
↻ Loop: production feedback → back to 1 Discover or 2 Plan. Long migrations / flaky CI nights → jump to Fable/Sol xhigh or Terra xhigh/max.
# Phase What you do Cloud model + effort Ollama / free
1 Discover & Brainstorm Ideas, problem framing, options, spikes Sol high · Fable high Qwen3.5
2 Plan & Architecture ADRs, APIs, data model, milestones, risks Sol high · Fable high · Terra high budget Qwen3.5
3 Review the Plan Adversarial check: gaps, failure modes, scope Fable high* · Sol high* (*≠ author) Qwen3.5
4 Implement Write code, PR features from approved plan Terra high default · Fable/Sol high · xhigh if multi-file / multi-hour GLM-5.2 · MiniMax M3
5 Review Code Bugs, security, tests, plan match Fable high* · Sol high* · Terra high routine GLM-5.2
6 Test & Debug Failing tests, flaky CI, root-cause Terra high · Sol high hard · Fable high/xhigh deep GLM-5.2 terminal
7 Docs & Release Changelog, runbooks, release notes, README Terra medium/high · Sol medium Qwen3.5
8 Maintain & Hotfix Small fixes daily; big incidents escalate Terra medium/high · Sol high incidents · Fable xhigh major outage GLM-5.2 · DeepSeek
Text flow (copy-friendly)
Discover (Sol/Fable high) → Plan (Sol/Fable high) → Review plan (other Sol/Fable high | Qwen) → Implement (Terra high | Fable/Sol high | xhigh if long) → Review code (Fable/Sol high | GLM) → Test/Debug (Terra/Sol high | GLM terminal) → Docs/Release (Terra med/high) → Maintain (Terra med/high) ↻ [outage → Sol/Fable xhigh]

⚡ Quick map: task → model + effort

Real situation Best pick Why
“Help me design auth for a SaaS” Sol high · Fable high Interactive strategy; high is enough
“Review this PR (200 lines)” Terra high · Sol high · GLM-5.2 Code critique, not multi-hour agent
“Implement the checkout ticket” Terra high (default) Clear scope → execute cheaply
“Migrate Express → Nest across 40 modules” Fable xhigh · Sol xhigh Long-horizon coding = xhigh territory
“Fix typo / rename variable” Terra medium · Sol medium Don’t burn high/xhigh on trivial work
“Red-team this architecture for failures” Sol high / xhigh · Fable high Adversarial planning skill

1. Claude Fable 5 — Medium · High · XHigh · Max

Price always $10/$50 per M. Effort only changes token volume. Anthropic default = high.

medium Claude Fable 5 Medium

When: routine Claude work where high feels slow/expensive but you still want Fable quality above older models.
Real scenario Summarize a 30-page RFC and list open questions for the team standup. Or rewrite a product FAQ in clearer English.
⚠ Avoid for multi-file coding agents or deep architecture debates.

high Claude Fable 5 High (default)

When: planning, brainstorming, plan review, normal PRs, design docs, most “think carefully” work. Start here.
Real scenario “Design multi-tenant billing for our SaaS.” “Review this plan for a mobile offline-first mode.” “Implement the profile settings API from this ticket.”
⚠ If the agent runs 1+ hours and keeps missing edge cases, step up to xhigh — don’t stay stuck on high forever.

xhigh Claude Fable 5 Extra High

When: long-horizon coding, big migrations, multi-repo work, agents that need deep self-verify (often 30+ min). This is why implement ≠ always high.
Real scenario “Migrate our 50-module Express app to NestJS with tests green.” “Refactor auth across frontend + backend + workers without downtime.” Stripe-style codebase-wide migrations.
⚠ Don’t use as daily chat default — 2–4× more tokens than high on many tasks.

max Claude Fable 5 Max

When: rare frontier problems where xhigh still fails and money is secondary. Absolute no cap on thinking spend.
Real scenario Novel distributed systems design with no internal playbook; or a one-shot eval where you must maximize score regardless of cost.
⚠ Often overthinks structured tasks. Never leave max as a standing default.

2. GPT-5.6 Sol — Medium · High · Extra High (xhigh)

Flagship OpenAI tier · $5/$30 per M · API default effort = medium. UI “Extra High” = xhigh. Also has max above xhigh for hardest jobs.

medium GPT-5.6 Sol Medium (API default)

When: interactive Sol work — design chat, product Q&A, medium coding, drafting UI copy. Already beats Fable on Agents’ Last Exam at much lower cost.
Real scenario “Suggest 3 IA structures for a B2B dashboard.” “Write the React component for this Figma button row.” Pair-programming in Codex on a feature branch.
⚠ If answers stay shallow on hard design/debug, step to high.

high GPT-5.6 Sol High (sweet spot)

When: hard / unclear / multi-system problems. Best community “daily driver” on the intelligence/price curve for Sol. Prefer for brainstorm + plan + serious review.
Real scenario “Our webhook pipeline drops events under load — diagnose and propose a redesign.” “Compare event-sourcing vs outbox for orders.” Deep PR review of a payment refactor.
⚠ Don’t jump straight to xhigh; OpenAI says raise only when evals show a measured gain.

xhigh GPT-5.6 Sol Extra High

When: long agent runs, terminal-heavy workflows, cyber-ish investigation, multi-angle verification, deep research. Coding Agent peak territory (~80).
Real scenario Overnight Codex agent: “bring green CI on this flaky monorepo.” Security-style “find auth bypass paths in this service.” Multi-hour research brief with sources.
⚠ Quota burner if always on. Compare with max only if xhigh still fails.

3. GPT-5.6 Terra — Medium · High · Extra High · Max

Balanced tier · $2.50/$15 per M (~½ Sol). Best everyday implementer once the plan is clear. Still supports medium → max; use Sol xhigh when you need the absolute ceiling.

medium GPT-5.6 Terra Medium

When: volume work — small tickets, renames, boilerplate, drafts. Post-GPT-5.5 daily driver for cost-aware teams.
Real scenario “Add a nullable column + migration.” “Write unit tests for this pure function.” “Convert this CSS module to Tailwind.”
⚠ If the PR needs real design judgment or multi-service changes, step to high.

high GPT-5.6 Terra High (implement default)

When: normal feature implementation after plan is written. Best “Terra executes” slot in the dual workflow (Sol plans → Terra builds).
Real scenario “Implement checkout steps 1–3 from the approved plan, with tests.” “Wire Stripe webhook + idempotency keys.” Day-to-day Codex feature work.
⚠ Open-ended product strategy → use Sol high first, not Terra as the only brain.

xhigh GPT-5.6 Terra Extra High

When: long implement jobs on a budget before paying Sol. Multi-file agents, larger refactors; Coding Agent ~77.4 ≈ Fable coding agent score.
Real scenario “Upgrade Next.js across the monorepo and fix all type errors.” “Split the admin app into packages without breaking CI.”
⚠ Need absolute terminal/cyber/UI ceiling? Jump to Sol xhigh instead of overpaying Terra forever.

max GPT-5.6 Terra Max

When: you want Terra’s higher ceiling without Sol’s full price; quality-first on a hard implement that still fits Terra’s role (execute a defined plan).
Real scenario Hard data migration with rollback plan already written; you need maximum care applying it, not a new product strategy.
⚠ If the problem is still ambiguous, max Terra won’t replace Sol high for planning.

🗓️ Sample day (real workflow)

Time You do Model + effort
09:00 Brainstorm checkout redesign with PM notes GPT-5.6 Sol high · or Fable high
10:00 Adversarial review of the plan (“find 5 failure modes”) Fable high / Sol high · different session · Ollama Qwen3.5
11:00–16:00 Implement tickets from the approved plan Terra high (default) · Terra xhigh if multi-file slog
16:30 PR review before merge Sol high or Fable high · or GLM-5.2 local
Night agent Long migration / flaky CI green-up Fable xhigh · Sol xhigh · (budget: Terra xhigh/max)

🔧 Tweaks & change requests — which model?

After the big SDLC loop (brainstorm → plan → implement → review → test → docs → maintain), most real work is incoming changes: bugs, small tweaks, medium features, big features. Size the change first, then pick model + effort. Do not default to Claude Fable 5 xhigh or GPT-5.6 Sol high for everything.

❌ Why not Fable / Opus xhigh for all?

  • Cost: Same top $/M, but xhigh often uses 2–4× more tokens. A typo fix can cost like a small feature.
  • Latency: Minutes of thinking for a 5-line CSS change wastes your day.
  • Over-engineering: Max agents expand scope — rewrite neighbors, “improve” unrelated code.
  • Wrong tool shape: xhigh = long-horizon agents (30+ min). Most tickets are not that. Opus 4.8 xhigh has the same trap as Fable xhigh.
  • Review bias: Same max model that wrote the code often rubber-stamps itself.

❌ Why not Sol high / Extra High for all?

  • Price vs Terra: Sol is ~2× Terra ($5/$30 vs $2.50/$15). Routine PRs rarely need flagship — Extra High multiplies that again.
  • Still not free: High/xhigh multiplies tokens. 20 small tickets/day on Sol Extra High burns quota fast.
  • Ceiling ≠ always better: Clear tickets → Terra high matches quality closely at half cost. Sol Extra High for long agents / hard multi-system work.
  • Repo migrations: Fable often still leads pure SWE-Pro-style work — Sol xhigh is not automatically best at everything.
  • Local free option: Ollama GLM-5.2 / Qwen3.5 absorb many tweaks at $0 API cost.

Change size ladder (issue → model)

XS · Tiny

Typo / copy / 1-line fix

Label text, CSS spacing, rename one variable, wrong import path, comment fix. <15 min human work.
Use: GPT-5.6 Terra medium · Ollama DeepSeek V4 Pro / Gemma 4 31B
Optional review: skip or quick human glance.
⚠ Skip Fable xhigh / Sol high — pure waste.
S · Small tweak

Bugfix / small feature tweak

1 file (maybe 2–3). Known root cause or clear ticket: “button disabled when form empty”, “wrong currency format”, “add filter dropdown option”. ~30–90 min.
Use: GPT-5.6 Terra high · Ollama GLM-5.2
Review: Terra high other session or GLM-5.2 · human for prod-critical.
⚠ Sol high only if root cause is unknown after one Terra pass.
M · Medium feature

New medium feature (scoped)

New endpoint + UI section, new table + CRUD, OAuth login for one provider, export CSV. Spec mostly known. Half-day to 2 days.
Plan (short): GPT-5.6 Sol high or Claude Fable 5 high
Implement: GPT-5.6 Terra high (→ Terra xhigh if multi-file slog)
Review: Sol/Fable high* · Ollama GLM-5.2
⚠ Don’t implement medium features on Fable xhigh by default — plan high, build Terra high.
L · Large feature

Big feature / multi-system change

Checkout redesign, multi-tenant billing, new mobile offline mode, permissions overhaul. Touches many modules, needs architecture choices. Days to weeks.
Brainstorm/Plan: GPT-5.6 Sol high · Claude Fable 5 high
Review plan: other Sol/Fable high · Ollama Qwen3.5
Implement: Terra high/xhigh · Fable/Sol high (xhigh if multi-day agent)
Review code: Fable/Sol high*
⚠ Sol high for all phases is OK quality-wise but 2× Terra on implement volume — split plan vs build.
XL · Epic / migration

Epic, rewrite, migration, major incident

Framework upgrade monorepo, Express→Nest, data migration with rollback, production outage root-cause + fix. Multi-day, high risk.
Plan: Claude Fable 5 / GPT-5.6 Sol high (→ xhigh deep research)
Implement agent: Claude Fable 5 xhigh · GPT-5.6 Sol xhigh · budget: Terra xhigh/max
Review: Fable/Sol high* · never same agent session only
⚠ Here xhigh earns its keep — still not for every tiny ticket next week.
Issue · Unclear

User report: “it’s broken” / vague request

No repro steps, “make it better”, “like competitor X”. Need diagnosis before coding.
Investigate first: GPT-5.6 Sol high · Claude Fable 5 high
Write repro + acceptance criteria → then route as S/M/L above. Do not jump to implement on Terra until the bug is sized.
⚠ Terra medium on vague issues → wrong fix, thrash.
User says… Size Model + effort Skip
“Fix typo on checkout button” XS GPT-5.6 Terra medium Claude Fable 5 xhigh · GPT-5.6 Sol high
“Date format wrong in VN locale” S GPT-5.6 Terra high · Ollama GLM-5.2 Claude Fable 5 xhigh
“Add export invoices to CSV” M Plan Sol/Fable high → Implement Terra high All-in Fable xhigh
“Rebuild permissions for multi-tenant orgs” L Sol/Fable high plan → Terra/Fable high–xhigh build Terra medium alone
“Migrate whole API to NestJS” XL Claude Fable 5 xhigh · GPT-5.6 Sol xhigh GPT-5.6 Terra medium
“Site is slow / something’s wrong” Unclear GPT-5.6 Sol high investigate first Blind implement on any model
“Make UI prettier like Linear” M–L GPT-5.6 Sol high (design) → Terra high implement Claude Fable 5 xhigh first pass
“Prod down after deploy” XL incident Sol/Fable high now → xhigh if multi-service Waiting for perfect cheap model
Decision recipe (30 seconds):
1) Is the request clear? No → Sol/Fable high diagnose.
2) How many files / risk? 1 file → Terra medium/high. Many modules / data / auth → Sol/Fable high plan, then Terra/Fable implement.
3) Will an agent run >30–60 min? → xhigh (Fable/Sol/Terra), not high forever.
4) Always ask: “Would a mid-level eng need a full day?” If no, don’t pay Fable xhigh / Sol high.

Remember

Start at the lowest effort that still succeeds. Escalate one step at a time (medium → high → xhigh → max). Prefer a second model for review. Plan with high; long implement with xhigh — not the other way around. For tweaks: size the change first — most issues are S/M, not Fable xhigh.

🃏 Model Cards — strengths at a glance

Each card is a configuration you can actually pick (model + effort where it matters). Order: Fable → Sol → Terra → other cloud → Ollama. Hover a top-right chip to highlight a family. Dual-capable cloud flagships (Fable, Opus, Sol, Terra) appear for both plan and code — specialists live mainly under Ollama.

C5
Claude Fable 5 high
Anthropic · Mythos-class · default
S-TIER
Context1M tokens
SWE-Bench Pro80.3% 🥇
Intel Index~60–65
Pricing$10/$50 per M
🏆 Best Coding (SWE-Pro) 🏆 Best Long-Horizon API default effort Vision SOTA $$$ Expensive Safeguard Fallback
Best for: Start here for most Fable work. Large-scale code migrations, multi-day agentic sessions, deep research. Raise to xhigh only when evals show headroom.
X5
Claude Fable 5 xhigh
Anthropic · extended exploration
S-TIER+
Context1M tokens
Horizon30+ min runs
Token budgetMillions
Pricing$10/$50 · volume↑
🏆 Deepest Fable thinking Self-verify / reflect More tool calls Claude Code ultracode 2–4× output tokens Needs high max_tokens
Best for: Capability-sensitive agentic coding, iterative tool use, detailed research, multi-hour autonomous loops. Always thinks deeply with extended exploration. Not for routine chat — use high/medium instead.
SM
GPT-5.6 Sol medium
OpenAI · flagship default
S-TIER
Context1M tokens
ALE (medium)beats Fable
Intel (est.)~56
Pricing$5/$30 per M
🏆 API default Sol 🏆 ALE @ ¼ Fable cost Design/UI Interactive coding Programmatic tools Still premium tier
Best for: Default Sol setting. Knowledge work, design/frontend, multi-turn coding. Medium reasoning already beats Claude Fable on Agents' Last Exam at roughly one-quarter estimated cost.
SH
GPT-5.6 Sol high
OpenAI · hard reasoning
S-TIER
Context1M tokens
FocusQuality > latency
Intel (est.)~58
Pricing$5/$30 per M
Hard debug / plans Security reviews Multi-file refactors Cyber workflows More latency & tokens
Best for: When medium misses hard cases: complex debugging, deep planning, high-value agent steps. OpenAI: evaluate medium vs high on representative workloads before locking defaults.
SX
GPT-5.6 Sol xhigh
OpenAI · Extra High / near-max
S-TIER+
Coding Agent80 AA 🥇
Terminal 2.188.8%
Intel (max)59
Pricing$5/$30 · ~$1.04/task
🏆 Coding Agent SOTA 🏆 Agents' Last Exam 53.6 🏆 BrowseComp 92.2% ExploitBench cyber Ultra multi-agent SWE-Pro 64.6% < Fable
Best for: Deep research, async agentic runs, terminal SOTA, cybersecurity, computer use. Pair with Ultra (4 agents) for +3pts Terminal at ~3× cost. Compare xhigh vs max on your hardest evals.
TM
GPT-5.6 Terra medium
OpenAI · balanced daily driver
A+/S- value
Context1M tokens
vs Sol price½ cost
Intel (est.)~52
Pricing$2.50/$15 per M
🏆 Team default Everyday coding Agentic OK MRCR ~90% Below Sol ceiling
Best for: Practical center of GPT-5.6. Most product engineering after GPT-5.5. Within 2–3 pts of Sol on many agentic benches at half the price.
TH
GPT-5.6 Terra high
OpenAI · quality step-up
A+/S- value
Context1M tokens
RoleEscalate path
Intel (est.)~54
Pricing$2.50/$15 per M
No-Sol escalate Harder plans Debug step-up Cost control Still not Sol ceiling
Best for: When Terra Medium quality slips but you refuse Sol spend. Harder reasoning, multi-step agents, higher-stakes knowledge work at Terra rates.
TX
GPT-5.6 Terra xhigh
OpenAI · budget frontier
S- value
Coding Agent77.4 ≈ Fable
DeepSWE69.6% ≈ Fable
Intel (max)55
Pricing$2.50/$15 · ~$0.55/task
🏆 Best $/long-horizon ≈ Fable Coding Agent Terminal 87.4% ALE 50.4 AA: Luna/Sol may dominate Pareto
Best for: Long-horizon engineering without Sol bill. Nearly ties Fable on Coding Agent Index & DeepSWE at a fraction of cost. Best "extra high" escalate before jumping to Sol.

4. Other cloud + Ollama models

G4
Grok 4.5 high
xAI · Grok app / Grok Build · your Grok plan
A-TIER
Context500K–1M*
SWE-Bench Pro64.7%
Coding Agent~76
Terminal-Bench83.3%
Intel Index54 (high)
Pricing~$2/$6 per M
🏆 Token + $ efficiency 🏆 SWE Marathon lead 29% Terminal 83.3% near SOTA Dual plan + code (high) Fast · low tokens/task Verify facts (QA risk) Trails Fable on SWE-Pro
Best for: Your Grok plan default. Fast dual plan+implement on a budget, agentic terminal work, cheap overnight coding agents (Grok Build). Escalate to Fable/Sol when you need max SWE-Pro repo quality or careful adversarial review. Always fact-check knowledge claims.
K3
Kimi K3
Moonshot AI · 2.8T params
A-TIER
Context1M tokens
Architecture2.8T MoE, KDA
MultimodalText + Image
Pricing$3/$15 per M
🏆 Best Value Cloud Complex Coding Long-Horizon Agent Repo Navigation Open-Weight Debugging w/ Images
Best for: Complex coding with large repos, long-horizon agentic workflows, debugging with visual feedback. Best price-performance among cloud models.
Q3
Qwen3.5
Alibaba · 397B (17B act)
A-TIER
Context1M tokens
SWE Verified76.4%
VisionNative VL
PricingFREE (Ollama)
🏆 Best Open Vision 🏆 Multilingual (201 langs) Strong Agent BrowseComp Lead 17B Active (Fast) Document Understanding
Best for: Multimodal tasks, multilingual work, visual document understanding, search agents, computer use. Best open-weight vision-language model.
GL
GLM-5.2
Zhipu AI · MIT License
A-TIER
Context1M tokens (solid)
Terminal-Bench81.0% 🥇
SWE-Bench Pro62.1%
PricingFREE (Ollama)
🏆 Best Terminal Agent 🏆 Long-Horizon Open IndexShare Arch Effort Level Control MIT License FrontierSWE #2
Best for: Terminal-based coding, long-horizon engineering (hours to days), post-training ML research, open-source agentic coding. #1 open model for terminal tasks.
K2
Kimi K2.7 Code
Moonshot AI · 1T (32B act)
B-TIER
Context256K tokens
MLS Bench Lite35.1
Agentic +10%vs K2.6
PricingFREE (Ollama)
Coding-Focused End-to-End Tasks Thinking Mode Multimodal Input 256K Context Smaller than K3
Best for: Focused coding tasks, end-to-end programming, multi-turn debugging. Good local coding assistant when K3 is overkill.
M3
MiniMax M3
MiniMax · MSA Architecture
A-TIER
Context1M tokens
SWE-Bench Pro59.0%
Terminal-Bench66.0%
PricingFREE (Ollama)
🏆 Claw-Eval 74.5% 🏆 Native Multimodal MSA Sparse Attention Computer Use Paper Reproduction CUDA Kernel Opt
Best for: Autonomous agents, computer use, paper reproduction, CUDA optimization, multimodal coding. Only open model with coding + 1M + native multimodal.
G4
Gemma 4 31B
Google · 31B params
B-TIER
Context128K tokens
ELO Rating1452
MultimodalText+Vision+Audio
PricingFREE (Ollama)
Small & Fast Audio Input #3 Open Model Google Ecosystem 128K Context 31B Only
Best for: Lightweight local inference, audio+text tasks, quick prototyping, resource-constrained environments. Best small open model.
D4
DeepSeek V4 Pro
DeepSeek · 1.6T (49B act)
B-TIER
Context1M tokens
Architecture1.6T MoE
Cost EfficiencyBest-in-class
PricingFREE (Ollama)
🏆 Cost Efficiency Agentic Coding Long Context Open Source SOTA Arena Underwhelms Mid-tier Coding
Best for: Cost-sensitive deployments, long-context tasks, general knowledge work. Best cost-performance ratio among all models.

🎯 Task Suitability Matrix

Qualitative fit by task type (not a single benchmark number). Green does not mean “always use xhigh” — use the Effort Guide for dialing. Hover Fable / GPT-5.6 / Ollama chips to isolate columns. Opus is dual-capable like Fable but is grouped under Claude in narrative tabs; matrix columns list primary published configs.

= Excellent   = Good   = Adequate   = Limited

👤 Your Use Case: Plan · Implement · Review

You're currently on DeepSeek V4 Pro. Four task modes: brainstorm/plan, implement code, review the plan, review the code — different roles and defaults, not always different models.

✅ Cloud flagships are dual-capable (Plan + Implement)

Yes — the same cloud model can plan and implement. Both columns list Fable, Opus, Sol, Terra, and Grok 4.5 for a reason: they are generalists. What changes is metric, rank, and default role.

Model (your access) Can plan? Can implement? Best default role
Claude Fable 5 high/xhigh · Max 5x ✓ Excellent ✓ Excellent (SWE-Pro) high plan + normal PR · xhigh migrations
Claude Opus 4.8 high · Max 5x ✓ Excellent ✓ Excellent Plan + review + hard code · second Claude after Fable
GPT-5.6 Sol high/xhigh · ChatGPT Plus ✓ Excellent ✓ Excellent high plan/debug · xhigh long agent / terminal
GPT-5.6 Terra high/xhigh · ChatGPT Plus ✓ Good ✓ Excellent ($/PR) high default implement · xhigh long multi-file
Grok 4.5 high · Grok plan ✓ Strong / fast ✓ Strong agent (CA~76 · TB 83.3%) Cheap dual plan+code · overnight agents · fact-check knowledge

Where columns differ: (1) metric — Intelligence vs Coding Agent / SWE-Pro · (2) defaults — Plan leans Fable/Opus/Sol high; Implement leans Terra/Grok high · (3) Ollama Cloud split — Qwen plans, GLM codes (not dual).

🪪 Your subscription stack (how to use what you already pay for)

Claude Max 5x
Fable 5 + Opus 4.8
Plan / hard review / migrations (Fable high→xhigh). Opus high for second-opinion plan & code review.
ChatGPT Plus
GPT-5.6 Sol · Terra
Sol high plan/debug/UI. Terra high default implement. Sol xhigh for terminal/agent ceiling.
Grok
Grok 4.5 High
Fast dual plan+code, cheap agents (CA~76, TB 83.3%). Prefer for volume work; escalate Claude/GPT when quality stalls. Fact-check.
Ollama Cloud
Qwen3.5 · GLM-5.2 · others
Privacy / free-ish volume. Qwen = plan & docs. GLM = terminal coding. Not dual — pick by skill.

Recommended daily loop on your plans: (1) Plan on Fable high or Sol high · (2) Optional adversarial review Opus high or Sol · (3) Implement Terra high or Grok 4.5 high (volume/cheap) · hard multi-file → Fable/Sol xhigh · (4) Review ≠ author · (5) Ollama Cloud for offline drafts / terminal GLM.

🧠 Planning & Brainstorming

Architecture design, creative thinking, research, strategy

Default effort: Fable/Opus/Sol high (not xhigh). Same model can later implement — prefer a fresh session for build.

☁️ Cloud ranking — Planning / Reasoning (AA Intelligence)
62
Claude Fable 5 xhigh
61
Claude Opus 4.8 xhigh
59
GPT-5.6 Sol xhigh
58
GPT-5.6 Sol high
56
GPT-5.6 Sol medium
55
GPT-5.6 Terra xhigh
53
GPT-5.6 Terra medium
58
Kimi K3
54
Grok 4.5 high
AA Intelligence Index · Grok 4.5 high = 54 (AA) · dual-capable
👑
Claude Fable 5 high · dual Plan+Code
Deep strategy · can implement later (high/xhigh) · prefer new session for build
~60–65
2
Claude Opus 4.8 high · dual Plan+Code
Strong all-in-one · plan then implement with same model if you want
~59–61
3
GPT-5.6 Sol high/xhigh · dual
ALE/design lead · also top coding agent (same model both roles)
~58–59
4
GPT-5.6 Sol high
Hard plans / high-stakes brainstorming
~58
5
GPT-5.6 Sol medium
Default Sol · interactive strategy sessions
~56
6
GPT-5.6 Terra xhigh
Budget frontier planning · ½ Sol price
55
7
GPT-5.6 Terra high / medium
Daily driver for product planning
~52–54
8
Grok 4.5 high · dual · your Grok plan
Fast cheap strategy · AA Intel 54 · verify facts
54
9
Kimi K3
Open-weight cloud · long context
~58
🦙 Ollama Cloud / Local
55
Qwen3.5
52
MiniMax M3
50
DeepSeek V4 Pro
48
GLM-5.2*
45
Gemma 4 31B
42
Kimi K2.7 Code
Est. Intelligence Index (*GLM is coding-specialist, not best planner)
👑
Qwen3.5
397B MoE · 201 langs · vision · best free brainstormer
~55
2
MiniMax M3
Native multimodal · computer use · agentic workflows
~52
3
DeepSeek V4 Pro
Your current model · solid generalist · cost-efficient
~50
4
GLM-5.2*
Coding specialist — weak for open brainstorm; use for implementation
~48
5
Gemma 4 31B
Lightweight local · 128K context only
~45
6
Kimi K2.7 Code
Code-focused · 256K context — not for long planning
~42
🤔 GLM-5.2 ranks mid for planning — it's a coding specialist, not a general reasoner.
Built for terminal agents & long-horizon engineering, not open-ended brainstorming.
💡 Why Qwen3.5 > GLM-5.2 for planning?
Skill GLM-5.2 Qwen3.5
General knowledge 🟡 🟢
Multilingual (201 langs) 🔴 🟢
Search / Research agent 🟡 🟢
Creative brainstorming 🟡 🟢
Vision understanding 🔴 🟢
Terminal coding 🟢🟢🟢 🟡

GLM-5.2 was RL-trained on coding environments. Qwen3.5 was RL-trained across all domains — reasoning, search, vision, agents, math. Breadth wins for brainstorming.

🏆 Cloud picks for planning
  • Claude Fable 5 high — default strategy · dual-capable (can implement later in new session) ($10/$50)
  • Claude Opus 4.8 high — dual-capable · plan + implement + review · great all-in-one Claude
  • GPT-5.6 Sol high (→ xhigh long agent) — dual-capable · ALE/design lead · also codes well ($5/$30)
  • GPT-5.6 Terra high — OK for clear product plans · shines more on implement (½ Sol) ($2.50/$15)
  • Grok 4.5 high — your Grok plan · fast dual plan+code · AA 54 · fact-check knowledge (~$2/$6)

↪ After plan is approved: keep dual model (Fable/Opus/Sol/Grok) in a new session, or switch to Terra high / Grok high to save cost.

🦙 Best Ollama: Qwen3.5
  • 397B params (17B active) — largest open reasoning model
  • ▸ 1M context, 201 languages, native vision
  • ▸ Strong on BrowseComp, search agents, general knowledge
  • ▸ DeepSeek V4 Pro is also solid — your current model is decent here!

🛠️ Coding Implementation

Writing code, debugging, refactoring, terminal agents

Default effort: Terra high for clear PRs. Opus / Sol / Fable high also implement well (dual). Raise xhigh only for multi-file / multi-hour jobs.

☁️ Cloud ranking — Coding / Agents (not the same order as Planning)
80
GPT-5.6 Sol xhigh
78.5
Claude Opus 4.8 xhigh
78
GPT-5.6 Sol high
77.4
GPT-5.6 Terra xhigh
77.2
Claude Fable 5 xhigh
76
GPT-5.6 Sol medium
75
GPT-5.6 Terra medium
~70
Kimi K3
76
Grok 4.5 high
Coding Agent Index · Grok Build ~76 · Terminal 83.3% · Fable SWE-Pro still 80.3% on pure repo
👑
GPT-5.6 Terra high · dual OK · best default PR
Everyday implement after plan is clear · ½ Sol price
CA ~75
2
GPT-5.6 Sol xhigh · dual
Ceiling: Coding Agent 80 · Terminal 88.8% · also plans well
CA 80
3
Claude Opus 4.8 high · dual Plan+Code
FrontierSWE lead · same model can plan then implement
CA ~78.5
4
Claude Fable 5 xhigh · dual
SWE-Pro king · migrations · high is enough for normal PRs
Pro 80.3
5
GPT-5.6 Sol high · dual
Hard debug / multi-file · also strong planner
CA ~78
6
GPT-5.6 Terra xhigh · dual long jobs
Long multi-file on budget · CA 77.4 ≈ Fable agent
CA 77.4
7
Grok 4.5 high · dual · your Grok plan
CA~76 · Terminal 83.3% · cheap agents · SWE Marathon lead · fact-check
CA 76
8
GPT-5.6 Sol medium
Interactive coding · design/UI
CA ~76
9
Kimi K3
Strong code · open-weight cloud
Pro ~70
🦙 Ollama Cloud / Local
81
GLM-5.2
76*
Qwen3.5
66
MiniMax M3
~55
Kimi K2.7 Code
~50
DeepSeek V4 Pro
~35
Gemma 4 31B
Terminal-Bench 2.1 (%) — GLM-5.2 leads · *Qwen SWE-Verified
👑
GLM-5.2
Terminal-Bench 81% · best local coding agent · MIT
TB 81
2
Qwen3.5
Strong SWE-Verified · vision · multilingual codebases
76*
3
MiniMax M3
Computer use · autonomous agents · native multimodal
TB 66
4
Kimi K2.7 Code
Code-tuned · 256K context — short coding tasks only
~55
5
DeepSeek V4 Pro
Your current model · generalist · upgrade to GLM-5.2 for terminal
~50
6
Gemma 4 31B
Lightweight local · not for long-horizon coding
~35
🔥 GLM-5.2 scores 81% on Terminal-Bench — higher than Claude Fable 5, GPT-5.6, and every other model.
The best coding agent you can run locally. Period.
🏆 Cloud picks for coding
  • GPT-5.6 Terra high (→ xhigh long jobs) — default implementer · dual OK · best $/PR
  • Claude Opus 4.8 high (→ xhigh) — dual Plan+Code · same model that planned can implement (prefer new session)
  • GPT-5.6 Sol high / xhighdual · CA 80 / Terminal 88.8% · use xhigh for long agent ceiling
  • Claude Fable 5 high → xhigh — dual · high normal PR · xhigh migrations (SWE-Pro 80.3%)
  • Grok 4.5 highdual · your Grok plan · CA~76 · TB 83.3% · cheap volume agents · fact-check

Same model, two jobs is fine (Fable/Opus/Sol/Terra/Grok). Cost path: plan Sol/Fable/Opus high → implement Terra high or Grok high.

🦙 Best Ollama: GLM-5.2
  • ▸ Terminal-Bench 2.1: 81.0% — beats every cloud model
  • ▸ Solid 1M context, effort level control, MIT license
  • ▸ FrontierSWE: #2 only behind Claude Opus 4.8
  • ▸ ⚡ Runner-up: MiniMax M3 (59% SWE-Bench Pro + native vision)

🔎 Review: Plan review vs Code review — which model?

Short answer: Review is closer to critique / adversarial reasoning than pure generation. Prefer a planning-class model for plan review, and a strong coding model for code review — ideally a different model (or fresh session) than the one that wrote the plan/code. Implementation models can review code, but they often miss strategy holes; planning models can skim code, but they miss bugs.

📋 Review the Plan

Goals: find missing requirements, bad trade-offs, risk, scope gaps, inconsistent steps, “sounds good but fails in production”.

Best models
  • Claude Fable 5 high — deepest critique, long docs (xhigh only if huge multi-doc research)
  • Claude Opus 4.8 high — classic adversarial reviewer · great second opinion after Fable/Sol plans
  • GPT-5.6 Sol high (→ xhigh if complex) — knowledge work + design judgment
  • Terra high — budget plan critique after Sol wrote it
  • Ollama: Qwen3.5 — best free plan reviewer
Skill family: Same as Planning/Brainstorming (reasoning, strategy, risk) — not the pure implementer. Use planning-class models; ask them to be adversarial (“find 5 failure modes”).
🧪 Review the Coding Implementation

Goals: bugs, edge cases, security, regressions, tests, API contracts, “matches the plan?”, maintainability.

Best models
  • Claude Fable 5 high (→ xhigh huge repo) — best thorough PR/repo review
  • Claude Opus 4.8 high (→ xhigh) — FrontierSWE-class review · alternate Claude when Fable wrote the code
  • GPT-5.6 Sol high (→ xhigh) — agentic review, tools, security-ish checks
  • Terra high / xhigh — routine PR review at ½ Sol cost
  • Ollama: GLM-5.2 — best free code/terminal review
Skill family: Same as Coding Implementation (code understanding + tools) — but run as reviewer, not author. Implementer models can review; better if different model/session than the one that wrote the code.
Task Closest skill family Best pick Can implementer / planner do it?
Review the plan Planning & Brainstorming Sol high / Fable high / Opus 4.8 high · Ollama Qwen3.5 Yes — use planning models. Avoid pure coding specialists (GLM) as primary plan reviewers. Prefer a model that did not write the plan (Opus is a strong second Claude).
Review the code Coding Implementation (+ critique) Fable/Opus/Sol high · Terra high routine · Ollama GLM-5.2 Yes — coding models can review. Terra that implemented can self-check, but a second model (Sol/Fable/Opus/GLM) catches more. Planning-only models are weak on deep code review.
Plan ↔ Code consistency Both (cross-check) Sol high · Fable high · Opus 4.8 high Needs strong reasoning + enough code skill. Best: paste plan + diff into Sol/Fable/Opus high.
Recommended 4-step pipeline:
(1) Plan — Sol high or Fable high or Opus 4.8 high (not xhigh by default) · (2) Review plan — different session: Sol/Fable/Opus high adversarial · Ollama Qwen3.5 · (3) Implement — Terra high (default) · Fable/Opus/Sol high for normal code · Fable/Opus/Sol/Terra xhigh only for multi-hour / multi-file / migrations · (4) Review code — Fable/Opus/Sol high, or GLM-5.2 local — not the same implementer session.

🧭 When to use Sol/Terra High vs Extra High

Research synthesis — OpenAI docs + AA cost curves + community defaults (July 2026). Extra High = API xhigh.

🧠 Brainstorm / Planning
  • Default: GPT-5.6 Sol high — best sweet spot for hard / unclear / multi-system design
  • Escalate: Sol Extra High (xhigh) — novel architecture, deep research, many trade-offs, async long runs
  • Budget: Terra high → Terra xhigh — good product planning when Sol is too expensive
  • Why Sol for brainstorm? Sol leads Agents' Last Exam & design judgment; investigates ambiguous goals better
🛠️ Implementation / Coding
  • Default: GPT-5.6 Terra medium / high — plan is clear → execute PRs, features, refactors at ½ Sol price
  • Harder code: Terra Extra High (xhigh) — long-horizon engineering, multi-file, CA ≈ Fable (77.4)
  • Ceiling: Sol high → Sol xhigh — terminal agents, cyber, UI polish, tool-heavy agents (CA 80)
  • Why Terra for implement? Within ~2–3 pts of Sol on many agentic coding benches at half cost once scope is defined
Config Use when Avoid when
Sol high Default hard reasoning: architecture, tough debug, multi-system plans. Community “Pareto daily” for intelligence/price. Routine chat / trivial edits (use medium or Terra).
Sol Extra High (xhigh) Measured quality gain needed: deep research, long agent loops, terminal SOTA, cyber, multi-angle verification. OpenAI: use high/xhigh when evals show gain. As blanket default — often only slight intel gain vs high, much higher tokens/cost. Prefer max only if xhigh still fails.
Terra high Execution after plan is clear; everyday implement when medium misses edge cases. Strong balance of speed/cost. Open-ended novel design with high ambiguity — step up to Sol high first.
Terra Extra High (xhigh) Budget long-horizon coding before paying Sol; CA ~77.4 ≈ Fable coding agent; multi-file migrations at Terra rates. You need absolute Sol ceiling (terminal 88.8%, ALE 53.6, cyber) — jump to Sol xhigh instead of overpaying Terra max for less ceiling.
Recommended dual-mode workflow: (1) Brainstorm/plan with Sol high, Fable high, or Opus 4.8 high (not xhigh by default; → xhigh only if stuck on deep research) → write acceptance criteria → (2) Implement with Terra high (default) or Fable/Opus/Sol high for normal PRs → raise xhigh only for multi-file / multi-hour / migrations; Sol xhigh for terminal/cyber/UI ceiling. Start API effort at medium; raise when quality slips. Full scenarios → Effort Guide tab.

🎯 Your Optimal Setup

☁️ GPT
GPT-5.6 Stack

Plan: Sol high
Review plan: Sol high* / Fable / Opus
Implement: Terra high
Review code: Sol high / Fable / Opus
*different session than author

☁️ Claude
Claude Stack

Plan: Fable high · Opus high
Review plan: Opus* / Fable* (swap)
Implement: Fable high · Opus high
Review code: other Claude / Sol
*Fable wrote → review with Opus (or Sol)

☁️ + 🦙
Hybrid (Best)

Plan + review plan: Sol / Fable / Opus
Implement: Terra · Ollama GLM-5.2
Review code: Fable / Opus / Sol / GLM
Best of both worlds

🦙 + 🦙
Ollama-Only

Planning: Qwen3.5
Coding: GLM-5.2
100% free & local

🔄
Current → Upgrade

From: DeepSeek V4 Pro
Cloud: Terra med → Sol xhigh
Local: GLM-5.2 + Qwen3.5
Clear upgrade on both fronts

❌ Why NOT the other Ollama models for these tasks?

Kimi K2.7 Code

Only 256K context — can't handle long-horizon coding. Outclassed by GLM-5.2 and M3 on every benchmark. Weakest of the four.

DeepSeek V4 Pro

Jack-of-all-trades, master of none. ~50% SWE-Bench Pro vs GLM-5.2's 81% Terminal-Bench. Good generalist but not specialized enough for either task.

Gemma 4 31B

Only 128K context and 31B params. Too small for serious coding or deep planning. Good for quick prototyping only.

💼 Role Picks — best model by job

Best starting pick for each role, based on benchmarks + practical dual-use. Many cloud flagships work for both plan and implement — the card shows the usual default, not the only option.

Pipeline reminder: Plan Sol/Fable/Opus high → Implement Terra high (or same dual model in a new session) → Review with a different model/session.
👨‍💻
Senior Software Engineer
Architecture, hard reviews, complex refactors
C5
Claude Fable 5 high → xhigh
SWE-Pro 80.3% · dual plan+code
O4
Claude Opus 4.8 high
Alt dual Claude · great second reviewer
🤖
AI Agent Developer
Autonomous agents, tool calling, long-horizon
GL
GLM-5.2 (Ollama)
Terminal-Bench 81% — best local agent coding
S
GPT-5.6 Sol high/xhigh
Cloud agent ceiling · Terminal ~88.8%
🎨
Full-stack Developer
Frontend + backend, UI/UX, product polish
S
GPT-5.6 Sol medium/high
Design/UI strength · dual plan+code
T
GPT-5.6 Terra high
Default implement after plan is clear
🏗️
Large codebase refactor
Multi-module migrations, system-wide restructure
X5
Claude Fable 5 xhigh
Long-horizon coding · SWE-Pro class work
🏭
SAP / MES / ERP
Enterprise systems, large DBs, complex business rules
C5
Claude Fable 5 high
Deep reasoning + 1M context · dual
T
GPT-5.6 Terra high
Clear tickets / feature slices after plan
🛒
WordPress / WooCommerce
Plugins, themes, store automation
S
GPT-5.6 Sol high
UI + web automation · dual
📚
Million-token RAG
Huge corpora, long-doc retrieval
K3
Kimi K3
1M context · strong value open-weight cloud
🦙
Local / Open Deployment
On-prem, offline, privacy-first
Q
Qwen3.5 — plan / docs / vision
G
GLM-5.2 — local coding agent
Ge
Gemma 4 31B — light / fast
💰
Strong on a budget
Best cost / performance defaults
T
GPT-5.6 Terra medium/high
Cloud default implementer · ~½ Sol
D4
DeepSeek V4 Pro
Free local generalist
📝
Content & Marketing
Copy, SEO, social, ads, multimodal assets
M3
MiniMax M3
Native multimodal — text + image + video

📋 Quick reference

You do… Best pick Why Type
Senior Software Engineer Fable 5 high · Opus 4.8 high SWE-Pro / dual Claude · review with the other ☁️ Cloud
AI Agent Developer GLM-5.2 · Sol high/xhigh Local TB 81% / cloud terminal ceiling 🦙 Ollama / ☁️ Cloud
Full-stack Developer Sol high · Terra high UI/design + default implement ☁️ Cloud
Large refactor Fable 5 xhigh Long-horizon multi-file migrations ☁️ Cloud
SAP / MES / ERP Fable high · Terra high Plan dual Claude · implement cheap clear PRs ☁️ Cloud
WordPress / WooCommerce Sol high UI + web automation ☁️ Cloud
Million-token RAG Kimi K3 1M context, strong value ☁️ Cloud
Local / Open Qwen3.5 · GLM-5.2 Plan vs code split (not dual) 🦙 Ollama
Budget strong Terra med/high · DeepSeek Cloud cheap implement + free local ☁️ Cloud / 🦙 Ollama
Content & Marketing MiniMax M3 Native multimodal assets 🦙 Ollama

🏆 Recommendations — the short list

If you only remember one pipeline: plan with Sol/Fable/Opus high → implement with Terra high (or the same dual model in a new session) → review with a different model. Raise to xhigh only for multi-file / multi-hour / migrations.

🏆 Best Overall (Cloud)

  • Claude Fable 5 high → xhigh — Best SWE-Pro repo coding & multi-day agents; dual plan+code; start high
  • Claude Opus 4.8 high → xhigh — Dual Claude flagship; plan, implement, adversarial review
  • GPT-5.6 Sol medium → xhigh — Best agentic/terminal/cyber/ALE; dual; xhigh for SOTA ceiling
  • GPT-5.6 Terra medium → xhigh — Best daily implementer / budget frontier (~½ Sol) · ChatGPT Plus
  • Grok 4.5 high — Your Grok plan · CA~76 · TB 83.3% · dual · best token/$ agent · fact-check

💰 Best Value (Cloud)

  • Grok 4.5 high — ~$2/$6 · AA ~$0.31/Intel task · CA~76 · TB 83.3% · SWE Marathon lead
  • GPT-5.6 Terra Medium/High — $2.50/$15, within ~2–3 pts of Sol; default PR implementer
  • Kimi K3 — $3/$15 per M, 2.8T params, 1M context, open-weight cloud

🦙 Best Open/Local (Ollama)

  • Qwen3.5 — Best free planner: vision-language, multilingual, docs
  • GLM-5.2 — Best free coder: terminal agent, long-horizon, MIT
  • MiniMax M3 — Computer use + native multimodal agents

⚡ Best for Specific Tasks

  • Plan / architecture → Fable/Opus/Sol high
  • Normal PR implement → Terra high · Grok 4.5 high · dual Fable/Opus/Sol high
  • Code Migration → Fable 5 xhigh · Sol/Opus xhigh
  • Design/UI/Frontend → Sol medium/high
  • Terminal/CLI (cloud) → Sol xhigh · Terra xhigh
  • Terminal/CLI (local) → GLM-5.2
  • Cybersecurity → Sol high/xhigh
  • Code review → Fable/Opus/Sol high ≠ author
  • Multilingual → Qwen3.5 (201 languages)
  • Computer Use → Sol xhigh or MiniMax M3 (local)
  • Lightweight Local → Gemma 4 31B

🔮 Recommended Pairings

  • Your plans stack: Claude Max 5x (Fable+Opus) · ChatGPT Plus (5.6) · Grok 4.5 high · Ollama Cloud
  • Default dual stack: Fable/Opus/Sol high (plan) + Terra/Grok high (build) + other session review
  • Same-model dual: Opus, Sol, or Grok high for plan and implement (new session for build)
  • Effort ladder: med → high → xhigh only if stuck · never start at max
  • Cloud + Local: Sol/Terra cloud + GLM-5.2 (terminal) + Qwen3.5 (plan/docs)
  • Budget Stack: Terra medium/high + DeepSeek V4 Pro
  • Multimodal Stack: Sol medium + Qwen3.5 (vision) + MiniMax M3 (computer use)
  • Orchestrator: Fable/Opus high plans · Terra workers execute

⚠️ Caveats

  • Claude Fable 5: Safeguard fallback to Opus 4.8 on ~5–8% of requests (cyber/bio); xhigh burns 2–4× tokens
  • Claude Opus 4.8: Excellent dual model — still not “always xhigh”; use high by default
  • Grok 4.5 high: Near-frontier agentic value; trails Fable on SWE-Pro (64.7% vs 80.3%); verify factual QA
  • GPT-5.6 Sol: Leads agentic benches, trails Fable on SWE-Bench Pro (64.6% vs 80.3%)
  • GPT-5.6 Terra: Best practical implementer; not the best open-ended strategist alone
  • Grok 4.5: Higher hallucination on factual QA — verify outputs
  • Ollama: Needs GPU VRAM; Qwen 397B class needs ~80GB+ — not dual (Qwen plans, GLM codes)
  • Gemma 4 31B: 128K context only — not for long-horizon tasks
Session tip