July 2026 field guide for your stack: Claude Max 5x (Fable 5 · Opus 4.8) · ChatGPT Plus (GPT-5.6) · Grok (Grok 4.5 High) · Ollama Cloud. Pick model then effort. Hover chips top-right to highlight. Default: plan Fable/Opus/Sol high → build Terra/Grok high → review another session.
How to read this tab: Charts compare published peaks — not always daily effort. Fable leads SWE-Pro (80.3%). Opus 4.8 = dual Claude Max 5x (SWE-Pro 69.2 · TB 74.6 Terminus-2 · AA 61.4 launch · $5/$25). Sol leads many agentic/terminal scores. Terra = cost-aware implementer. Grok 4.5 high = your Grok plan (CA~76, TB 83.3%, cheap dual). Ollama Cloud = Qwen plan · GLM code.
Real multi-file GitHub tasks. Fable leads; Sol/Terra trail on this bench even when they win agentic scores.
Artificial Analysis composite (reasoning, agents, science). Good proxy for planning / hard thinking.
Ordinal guide (not a lab metric). Higher = more capability per dollar. Free Ollama scores high on value; flagships trade money for ceiling. Opus $5/$25 ≈ better $/M than Fable $10/$50 at near-frontier peaks.
Peer-scale 1M quality (not AA-LCR %). Critical for big repos & RAG. Opus AA-LCR is roughly flat vs 4.7 (~68% on public AA-LCR boards — different scale).
| Model | Provider | Type | Params | Context | SWE-Bench Pro | Terminal-Bench 2.1 | Intel Index | Input $/M | Output $/M |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 high | Anthropic | ☁️ Cloud | — | 1M | 80.3 | ~84–86 | 60–65 | $10 | $50 |
| Claude Fable 5 xhigh | Anthropic | ☁️ Cloud | — | 1M | 80.3+ | ~86 | ~60 max | $10 | $50 |
| Claude Opus 4.8 high → max | Anthropic | ☁️ Cloud | — | 1M | 69.2 | 74.6† | 61.4‡ | $5 | $25 |
| GPT-5.6 Sol medium | OpenAI | ☁️ Cloud | — | 1M | ~65 | ~86 | ~56 | $5 | $30 |
| GPT-5.6 Sol high | OpenAI | ☁️ Cloud | — | 1M | ~66 | ~88 | ~58 | $5 | $30 |
| GPT-5.6 Sol xhigh / max | OpenAI | ☁️ Cloud | — | 1M | 64.6 Pro | 88.8 | 59 max | $5 | $30 |
| GPT-5.6 Terra medium | OpenAI | ☁️ Cloud | — | 1M | ~62 | ~85 | ~52 | $2.50 | $15 |
| GPT-5.6 Terra high | OpenAI | ☁️ Cloud | — | 1M | ~63 | ~86 | ~54 | $2.50 | $15 |
| GPT-5.6 Terra xhigh / max | OpenAI | ☁️ Cloud | — | 1M | ~64 | 87.4 | 55 max | $2.50 | $15 |
| Grok 4.5 high | xAI / SpaceXAI | ☁️ Cloud | MoE | 500K–1M* | 64.7 | 83.3 | 54 (high) | $2 | $6 |
| Kimi K3 | Moonshot AI | ☁️ Cloud | 2.8T (MoE) | 1M | ~70 | — | — | $3 | $15 |
| Qwen3.5 | Alibaba | 🦙 Ollama | 397B (17B act) | 1M | 76.4* | 52.5 | — | Free | Free |
| GLM-5.2 | Zhipu AI | 🦙 Ollama | — (MoE) | 1M | 62.1 | 81.0 | — | Free | Free |
| Kimi K2.7 Code | Moonshot AI | 🦙 Ollama | 1T (32B act) | 256K | ~55 | — | — | Free | Free |
| MiniMax M3 | MiniMax | 🦙 Ollama | — (MoE) | 1M | 59.0 | 66.0 | — | Free | Free |
| Gemma 4 31B | 🦙 Ollama | 31B | 128K | ~35 | — | — | Free | Free | |
| DeepSeek V4 Pro | DeepSeek | 🦙 Ollama | 1.6T (49B act) | 1M | ~50 | — | — | Free | Free |
* Qwen3.5 uses SWE-bench Verified where SWE-Pro unpublished. Claude Opus 4.8 (sources, May 2026): SWE-Bench Pro 69.2% · SWE-Verified 88.6% · Terminal-Bench 2.1 74.6% on Terminus-2 harness († harness-sensitive — GPT-5.5 Codex CLI can read higher on the same suite) · AA Intelligence Index 61.4 at launch max effort (‡ AA May 28 analysis; later AA index revisions may re-rank) · GDPval-AA Elo 1890 · OSWorld-Verified 83.4% · HLE no-tools 49.8% · list $5/$25 per M (fast mode $10/$50). No separate published SWE-Pro for xhigh — raise effort only when quality fails. Grok 4.5 (high): AA Intel 54 · Coding Agent ~76 · Terminal-Bench 2.1 83.3% · SWE-Pro 64.7% · ~$2/$6. Sol Coding Agent max ≈ 80; Fable SWE-Pro 80.3%. Context windows vary by product surface.
Same model family, different intelligence / latency / cost trade-offs. Effort does not change the $/M token rate — it changes how many tokens the model spends (thinking, tool calls, verification). Rule of thumb: start medium/high; only raise to xhigh/max when quality fails after 2–3 attempts.
Anthropic Fable 5 (Claude Max 5x): API default is high. Use xhigh only for long-horizon work (30+ min agents, multi-million token budgets). Dual-capable: plan and implement.
Claude Opus 4.8 (Claude Max 5x): Prior Claude flagship — plan, implement, adversarial review. Published peaks: SWE-Pro 69.2%, TB 2.1 74.6% (Terminus-2), AA Intel 61.4 (launch), $5/$25. Prefer high daily; raise for hard multi-file. Best second Claude when Fable authored the work.
OpenAI GPT-5.6 (ChatGPT Plus / API): Tiers Sol ($5/$30) · Terra ($2.50/$15) · Luna ($1/$6). Effort: none → low → medium (default) → high → xhigh → max. Ultra = multi-agent, not an effort level. Terra high = default implementer after a clear plan.
Grok 4.5 High (Grok / xAI): Your practical Grok setting. AA Intelligence ≈ 54 (high) · Coding Agent ≈ 76 in Grok Build · Terminal-Bench 2.1 83.3% · SWE-Pro 64.7% · SWE Marathon 29% (leads Fable/Opus on that long-horizon harness). List ~$2/$6 per M and very token-efficient (~⅓–½ the tokens of peers on AA tasks). Dual-capable for fast plan+code; verify factual claims (higher hallucination risk than Claude/GPT on pure QA).
Pareto note: Grok 4.5 often sits on the cost–capability frontier (~$0.31/AA task, ~$2.5/Coding Agent task in public AA writeups). Sol (max) Intel ≈ 59 vs Fable ≈ 60; Terra ≈ 55. Use Grok high for cheap near-frontier agents; escalate to Fable/Sol when SWE-Pro-class repo quality is mandatory.
effortmax_tokens high (start ~64k) — hard ceiling on think+textultracode = xhigh + multi-agent permission (not a separate API level)reasoning.effort: "xhigh"xhigh vs max on hard evals — max explores even longer| Config | API effort | $/M in | $/M out | AA Intel | Coding Agent | Terminal 2.1 | SWE-Pro | ALE | When to use |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 xhigh | xhigh | $10 | $50 | ~60 | 77.2 | ~84–86 | 80.3 | 40.5 | Repo migrations, multi-day agents, deep research |
| GPT-5.6 Sol medium | medium | $5 | $30 | ~56 | ~76 | ~86 | ~63 | ~51+ | Default Sol interactive work |
| GPT-5.6 Sol high | high | $5 | $30 | ~58 | ~78 | ~88 | ~64 | ~52 | Hard debug, high-stakes plans |
| GPT-5.6 Sol Extra High | xhigh→max | $5 | $30 | 59 | 80 | 88.8 | 64.6 | 53.6 | Cyber, terminal SOTA, multi-agent Ultra |
| GPT-5.6 Terra medium | medium | $2.50 | $15 | ~52 | ~74 | ~85 | ~61 | ~49 | Daily driver after GPT-5.5 |
| GPT-5.6 Terra high | high | $2.50 | $15 | ~54 | ~76 | ~86 | ~63 | ~50 | Escalate without Sol bill |
| GPT-5.6 Terra Extra High | xhigh→max | $2.50 | $15 | 55 | 77.4 | 87.4 | ~64 | 50.4 | Budget long-horizon frontier |
Sources: Anthropic effort docs + Fable 5 launch; OpenAI GPT-5.6 GA (Jul 9, 2026); Artificial Analysis Intelligence/Coding Agent Index & cost-per-task (Jul 2026); Vellum / independent writeups for Terminal-Bench & ALE. Medium/high midpoints are interpolated where labs only publish max-effort curves — treat as decision guide, re-eval on your harness.
Default stack: GPT-5.6 Terra Medium → GPT-5.6 Terra High if quality slips → GPT-5.6 Sol High/xhigh for agentic/cyber/UI ceiling → Claude Fable 5 High/xhigh for SWE-Pro-class repo migrations & multi-day coding.
Cost trap: Fable xhigh and Sol max can burn 2–7× tokens vs medium/high on the same prompt. Measure thinking_tokens / reasoning tokens; raise max_tokens before blaming quality.
Split brain: Fable orchestrator + cheaper workers (Sonnet/Haiku or Terra) can hit ~96% of all-Fable BrowseComp quality at ~46% cost (Anthropic multi-agent guidance).
Effort does not change the $/million token price. It changes how much the model thinks, how many tools it uses, and how long it runs. Higher effort = more tokens, latency, and cost. Start low enough to finish fast; raise only when quality fails.
❌ Why not Fable xhigh / Opus 4.8 xhigh / Sol Extra High for everything? → jump to answer · 🧭 How to choose
Stop guessing. Answer 3 questions (or use the cheat sheet). Two knobs matter: which model (Fable / Sol / Terra / Ollama) and how hard it thinks (medium · high · xhigh). Effort does not change $/M token price — it burns more tokens.
Fable · Sol · Terra · Ollama
medium · high (default work) · xhigh · max
Dual workflow: Sol/Fable plans → Terra builds → Sol/Fable reviews
“Best model + max effort on everything” feels safe. It is usually slower, more expensive, and sometimes worse for small jobs. Max capability is a tool for hard, long jobs — not a default lifestyle. Direct link: #why-not-always-max
| Situation | Use | Why |
|---|---|---|
| Unclear / design / “how should we…?” | Sol high or Fable high | Interactive strategy; high is enough |
| Plan / architecture | Sol high · Fable high | Not xhigh by default |
| Review plan or PR | Fable/Sol high · other session | Author ≠ reviewer |
| Clear implement (normal PR) | Terra high | Default builder; ~½ Sol price |
| XS typo / rename / 1-line | Terra medium | Don’t burn high/xhigh |
| Long multi-file / migration | Fable/Sol xhigh · budget Terra xhigh | Long-horizon agent coding |
| Docs / release notes | Terra medium/high | Known facts → cheap work |
| Local / privacy / offline | GLM-5.2 code · Qwen3.5 plan | Free; escalate cloud if hard |
| Stuck after 2–3 fails | Raise one step only | Terra→Sol high→Fable/Sol xhigh |
Plan = high is not “plan is less important.” xhigh is for long coding agents, not every design chat.
Terra high → Sol high → Fable/Sol xhigh↑ Back to: Why not always Fable/Opus xhigh / Sol Extra High?
Not because plan is less important. Anthropic’s xhigh is built for long-horizon agentic coding (often 30+ min, many tools, self-verify). Planning is usually short interactive turns — high is the API default and already strong. Using xhigh on every plan wastes 2–4× tokens. Using only high on a day-long migration under-explores. Rule: Fable high for plan & normal code; Fable xhigh when the coding job is multi-hour / multi-file / migration-class.
Follow the flow left → right. Author ≠ reviewer when quality matters. Effort: start medium/high; raise to xhigh only on long hard jobs.
| # | Phase | What you do | Cloud model + effort | Ollama / free |
|---|---|---|---|---|
| 1 | Discover & Brainstorm | Ideas, problem framing, options, spikes | Sol high · Fable high | Qwen3.5 |
| 2 | Plan & Architecture | ADRs, APIs, data model, milestones, risks | Sol high · Fable high · Terra high budget | Qwen3.5 |
| 3 | Review the Plan | Adversarial check: gaps, failure modes, scope | Fable high* · Sol high* (*≠ author) | Qwen3.5 |
| 4 | Implement | Write code, PR features from approved plan | Terra high default · Fable/Sol high · xhigh if multi-file / multi-hour | GLM-5.2 · MiniMax M3 |
| 5 | Review Code | Bugs, security, tests, plan match | Fable high* · Sol high* · Terra high routine | GLM-5.2 |
| 6 | Test & Debug | Failing tests, flaky CI, root-cause | Terra high · Sol high hard · Fable high/xhigh deep | GLM-5.2 terminal |
| 7 | Docs & Release | Changelog, runbooks, release notes, README | Terra medium/high · Sol medium | Qwen3.5 |
| 8 | Maintain & Hotfix | Small fixes daily; big incidents escalate | Terra medium/high · Sol high incidents · Fable xhigh major outage | GLM-5.2 · DeepSeek |
| Real situation | Best pick | Why |
|---|---|---|
| “Help me design auth for a SaaS” | Sol high · Fable high | Interactive strategy; high is enough |
| “Review this PR (200 lines)” | Terra high · Sol high · GLM-5.2 | Code critique, not multi-hour agent |
| “Implement the checkout ticket” | Terra high (default) | Clear scope → execute cheaply |
| “Migrate Express → Nest across 40 modules” | Fable xhigh · Sol xhigh | Long-horizon coding = xhigh territory |
| “Fix typo / rename variable” | Terra medium · Sol medium | Don’t burn high/xhigh on trivial work |
| “Red-team this architecture for failures” | Sol high / xhigh · Fable high | Adversarial planning skill |
Price always $10/$50 per M. Effort only changes token volume. Anthropic default = high.
Flagship OpenAI tier · $5/$30 per M · API default effort = medium. UI “Extra High” = xhigh. Also has max above xhigh for hardest jobs.
Balanced tier · $2.50/$15 per M (~½ Sol). Best everyday implementer once the plan is clear. Still supports medium → max; use Sol xhigh when you need the absolute ceiling.
| Time | You do | Model + effort |
|---|---|---|
| 09:00 | Brainstorm checkout redesign with PM notes | GPT-5.6 Sol high · or Fable high |
| 10:00 | Adversarial review of the plan (“find 5 failure modes”) | Fable high / Sol high · different session · Ollama Qwen3.5 |
| 11:00–16:00 | Implement tickets from the approved plan | Terra high (default) · Terra xhigh if multi-file slog |
| 16:30 | PR review before merge | Sol high or Fable high · or GLM-5.2 local |
| Night agent | Long migration / flaky CI green-up | Fable xhigh · Sol xhigh · (budget: Terra xhigh/max) |
After the big SDLC loop (brainstorm → plan → implement → review → test → docs → maintain), most real work is incoming changes: bugs, small tweaks, medium features, big features. Size the change first, then pick model + effort. Do not default to Claude Fable 5 xhigh or GPT-5.6 Sol high for everything.
| User says… | Size | Model + effort | Skip |
|---|---|---|---|
| “Fix typo on checkout button” | XS | GPT-5.6 Terra medium | Claude Fable 5 xhigh · GPT-5.6 Sol high |
| “Date format wrong in VN locale” | S | GPT-5.6 Terra high · Ollama GLM-5.2 | Claude Fable 5 xhigh |
| “Add export invoices to CSV” | M | Plan Sol/Fable high → Implement Terra high | All-in Fable xhigh |
| “Rebuild permissions for multi-tenant orgs” | L | Sol/Fable high plan → Terra/Fable high–xhigh build | Terra medium alone |
| “Migrate whole API to NestJS” | XL | Claude Fable 5 xhigh · GPT-5.6 Sol xhigh | GPT-5.6 Terra medium |
| “Site is slow / something’s wrong” | Unclear | GPT-5.6 Sol high investigate first | Blind implement on any model |
| “Make UI prettier like Linear” | M–L | GPT-5.6 Sol high (design) → Terra high implement | Claude Fable 5 xhigh first pass |
| “Prod down after deploy” | XL incident | Sol/Fable high now → xhigh if multi-service | Waiting for perfect cheap model |
Start at the lowest effort that still succeeds. Escalate one step at a time (medium → high → xhigh → max). Prefer a second model for review. Plan with high; long implement with xhigh — not the other way around. For tweaks: size the change first — most issues are S/M, not Fable xhigh.
Each card is a configuration you can actually pick (model + effort where it matters). Order: Fable → Sol → Terra → other cloud → Ollama. Hover a top-right chip to highlight a family. Dual-capable cloud flagships (Fable, Opus, Sol, Terra) appear for both plan and code — specialists live mainly under Ollama.
Qualitative fit by task type (not a single benchmark number). Green does not mean “always use xhigh” — use the Effort Guide for dialing. Hover Fable / GPT-5.6 / Ollama chips to isolate columns. Opus is dual-capable like Fable but is grouped under Claude in narrative tabs; matrix columns list primary published configs.
= Excellent = Good = Adequate = Limited
You're currently on DeepSeek V4 Pro. Four task modes: brainstorm/plan, implement code, review the plan, review the code — different roles and defaults, not always different models.
Yes — the same cloud model can plan and implement. Both columns list Fable, Opus, Sol, Terra, and Grok 4.5 for a reason: they are generalists. What changes is metric, rank, and default role.
| Model (your access) | Can plan? | Can implement? | Best default role |
|---|---|---|---|
| Claude Fable 5 high/xhigh · Max 5x | ✓ Excellent | ✓ Excellent (SWE-Pro) | high plan + normal PR · xhigh migrations |
| Claude Opus 4.8 high · Max 5x | ✓ Excellent | ✓ Excellent | Plan + review + hard code · second Claude after Fable |
| GPT-5.6 Sol high/xhigh · ChatGPT Plus | ✓ Excellent | ✓ Excellent | high plan/debug · xhigh long agent / terminal |
| GPT-5.6 Terra high/xhigh · ChatGPT Plus | ✓ Good | ✓ Excellent ($/PR) | high default implement · xhigh long multi-file |
| Grok 4.5 high · Grok plan | ✓ Strong / fast | ✓ Strong agent (CA~76 · TB 83.3%) | Cheap dual plan+code · overnight agents · fact-check knowledge |
Where columns differ: (1) metric — Intelligence vs Coding Agent / SWE-Pro · (2) defaults — Plan leans Fable/Opus/Sol high; Implement leans Terra/Grok high · (3) Ollama Cloud split — Qwen plans, GLM codes (not dual).
Recommended daily loop on your plans: (1) Plan on Fable high or Sol high · (2) Optional adversarial review Opus high or Sol · (3) Implement Terra high or Grok 4.5 high (volume/cheap) · hard multi-file → Fable/Sol xhigh · (4) Review ≠ author · (5) Ollama Cloud for offline drafts / terminal GLM.
Architecture design, creative thinking, research, strategy
Default effort: Fable/Opus/Sol high (not xhigh). Same model can later implement — prefer a fresh session for build.
| Skill | GLM-5.2 | Qwen3.5 |
| General knowledge | 🟡 | 🟢 |
| Multilingual (201 langs) | 🔴 | 🟢 |
| Search / Research agent | 🟡 | 🟢 |
| Creative brainstorming | 🟡 | 🟢 |
| Vision understanding | 🔴 | 🟢 |
| Terminal coding | 🟢🟢🟢 | 🟡 |
GLM-5.2 was RL-trained on coding environments. Qwen3.5 was RL-trained across all domains — reasoning, search, vision, agents, math. Breadth wins for brainstorming.
↪ After plan is approved: keep dual model (Fable/Opus/Sol/Grok) in a new session, or switch to Terra high / Grok high to save cost.
Writing code, debugging, refactoring, terminal agents
Default effort: Terra high for clear PRs. Opus / Sol / Fable high also implement well (dual). Raise xhigh only for multi-file / multi-hour jobs.
↪ Same model, two jobs is fine (Fable/Opus/Sol/Terra/Grok). Cost path: plan Sol/Fable/Opus high → implement Terra high or Grok high.
Short answer: Review is closer to critique / adversarial reasoning than pure generation. Prefer a planning-class model for plan review, and a strong coding model for code review — ideally a different model (or fresh session) than the one that wrote the plan/code. Implementation models can review code, but they often miss strategy holes; planning models can skim code, but they miss bugs.
Goals: find missing requirements, bad trade-offs, risk, scope gaps, inconsistent steps, “sounds good but fails in production”.
Goals: bugs, edge cases, security, regressions, tests, API contracts, “matches the plan?”, maintainability.
| Task | Closest skill family | Best pick | Can implementer / planner do it? |
|---|---|---|---|
| Review the plan | Planning & Brainstorming | Sol high / Fable high / Opus 4.8 high · Ollama Qwen3.5 | Yes — use planning models. Avoid pure coding specialists (GLM) as primary plan reviewers. Prefer a model that did not write the plan (Opus is a strong second Claude). |
| Review the code | Coding Implementation (+ critique) | Fable/Opus/Sol high · Terra high routine · Ollama GLM-5.2 | Yes — coding models can review. Terra that implemented can self-check, but a second model (Sol/Fable/Opus/GLM) catches more. Planning-only models are weak on deep code review. |
| Plan ↔ Code consistency | Both (cross-check) | Sol high · Fable high · Opus 4.8 high | Needs strong reasoning + enough code skill. Best: paste plan + diff into Sol/Fable/Opus high. |
Research synthesis — OpenAI docs + AA cost curves + community defaults (July 2026). Extra High = API xhigh.
| Config | Use when | Avoid when |
|---|---|---|
| Sol high | Default hard reasoning: architecture, tough debug, multi-system plans. Community “Pareto daily” for intelligence/price. | Routine chat / trivial edits (use medium or Terra). |
| Sol Extra High (xhigh) | Measured quality gain needed: deep research, long agent loops, terminal SOTA, cyber, multi-angle verification. OpenAI: use high/xhigh when evals show gain. | As blanket default — often only slight intel gain vs high, much higher tokens/cost. Prefer max only if xhigh still fails. |
| Terra high | Execution after plan is clear; everyday implement when medium misses edge cases. Strong balance of speed/cost. | Open-ended novel design with high ambiguity — step up to Sol high first. |
| Terra Extra High (xhigh) | Budget long-horizon coding before paying Sol; CA ~77.4 ≈ Fable coding agent; multi-file migrations at Terra rates. | You need absolute Sol ceiling (terminal 88.8%, ALE 53.6, cyber) — jump to Sol xhigh instead of overpaying Terra max for less ceiling. |
medium; raise when quality slips. Full scenarios → Effort Guide tab.
Plan: Sol high
Review plan: Sol high* / Fable / Opus
Implement: Terra high
Review code: Sol high / Fable / Opus
*different session than author
Plan: Fable high · Opus high
Review plan: Opus* / Fable* (swap)
Implement: Fable high · Opus high
Review code: other Claude / Sol
*Fable wrote → review with Opus (or Sol)
Plan + review plan: Sol / Fable / Opus
Implement: Terra · Ollama GLM-5.2
Review code: Fable / Opus / Sol / GLM
Best of both worlds
Planning: Qwen3.5
Coding: GLM-5.2
100% free & local
From: DeepSeek V4 Pro
Cloud: Terra med → Sol xhigh
Local: GLM-5.2 + Qwen3.5
Clear upgrade on both fronts
Only 256K context — can't handle long-horizon coding. Outclassed by GLM-5.2 and M3 on every benchmark. Weakest of the four.
Jack-of-all-trades, master of none. ~50% SWE-Bench Pro vs GLM-5.2's 81% Terminal-Bench. Good generalist but not specialized enough for either task.
Only 128K context and 31B params. Too small for serious coding or deep planning. Good for quick prototyping only.
Best starting pick for each role, based on benchmarks + practical dual-use. Many cloud flagships work for both plan and implement — the card shows the usual default, not the only option.
| You do… | Best pick | Why | Type |
|---|---|---|---|
| Senior Software Engineer | Fable 5 high · Opus 4.8 high | SWE-Pro / dual Claude · review with the other | ☁️ Cloud |
| AI Agent Developer | GLM-5.2 · Sol high/xhigh | Local TB 81% / cloud terminal ceiling | 🦙 Ollama / ☁️ Cloud |
| Full-stack Developer | Sol high · Terra high | UI/design + default implement | ☁️ Cloud |
| Large refactor | Fable 5 xhigh | Long-horizon multi-file migrations | ☁️ Cloud |
| SAP / MES / ERP | Fable high · Terra high | Plan dual Claude · implement cheap clear PRs | ☁️ Cloud |
| WordPress / WooCommerce | Sol high | UI + web automation | ☁️ Cloud |
| Million-token RAG | Kimi K3 | 1M context, strong value | ☁️ Cloud |
| Local / Open | Qwen3.5 · GLM-5.2 | Plan vs code split (not dual) | 🦙 Ollama |
| Budget strong | Terra med/high · DeepSeek | Cloud cheap implement + free local | ☁️ Cloud / 🦙 Ollama |
| Content & Marketing | MiniMax M3 | Native multimodal assets | 🦙 Ollama |
If you only remember one pipeline: plan with Sol/Fable/Opus high → implement with Terra high (or the same dual model in a new session) → review with a different model. Raise to xhigh only for multi-file / multi-hour / migrations.