← All posts Abstract illustration of three wireframe tools of different sizes standing side by side

At a glance

  • Claude Fable 5 is strongest where difficult coding work justifies paying for frontier-model quality.
  • Less expensive models can be better value for routine implementation and high-volume workloads.
  • Enterprise selection should compare quality, latency, refusal behaviour and security—not benchmark score alone.
  • The right answer is usually model routing by task rather than using one premium model for everything.

Claude Fable 5 leads every published coding benchmark in this comparison. It also costs $10 per million input tokens and $50 per million output. That is twice Anthropic's price for Opus 4.8. A leaderboard position alone does not justify the premium; the question is which work benefits enough to pay for it.

The short answer: yes for complex, multi-file work where correctness dominates cost. Fable 5 scores 90.0 percent on FrontierSWE. This benchmark measures whether a model resolves real GitHub issues under a standardised agent scaffold. Fable leads Claude Opus 4.8, GLM-5.2 and GPT-5.5 by about 15 points. Those three cluster within 2.5 points in the mid-70s.

The premium is harder to defend for high-volume, lower-stakes work. Z.ai's GLM-5.2 costs roughly one-seventh as much and performs close to Opus 4.8. A split estate is therefore credible: use Fable 5 for difficult work and a cheaper model for volume. Data residency and supplier jurisdiction still need separate assessment.


Read the Benchmark Numbers Sceptically

Fable 5's headline 80.3 on SWE-bench Pro comes from a tuned vendor scaffold. Most published scores on this harder benchmark are self-reported in the same way. Scale AI's standardised leaderboard runs models on identical scaffolding and produces lower results.

The SWE-bench Pro tracking gives Opus 4.8 a 69.2 percent vendor aggregate. Scale's best standardised Claude run scored 51.9 on the public set. A result 10 to 30 points above a standardised run may reflect the scaffold as much as the model. Fable 5's 80.3 needs the same caution.

We therefore ran a 28-task evaluation with identical prompts and grading. Fable 5 produced the strongest answers when it attempted the task, but refusals reduced completed work.

GLM-5.2's 62.1 has its own caveat: it was measured by third parties, as Z.ai published no SWE-bench figure at launch. The honest reading is narrower than the headlines: GLM-5.2 beats GPT-5.5 on like-for-like runs and sits within a point of Opus 4.8 on FrontierSWE, where the scores come from one analysis. Fable 5's FrontierSWE lead is not a scaffold artefact, though — it holds on the standardised comparison the other numbers lack.

FrontierSWE scores across four models Four horizontal bars on a 0 to 100 scale. Fable 5 leads at 90; Opus 4.8, GLM-5.2 and GPT-5.5 sit within 2.5 points of each other in the mid-70s. Fable 5 Opus 4.8 GLM-5.2 GPT-5.5 Claude Fable 5 — FrontierSWE: 90.0 Claude Opus 4.8 — FrontierSWE: 75.1 GLM-5.2 — FrontierSWE: 74.4 GPT-5.5 — FrontierSWE: 72.6 90.0 75.1 74.4 72.6 0 20 40 60 80 100
FrontierSWE, percent. GLM-5.2 sits level with Opus 4.8; Fable 5 is the outlier. Sources: BenchLM and VentureBeat, June 2026.

The practical rule: trust orderings within a single standardised run, treat cross-source comparisons as directional, and pilot on your own repositories before believing any of it.


What Each Model Offers an Engineering Organisation

Claude Fable 5: the quality ceiling, priced accordingly

Fable 5 leads FrontierSWE at 90.0 percent, about 15 points above every other model here (BenchLM). It is the strongest choice when correctness on complex, multi-file work dominates cost.

The ecosystem strengthens that case. Claude Code received a 46 percent “most loved” rating in The Pragmatic Engineer's survey of 906 developers. Community reports also favour the integration between a tool and model produced by the same lab.

The premium is steep: $10/$50 per million tokens, double Opus 4.8, which stays available at $5/$25 as Anthropic's own mid-tier (batch processing halves Fable to $5/$25). The economical pattern is a split estate: Fable for planning, architecture, and the hardest debugging; a cheaper model for volume. Until June the volume tier meant Haiku- or Sonnet-class models. GLM-5.2 arriving at Opus-class quality changes what the volume tier can be.

One caveat without precedent: on 12 June a US export-control directive forced Anthropic to suspend Fable 5 for all customers. Access returned on 30 June when the controls were lifted. Two weeks of a leading model switched off by government order is a concentration-risk data point for any single-vendor estate, whichever jurisdiction it sits in.

GLM-5.2: the budget alternative, close enough to matter

GLM-5.2 costs $1.40 per million input tokens and $4.40 for output. Cached input costs $0.26. Fable 5 costs $10/$50 and Opus 4.8 costs $5/$25 (llm-stats comparison).

It trails Fable 5 by 15.6 points on FrontierSWE but sits close to Opus 4.8. For long refactors and repository-wide analysis, that price gap compounds daily. GLM-5.2 is worth testing as the volume model in a split estate.

Switching costs are the under-reported fact. Z.ai exposes an Anthropic-compatible endpoint, so GLM-5.2 drops into Claude Code, Cline, and Cursor by changing a base URL and key. In Claude Code the 1M-context variant is glm-5.2[1m] (setup guide). Teams with agentic tooling already built around Claude, as in our coding tool comparison, can trial GLM-5.2 alongside Fable 5 in the same workflow in an afternoon. The compatibility also means adopting it creates no lock-in — it slots in or out without a tooling migration.

The weaknesses mirror the price. No vendor with a UK legal presence stands behind the output, the benchmark record is thinner, and the data-residency analysis must precede any pilot on real code. The managed-cloud route is not built yet. At the time of writing, AWS Bedrock's catalogue carries GLM 5 rather than 5.2, and Microsoft Foundry's Fireworks listing only reaches GLM 5.1 — so procurement cannot simply tick the usual hyperscaler box for 5.2. Early adopters also report migration friction: model-id conventions differ by tool, and routing through intermediaries has produced tool-call errors in Cursor.

GPT-5.5: behind on these benchmarks, ahead on distribution

GPT-5.5 trails Fable 5 by a wide margin and GLM-5.2 by a smaller one — 58.6 on SWE-bench Pro, 72.6 percent on FrontierSWE. Coding benchmarks are not the whole of engineering work, though. Its strength is sustained multi-step execution, as we covered in our GPT-5.5 analysis, and it is the default model in more enterprise tooling than either rival. It is not the budget option: at $5/$30 per million tokens its output costs more than Opus 4.8's, so the case is the ecosystem, not the price. For organisations standardised on OpenAI through Microsoft, the question is not whether GPT-5.5 tops a leaderboard but whether the gap justifies leaving an integrated estate.


The Comparison at a Glance

Claude Fable 5 GLM-5.2 Claude Opus 4.8 GPT-5.5
FrontierSWE 90.0% 74.4% 75.1% 72.6%
SWE-bench Pro 80.3 (vendor scaffold) 62.1 (third-party run) 69.2 (vendor scaffold — not comparable) 58.6
Price per MTok (in/out) $10 / $50 (~$1 cached; batch $5/$25) $1.40 / $4.40 ($0.26 cached) $5 / $25 ($0.50 cached) $5 / $30 ($0.50 cached)
Context window 1M tokens 1M tokens 1M tokens 1M listed; 2x input price above 272K
Access API and vendor tools only API, coding plans, MIT weights; no first-party hyperscaler offering yet API and vendor tools only API and vendor tools only
Vendor accountability US/UK contractual presence Chinese jurisdiction; DPA offered US/UK contractual presence US/UK contractual presence

Cost per Token Is Not Cost per Task

The price gap is per token, but engineering organisations buy completed work. If a cheaper model needs two attempts where a stronger model needs one, much of the saving disappears. Engineer review time narrows it further.

On FrontierSWE, GLM-5.2 sits 0.7 points behind Opus 4.8 and 15.6 behind Fable 5. The gap on your repositories may differ, and no public retry-adjusted cost data exists. A two-week pilot can measure completed-task cost directly.

Price per million tokens across four models Horizontal bars on a 0 to 50 dollar scale. Fable 5 output costs $50, GPT-5.5 $30, Opus 4.8 $25, GLM-5.2 $4.40. INPUT, $/MTOK OUTPUT, $/MTOK GLM-5.2 Opus 4.8 GPT-5.5 Fable 5 GLM-5.2 Opus 4.8 GPT-5.5 Fable 5 GLM-5.2 — input: $1.40 per million tokens Claude Opus 4.8 — input: $5.00 per million tokens GPT-5.5 — input: $5.00 per million tokens Claude Fable 5 — input: $10.00 per million tokens GLM-5.2 — output: $4.40 per million tokens Claude Opus 4.8 — output: $25.00 per million tokens GPT-5.5 — output: $30.00 per million tokens Claude Fable 5 — output: $50.00 per million tokens $1.40 $5.00 $5.00 $10.00 $4.40 $25.00 $30.00 $50.00 $0 $10 $20 $30 $40 $50
List prices per million tokens. Cached input: GLM-5.2 $0.26, Opus 4.8 $0.50, GPT-5.5 $0.50, Fable 5 about $1. Fable 5 batch runs $5/$25. Sources: llm-stats, OpenAI's published price table and Anthropic pricing docs, June 2026.

Caching changes the totals, but not the ordering. Agentic coding repeatedly reads the same files, and all four models discount cached input. Fable 5 falls to about $1 per million tokens, Opus 4.8 and GPT-5.5 to $0.50, and GLM-5.2 to $0.26.

GLM-5.2 therefore retains its price advantage. Real workloads can run well below list price, as our coding-tool comparison also shows.


Which Model for Which Team


New posts land here first. Sign up for email updates.

Frequently asked questions

Is Claude Fable 5 worth the price for coding?

On published benchmarks, yes for complex work. Fable 5 leads FrontierSWE at 90.0 percent, some 15 points clear of Claude Opus 4.8, GLM-5.2 and GPT-5.5, which all sit within 2.5 points of each other in the mid-70s. The premium is steep — $10/$50 per million tokens, double Opus 4.8 — and is best justified for planning, architecture and complex, multi-file work where correctness dominates cost. For high-volume, lower-stakes work, a cheaper model in a split estate is usually the better economics.

How much does Claude Fable 5 cost compared to alternatives?

Fable 5 costs $10 per million input tokens and $50 per million output — roughly 7x GLM-5.2's input price and 11x its output price, and double Claude Opus 4.8's $5/$25. Cached input narrows this to about $1 for Fable 5, against $0.50 for Opus 4.8 and GPT-5.5, and $0.26 for GLM-5.2. Batch processing brings Fable 5 down to $5/$25 — Opus 4.8's standard list price, though Opus 4.8's own batch rate is cheaper still at $2.50/$12.50. Per-token price is not per-task cost — a model that gets a complex task right first time can be cheaper overall than a cheaper model that needs several attempts.

Why do Claude Fable 5's benchmark scores differ between sources?

Because the scores are produced on different scaffolds. Fable 5's 80.3 on SWE-bench Pro is a vendor-reported figure from a tuned agent scaffold. Scale AI's standardised leaderboard, which runs every model on identical scaffolding, produces substantially lower scores across the board — Opus 4.8's vendor aggregate of 69.2 falls to 51.9 on Scale's standardised public set. Treat cross-model comparisons as directional unless the scores come from the same standardised run.

Is GLM-5.2 a viable alternative to Claude Fable 5 for coding?

For high-volume, cost-sensitive workloads, yes. GLM-5.2 trails Fable 5 by 15.6 points on FrontierSWE but costs a fraction of the price — $1.40/$4.40 per million tokens against Fable 5's $10/$50. Z.ai's Anthropic-compatible endpoint lets it drop into Claude Code, Cline and Cursor by changing a base URL and key. It is not a replacement for Fable 5 on the hardest work, but a credible second model in a split estate where GLM-5.2 handles volume and Fable 5 handles what needs to be right first time.

← All posts