Claude Opus 5 review 2026 — Anthropic's agentic coding value king
⚡ TL;DR

Claude Opus 5 is the most rational flagship buy of 2026: near-Fable performance at exactly half the price. Released July 24, 2026, Anthropic's new Opus flagship delivers Fable-5-level agentic coding — #1 on BenchLM agentic (80.5), #1 on Frontier-Bench, IMO 2026 perfect 42/42, 1M context — while matching Opus 4.8's $5/$25 per-million-token price. In CursorBench 3.2 it lands just 0.5% behind Fable 5's peak at less than half the task cost. The trade-offs: no official China access, still the priciest tier on a per-task basis (about 39x DeepSeek V4 Pro), and a tendency to stubbornly persist on wrong interpretations in long analytical runs. Buy it for long-horizon agent work and hard engineering; for high-volume budget calls, use a domestic model.

1. What Is Claude Opus 5?

On July 24, 2026 — about three weeks after the Fable 5 and Sonnet 5 releases — Anthropic shipped Claude Opus 5, its new top-of-the-Opus-line flagship. The positioning is deliberately simple: near-Fable-5 capability at half the price, on the same price point as the previous Opus 4.8 ($5/$25).

It is built for long-horizon autonomous agents, complex engineering, and knowledge work: cross-stage planning, sub-agent delegation, and self-verification across tasks that run for hours to days. It's the default "smart" choice inside Claude Code when you want maximum capability per dollar.

📌 Quick Numbers
  • 1M token context, max output 128k
  • 5 effort tiers (low/medium/high/xhigh/max) — pay for what you need
  • #1 BenchLM agentic (80.5), #1 Frontier-Bench, IMO 2026 perfect 42/42
  • $5 / $25 per million tokens — same as Opus 4.8, half of Fable 5
  • FastMode: 2x price for ~2.5x speed (API research preview, platform-dependent)

One-line take: "The ceiling belongs to Fable 5; every dollar below the ceiling is Opus 5's."

2. Core Features (2026)

2.1 Long-Horizon Agents — the Real Differentiator

Opus 5 is built to finish hours-long autonomous jobs: it plans across stages, delegates to sub-agents, and self-verifies. Independent tests show it refactoring coupled codebases and spinning up containers on its own when a service is missing — end to end with no human hand-holding.

2.2 Five Effort Tiers

Choose low → max per task. Daily Q&A uses the cheap tiers; cross-module migrations get xhigh/max. Even the low tier beats early Opus models, so don't copy old per-tier habits. Note: disabling thinking caps you at high — max power requires thinking enabled.

2.3 1M-Token Context + FastMode

Read an entire mid-size repository (≈75k lines) in one pass. For latency-sensitive production, FastMode doubles the price (~$10/$50) for ~2.5x speed — currently an API research preview, not available everywhere.

2.4 Benchmark Snapshot

BenchmarkOpus 5vs. Rivals
BenchLM Agentic80.5 (#1)Qwen3.8 Max 78.8, Mythos 5 75.6
Artificial Analysis Index v4.1.163 (#1)GPT-5.6 Sol 61, Kimi K3 59.7
IMO 202642/42perfect
Chinese eval (ReLE ~15k)76.8% (#2)up from 74.7% (#9)
Frontier-Bench v0.1#12x Opus 4.8 at equal cost
CursorBench 3.270.0%Fable 5 70.5%, half task cost
OSWorld 2.0#1 at any cost1/3 of Fable 5's cost to match

2.5 Chinese-Language Leap

On a ~15k-question Chinese benchmark, Opus 5 jumped from #9 to #2 (76.8%), with big gains in education (+5.0pt), medical (+3.7pt), and legal (+3.4pt). Response speed improved 21% and token usage dropped 25% vs. the previous generation.

3. Claude Opus 5 Pricing (2026)

API per-million-token (USD)

ModeInputOutputNote
Standard$5$25same as Opus 4.8, half of Fable 5
Fast (preview)$10$50~2.5x speed, platform-dependent

Subscriptions (include Claude Code)

PlanPrice/moUsage
Pro$20base quota
Max 5x$1005x quota, set Opus 5 as default
Max 20x$20020x quota, Agent Teams
Team$100includes Claude Code

Real per-task cost (the honest column)

ModelCost / task
Claude Opus 5$2.34
GPT-5.6 Sol$1.23
Kimi K3$0.84
GLM-5.3$0.68
DeepSeek V4 Pro$0.06

Opus 5 is still one of the most expensive mainstream models per task — about 39x DeepSeek V4 Pro. You pay for getting the hardest job right the first time.

⚠️ Cost Reality Check

At $2.34 per task, high-volume batch calls with Opus 5 will drain budgets fast. For frequent, repetitive workloads, route them to a cheap model (DeepSeek V4, GLM) and save Opus 5 for the tasks that actually need frontier reasoning.

4. Strengths — Where It Shines

✅ Best Performance-per-Dollar Among Flagships

Near-Fable ceiling (CursorBench 0.5% gap) at exactly half the token price — the rational pick when you want flagship power without the flagship bill.

✅ Long-Horizon Agents That Finish

Hours-long autonomous refactors, migrations, and debugging with sub-agent delegation and self-verification.

✅ Frontier Reasoning

IMO 2026 perfect 42/42; ARC-AGI 3 scores 3x the runner-up on unknown-problem reasoning.

✅ Strong Chinese-Language Ability

Now #2 on the Chinese comprehensive benchmark — a big deal for Chinese-speaking teams evaluating overseas flagships.

✅ Claude Code Integration

One subscription covers terminal CLI, IDE, and GitHub Actions; Max users can make Opus 5 the default model.

5. Weaknesses — The Fine Print

❌ No Official China Access

claude.ai is blocked in mainland China (availability scored ~2/10). You need an overseas node + payment, or a relay/reseller. For mainland teams, GLM / Kimi / DeepSeek remain the low-friction path.

❌ Still the Priciest Tier Per Task

~39x DeepSeek V4 Pro on per-task cost. Great for hard problems; wasteful for simple Q&A or summaries.

❌ "Stubbornness" on Long Analytical Runs

In an AA-AnalystAgent test requiring 5 consecutive correct answers, Opus 5 passed only 54% of questions; ~57% of failures came from misinterpreting the question early and persisting. Confident answers still need human review.

❌ Closed-Source, No Free Tier

API/subscription only, not self-hostable — no data sovereignty like open-weight models.

❌ Thinking-Lock & FastMode Limits

Disabling thinking caps effort at high; FastMode is a research preview and not available on every cloud platform.

6. Opus 5 vs Fable 5 vs GPT-5.6 vs Kimi K3 vs DeepSeek V4 (2026)

DimensionOpus 5Fable 5GPT-5.6 SolKimi K3DeepSeek V4 Pro
Intelligence index63 (#1)higher tier6159.753.2
Long-horizon agentStrongStrongestStrongStrongMedium
Context1M1M1M1M1M
Cost / task$2.34~$5+$1.23$0.84$0.06
Input $/1M$5$10$3$1.32
Output $/1M$25$50$15$3.96
China direct
RolePerformance+valueCeilingAll-round flagshipChina perf kingValue king
🧭 How to Choose
  • Hard engineering / long agents at sane costOpus 5 (the sweet spot)
  • Absolute ceiling, unlimited budgetFable 5
  • Budget-sensitive, high-volume, or self-hostDeepSeek V4 / GLM / Kimi
  • Mainland users wanting frontier-ish codingKimi K3 (open weights, #1 Chinese open model)

7. Who Should (and Shouldn't) Use Opus 5

Full-stack / backend / DevOps
Long-horizon agents + Claude Code: big refactors, migrations, complex debugging.
Agent builders
Multi-hour autonomous workflows, sub-agent delegation, cross-module orchestration.
Enterprise docs / finance / legal
Chinese-language document processing jumped; evaluate via relay for compliance.
Research / math / hard reasoning
IMO 42/42, ARC-AGI 3x the runner-up — frontier math and unknown-problem reasoning.
Budget-sensitive devs
~39x DeepSeek V4 Pro per task; high-volume calls will hurt.
Mainland users without overseas setup
No official access; strong risk-control on accounts; compliance questions.
Pure GUI beginners
CLI/API-first; start with Cursor or Trae if you don't do terminal.

8. Final Verdict & Scores

Coding / Agent Ability
★★★★★
Value for Money
★★★★★
Chinese Language
★★★★
Ease of Use
★★★
China Accessibility
★★
🏁 Bottom Line

Claude Opus 5 is the flagship to actually buy in 2026. It delivers ~95% of Fable 5's practical capability at half the token price and the same price as the previous Opus — a rare "get more for the same" upgrade. Use it for long-horizon agents, hard engineering, and frontier reasoning; keep high-volume budget calls on domestic models. The stubbornness on long analytical runs and the China access gap are real, but they don't change the calculus: this is the value king of the flagship tier.

9. FAQ

Q1: Opus 5 or Fable 5 — which should I pick?
If you need the absolute ceiling and don't care about cost, Fable 5. If you want flagship capability at a sane price, Opus 5 — it trails Fable 5's peak by only 0.5% on CursorBench 3.2 at less than half the task cost.
Q2: How much does Claude Opus 5 cost?
API: $5 input / $25 output per million tokens (same as Opus 4.8, half of Fable 5). FastMode is $10/$50. Subscriptions: Pro $20/mo, Max 5x $100/mo, Max 20x $200/mo — all include Claude Code.
Q3: When was Claude Opus 5 released?
July 24, 2026 — about three weeks after Fable 5 and Sonnet 5. It replaces Opus 4.8 as Anthropic's Opus-tier flagship.
Q4: Is Claude Opus 5 good at coding?
Yes — arguably its best use case. #1 on Frontier-Bench and BenchLM agentic, IMO 2026 perfect, deep Claude Code integration, and long-horizon agents that finish multi-hour refactors and migrations on their own.
Q5: Can I use Claude Opus 5 from China?
Officially no direct access. You need an overseas node + account + payment, or a relay/reseller. For mainland teams, GLM, Kimi K3, and DeepSeek V4 remain the low-friction, CNY-priced alternatives.
Q6: What are Opus 5's weaknesses?
No official China access, still the priciest tier per task (~39x DeepSeek V4 Pro), a tendency to stubbornly persist on wrong interpretations in long analytical runs, closed-source with no free tier, and FastMode limits on some cloud platforms.
Q7: How does Opus 5 compare with DeepSeek V4?
Opus 5 is the intelligence and agent ceiling; DeepSeek V4 Pro is the value king (about $0.06 vs $2.34 per task — a ~39x gap). Unlimited budget → Opus 5; pragmatic high-volume or self-hosted → DeepSeek.
🏷 Tags: Claude Claude Opus 5 Anthropic AI Review AI Tools LLM AI Coding AI Agent Claude Opus 5 Review AI Review 2026