Meta Muse Spark 1.2 Review 2026: Superintelligence Labs' First Model, Coding #2 to Opus 5, $0.10 Contributor Tier
Verdict first: Muse Spark 1.2 is Meta Superintelligence Labs' first product — a coding-focused upgrade of Muse Spark 1.1 shipped alongside the Muse Code terminal agent. On Meta's self-reported tests it lands coding at #2 to Claude Opus 5 (Terminal-Bench 82.9% vs 86.7%), with a 1M context window and a shockingly cheap Contributor tier at $0.10/$0.20 (roughly 12-21x below standard). The catch: it's closed-weights, benchmarks are vendor-reported and measured inside Meta's own Muse Code framework, and the Contributor deal trades your data for price. Take the headline numbers with a grain of salt, but the economics are real.
Muse Spark 1.2 — Meta Superintelligence Labs' first product, built for coding and agents
What Is Muse Spark 1.2?
Released on August 5, 2026, Muse Spark 1.2 is the first product from Meta Superintelligence Labs — a coding-focused upgrade of Muse Spark 1.1, launched alongside the Muse Code terminal coding agent (installable on macOS and Linux with one command). It's available via Meta Model API, OpenRouter, and seven other providers.
Specifications at a Glance
| Item | Value |
|---|---|
| Vendor | Meta Superintelligence Labs |
| Release date | August 5, 2026 |
| Positioning | Coding-focused flagship + Muse Code terminal agent |
| Parameters | Not disclosed by Meta |
| Context window | 1M tokens / up to 131,072 output |
| Inputs | Text, image, video, audio, PDF → text output |
| Reasoning | Always on; minimal to xhigh (5 levels) |
| Open weights | No (closed); "soon" announced, no date |
| Availability | Meta Model API, OpenRouter + 7 providers |
Pricing: The Contributor Deal
| Tier | Input / 1M | Output / 1M | Cache read / 1M |
|---|---|---|---|
| Standard | $1.25 | $4.25 | $0.15 |
| Contributor | $0.10 | $0.20 | $0.002 |
The Contributor tier (12-21x cheaper) costs you: Meta trains on your data, much stricter rate limits (60 req/min vs 3,000), payment method required, and country restrictions. It's widely read as the continuation of Meta's Llama open strategy — cheap tokens in exchange for training data. No long-context premium; injected steering tokens are free.
Benchmarks: Coding #2 to Opus 5
| Benchmark | Score |
|---|---|
| Terminal-Bench 2.1 | 82.9% (self-reported; +6.7 vs 1.1; 80.1 under Artificial Analysis) |
| DeepSWE v1.1 | 59.3% (up from 53.0) |
| Meta internal coding bench | 70.6% |
| AA Intelligence Index | 54 (57 re-scored; #12/185) |
| Vals AI independent | Overall 5th (71.88%); #1 finance/tax/legal documents |
In Meta's own comparison, Muse Spark 1.2 lands #2 to Claude Opus 5 across three coding benchmarks (Opus 5: Terminal-Bench 86.7%). Note: scores were measured inside the co-trained Muse Code framework and are vendor-reported.
The Harness-Lock Controversy
The headline "coding #2 to Opus 5" comes with a big asterisk. Muse Spark 1.2's benchmarks were measured inside its co-trained Muse Code framework, and developers report it is "almost useless" or token-inefficient in generic frameworks like Opencode. This is a classic harness-lock concern: part of the gain may come from the framework, not the raw model. The independent Vals AI test gives overall 5th place but #1 for finance/tax/legal document agents — evidence of real strength in specific domains even outside Meta's harness.
Pros & Cons
Pros
- ✅ Coding #2 to Opus 5 (self-reported, Terminal-Bench 82.9%)
- ✅ Cheap: standard $1.25/$4.25; Contributor $0.10/$0.20
- ✅ 1M context + multimodal input
- ✅ #1 on finance/tax/legal document agents (independent)
- ✅ No long-context premium; free steering tokens
Cons
- ❌ Closed weights — no self-hosting
- ❌ Benchmarks vendor-reported, no independent replication
- ❌ Harness-lock: strong mainly inside Meta's Muse Code
- ❌ No safety disclosure / model card
- ❌ Contributor tier costs your data + rate limits
Who Should Use It (And Who Shouldn't)
Great fit
- Developers using Muse Code for terminal coding
- High-value-agent / document workflows (finance, tax, legal)
- Teams happy to trade data for the ultra-cheap Contributor tier
- Anyone needing a 1M context window
Poor fit
- Users needing open/self-hosted weights (use Muse Glimmer instead)
- Data-privacy-sensitive users (Contributor tier)
- Teams needing reliable performance outside Meta's ecosystem
Frequently Asked Questions
Is Muse Spark 1.2 open source?
No — closed weights. Open weights announced "soon" with no date; open-source Muse Glimmer (30B, Apache 2.0) is available now.
How much does it cost?
Standard $1.25/$4.25 per M; Contributor $0.10/$0.20 (trade data for price, strict rate limits).
How big is it?
Parameter count undisclosed. 1M context, 131,072 max output.
Is it really #2 to Opus 5 for coding?
Self-reported Terminal-Bench 82.9% vs Opus 5's 86.7% — yes, but measured in Meta's own framework without independent replication.
What is the context window?
1M tokens; no long-context premium, free steering tokens.
Is it multimodal?
Inputs: text/image/video/audio/PDF; output text. Always-on reasoning with 5 effort levels.
What is Muse Code?
Meta's terminal coding agent, co-trained with Spark 1.2, installed on macOS/Linux with one command.
没有评论:
发表评论