2026年8月21日星期五

Meta Muse Spark 1.2 Review 2026: Superintelligence Labs' First Model, Coding #2 to Opus 5, $0.10 Contributor Tier

Meta Muse Spark 1.2 Review 2026: Superintelligence Labs' First Model, Coding #2 to Opus 5, $0.10 Contributor Tier

Meta Muse Spark 1.2 Review 2026: Superintelligence Labs' First Model, Coding #2 to Opus 5, $0.10 Contributor Tier

Verdict first: Muse Spark 1.2 is Meta Superintelligence Labs' first product — a coding-focused upgrade of Muse Spark 1.1 shipped alongside the Muse Code terminal agent. On Meta's self-reported tests it lands coding at #2 to Claude Opus 5 (Terminal-Bench 82.9% vs 86.7%), with a 1M context window and a shockingly cheap Contributor tier at $0.10/$0.20 (roughly 12-21x below standard). The catch: it's closed-weights, benchmarks are vendor-reported and measured inside Meta's own Muse Code framework, and the Contributor deal trades your data for price. Take the headline numbers with a grain of salt, but the economics are real.

Meta Muse Spark 1.2 review 2026 — coding-focused, 1M context, $0.10 Contributor tier, #2 to Opus 5

Muse Spark 1.2 — Meta Superintelligence Labs' first product, built for coding and agents

What Is Muse Spark 1.2?

Released on August 5, 2026, Muse Spark 1.2 is the first product from Meta Superintelligence Labs — a coding-focused upgrade of Muse Spark 1.1, launched alongside the Muse Code terminal coding agent (installable on macOS and Linux with one command). It's available via Meta Model API, OpenRouter, and seven other providers.

Specifications at a Glance

ItemValue
VendorMeta Superintelligence Labs
Release dateAugust 5, 2026
PositioningCoding-focused flagship + Muse Code terminal agent
ParametersNot disclosed by Meta
Context window1M tokens / up to 131,072 output
InputsText, image, video, audio, PDF → text output
ReasoningAlways on; minimal to xhigh (5 levels)
Open weightsNo (closed); "soon" announced, no date
AvailabilityMeta Model API, OpenRouter + 7 providers

Pricing: The Contributor Deal

TierInput / 1MOutput / 1MCache read / 1M
Standard$1.25$4.25$0.15
Contributor$0.10$0.20$0.002

The Contributor tier (12-21x cheaper) costs you: Meta trains on your data, much stricter rate limits (60 req/min vs 3,000), payment method required, and country restrictions. It's widely read as the continuation of Meta's Llama open strategy — cheap tokens in exchange for training data. No long-context premium; injected steering tokens are free.

Benchmarks: Coding #2 to Opus 5

BenchmarkScore
Terminal-Bench 2.182.9% (self-reported; +6.7 vs 1.1; 80.1 under Artificial Analysis)
DeepSWE v1.159.3% (up from 53.0)
Meta internal coding bench70.6%
AA Intelligence Index54 (57 re-scored; #12/185)
Vals AI independentOverall 5th (71.88%); #1 finance/tax/legal documents

In Meta's own comparison, Muse Spark 1.2 lands #2 to Claude Opus 5 across three coding benchmarks (Opus 5: Terminal-Bench 86.7%). Note: scores were measured inside the co-trained Muse Code framework and are vendor-reported.

The Harness-Lock Controversy

The headline "coding #2 to Opus 5" comes with a big asterisk. Muse Spark 1.2's benchmarks were measured inside its co-trained Muse Code framework, and developers report it is "almost useless" or token-inefficient in generic frameworks like Opencode. This is a classic harness-lock concern: part of the gain may come from the framework, not the raw model. The independent Vals AI test gives overall 5th place but #1 for finance/tax/legal document agents — evidence of real strength in specific domains even outside Meta's harness.

Pros & Cons

Pros

  • ✅ Coding #2 to Opus 5 (self-reported, Terminal-Bench 82.9%)
  • ✅ Cheap: standard $1.25/$4.25; Contributor $0.10/$0.20
  • ✅ 1M context + multimodal input
  • ✅ #1 on finance/tax/legal document agents (independent)
  • ✅ No long-context premium; free steering tokens

Cons

  • ❌ Closed weights — no self-hosting
  • ❌ Benchmarks vendor-reported, no independent replication
  • ❌ Harness-lock: strong mainly inside Meta's Muse Code
  • ❌ No safety disclosure / model card
  • ❌ Contributor tier costs your data + rate limits

Who Should Use It (And Who Shouldn't)

Great fit

  • Developers using Muse Code for terminal coding
  • High-value-agent / document workflows (finance, tax, legal)
  • Teams happy to trade data for the ultra-cheap Contributor tier
  • Anyone needing a 1M context window

Poor fit

  • Users needing open/self-hosted weights (use Muse Glimmer instead)
  • Data-privacy-sensitive users (Contributor tier)
  • Teams needing reliable performance outside Meta's ecosystem

Frequently Asked Questions

Is Muse Spark 1.2 open source?

No — closed weights. Open weights announced "soon" with no date; open-source Muse Glimmer (30B, Apache 2.0) is available now.

How much does it cost?

Standard $1.25/$4.25 per M; Contributor $0.10/$0.20 (trade data for price, strict rate limits).

How big is it?

Parameter count undisclosed. 1M context, 131,072 max output.

Is it really #2 to Opus 5 for coding?

Self-reported Terminal-Bench 82.9% vs Opus 5's 86.7% — yes, but measured in Meta's own framework without independent replication.

What is the context window?

1M tokens; no long-context premium, free steering tokens.

Is it multimodal?

Inputs: text/image/video/audio/PDF; output text. Always-on reasoning with 5 effort levels.

What is Muse Code?

Meta's terminal coding agent, co-trained with Spark 1.2, installed on macOS/Linux with one command.

Sources

没有评论:

发表评论

Microsoft MAI Review 2026: The 7-Model Family That Says Goodbye to OpenAI — Trillion-Parameter Flagship, Now Default in GitHub Copilot

Microsoft MAI Review 2026: The 7-Model Family That Says Goodbye to OpenAI — Trillion-Parameter Flagship, Now Default in GitHub ...