
On September 2, Meta launched its most powerful AI model to date, Muse Spark 1.3, designed specifically for extended agentic workflows and enhanced coding capabilities. Meta's Chief AI Officer Alexandr Wang stated that the new model outperforms OpenAI's GPT-5.6 Sol in coding and performs on par with Anthropic's Claude Fable 5.1.
Performance Leap: Knowledge Work, Long Context, and Coding All Improved
Across multiple benchmarks, Muse Spark 1.3 demonstrates comprehensive advantages over its predecessor and competitors:
| Benchmark | Muse Spark 1.3 | Muse Spark 1.2 | GPT-5.6 Sol | Opus 5 |
|---|---|---|---|---|
| GDPVal-AA v2 (Knowledge Work) | 1754 | 1615 | 1710 | 1824 |
| DeepSWE v1.1 (Long-Horizon Coding) | 75.4 | 55.0 | 73.0 | 74.0 |
| Terminal-Bench 2.1 (Terminal Coding) | 88.8 | 82.9 | 88.8 | 86.7 |
| MRCR 256K-512K (Long Context) | 98.5 | 66.3 | 91.5 | — |
| MRCR 512K-1M (Long Context) | 98.1 | 55.5 | 73.8 | — |
| OSWorld 2.0 (Agentic Computer Use) | 66.9 | 47.6 | 62.7 | 68.3 |
| Agentic IF Index (Instruction Following) | 57.8 | 46.2 | 60.5 | 59.1 |
According to Artificial Analysis, Muse Spark 1.3 (xhigh) achieved an Intelligence Index of 61, placing it on par with GPT-5.6 Sol (max) and Grok 4.6 (high). The limited-preview max version reached 62, trailing only select variants of Claude Fable 5.1 and Claude Opus 5.
Notably, Meta acknowledges that competitors still lead in several benchmarks, and direct comparisons are limited—different models excel at different tasks, and benchmark parameters may be optimized selectively.
Efficiency Revolution: 20% Fewer Tool Calls, 25% Less Token Consumption
Muse Spark 1.3 delivers significant runtime efficiency gains:
Tool calls reduced by approximately 20%: The model takes fewer detours in coding tasks, reducing unnecessary iteration loops
Token consumption reduced by approximately 25%: Fewer tokens needed to complete the same tasks
Cleaner, more concise code: Overall coding style is more streamlined
For long-horizon agentic tasks, the model can manage multiple workflows within a single thread, proactively fixing planning gaps and tracking learning outcomes. When facing ambiguous instructions, it proactively asks clarifying questions; when encountering obstacles, it requests user assistance; before executing critical actions, it seeks confirmation.
In an aerospace engineering fluid simulation example, the model demonstrated complete long-horizon execution: starting from CFD simulation results and CAD models in STEP format, it autonomously organized analysis goals, extracted parameters, conducted lift and drag analysis, and output a structured PDF report.
Unchanged Pricing, "Contributor SKU" for Data Access
Pricing remains unchanged from the previous version:
Input: $1.25 per million tokens
Cache-hit input: $0.15 per million tokens
Output: $4.25 per million tokens
The new model is now available on Muse Code and the Meta Model API, and will be gradually integrated into Instagram, Facebook, and Meta AI Assistant. Some developers are already consuming trillions of tokens per week via the API.
Meta also introduced a "Contributor SKU" priced at 1/12 to 1/21 of the standard rate ($0.10/million input, $0.20/million output), in exchange for usage data that may be used to train Meta's models. Alexandr Wang noted that a "double-digit percentage" of developers have already opted for this version.
Muse Spark 1.3's core upgrade isn't about benchmark breakthroughs—it's about "getting the same work done with fewer detours." The dual reduction in tool calls and token consumption is precisely what makes agents practical. Four updates in five months show Meta is using high-frequency iteration to catch up. The Contributor SKU reveals another trend: as model capabilities converge, data itself is becoming AI companies' scarcest strategic resource.