Meta logo

Muse Spark 1.3

Multimodal
Meta

Muse Spark 1.3 is Meta Superintelligence Labs' proprietary multimodal reasoning model for long-running agentic, multi-agent, and coding workflows. It improves on Muse Spark 1.2 in long-horizon collaboration, multitasking, tool use, failure recovery, and concise coding execution, supports a 1M-token context window, and is available through Muse Code and the Meta Model API at $0.10 per million input tokens.

Key Specifications

Parameters
-
Context
1.0M
Release Date
September 2, 2026
Average Score
73.4%

Timeline

Key dates in the model's history
Announcement
September 2, 2026
Last Update
September 3, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.20
Max Input Tokens
1.0M
Max Output Tokens
943.7K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
Agentic IF Index (Internal)
Meta internal composite instruction-following evaluation for agentic workflows.Self-reported
57.8%
AutomationBench
AutomationBench public v3 task set; deterministic end-state grading; pass@1.Self-reported
49.4%
DeepSearchQA
Official 900-question dataset with a common browser harness; F1 score.Self-reported
89.4%
DeepSWE 1.1
DeepSWE v1.1 official 113-task set with mini-swe-agent; max reasoning; task pass rate.Self-reported
75.4%
GDPval-AA
GDPval-AA v2 Elo in the Artificial Analysis Stirrup agentic harness; human baseline 1,000 (max_score 3000).Self-reported
58.5%
Job Bench
Official 65-task set with a file-aware rubric grader; mean rubric score.Self-reported
64.9%
MRCR v2 (8-needle)
OpenAI MRCR v2 8-needle, 256K-512K context range; 100 examples; mean sequence-matcher ratio.Self-reported
98.5%
MRCR v2 (8-needle, 512K-1M)
OpenAI MRCR v2 8-needle, 512K-1M context range; 100 examples; mean sequence-matcher ratio.Self-reported
98.1%
OSWorld 2.0
OSWorld 2.0 version 08.08 in Meta's common internal computer-use evaluation framework; mean per-task partial score.Self-reported
66.9%
SWE Atlas - Codebase QnA
SWE-Atlas public Codebase QnA split with mini-swe-agent and task-specific rubric grading; mean pass@1.Self-reported
59.4%
Terminal-Bench 2.1
Terminal-Bench 2.1 official 89-task set in an isolated cloud sandbox; mean pass@1.Self-reported
88.8%

License & Metadata

License
proprietary
Announcement Date
September 2, 2026
Last Updated
September 3, 2026

Compare Muse Spark 1.3

All comparisons

Articles about Muse Spark 1.3

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.