Meta Superintelligence Labs shipped the fourth Muse Spark update in five months on September 3. The company says Muse Spark 1.3 completes agent tasks with about a fifth fewer tool calls and a quarter fewer tokens than version 1.2, and it now appears in Muse Code and the Meta Model API.
Weights stay closed, and the strongest reasoning tier remains locked behind further safety testing. To keep skills from locking onto a single toolchain, Meta spread training across several agent harnesses.
Collaboration is the focus of this release. The model asks for clarification on vague prompts, pulls the user in when it stalls, confirms before consequential actions, and adapts its update style on long runs. Meta also says calibration improved, so the model flags its own limits instead of inventing outcomes.
The benchmark story is competitive. Meta reports 75.4 on the DeepSWE v1.1 coding test, ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 72.7, and wide leads on long-context retrieval, where an MRCR v2 score of 98.5 at 256K to 512K tokens beats Sol’s 91.5. Artificial Analysis independently places the shipping variant level with GPT-5.6 Sol on its intelligence index.
For agent workloads the efficiency numbers carry the real weight, since fewer round trips and fewer billed tokens cut the cost of each completed task.