LOGBOOK

Muse Spark 1.2 deserves a place in our AI toolkit

Meta's new model combines strong coding performance, a one-million-token context window, and competitive costs. Interesting enough to use, but not to trust blindly.

Muse Spark 1.2 deserves a place in our AI toolkit

Meta released Muse Spark 1.2 on 5 August 2026. The model was specifically improved for software development and works with Muse Code, Meta's new terminal coding agent.

Muse Spark 1.2 is not a new undisputed market leader. It is, however, a remarkably efficient specialist that comes close to the strongest models on some coding tasks. That is why we want to add it to our toolkit alongside models from OpenAI, Anthropic, Google, and Moonshot.

What makes Muse Spark 1.2 interesting?

Muse Code can investigate large codebases, plan changes, write code, and validate the result. It can run multiple background agents and keeps model calls, tool actions, approvals, and changes in an event log. This makes interrupted tasks resumable and improves traceability.

Meta trained Spark 1.2 together with this environment. The model learned to work with background agents, context compaction, and long-running tasks. Meta says some optimisation tasks involved more than one thousand tool calls and ran for up to 24 hours. That makes it particularly interesting for larger investigations, migrations, debugging, and improvements that require many steps.

Its one-million-token context window helps with large repositories and extensive documentation. This does not mean the model automatically understands an entire codebase, but it can process considerably more relevant information at once.

What do independent benchmarks say?

Artificial Analysis gives Muse Spark 1.2 an Intelligence Index score of 54. That is three points above Spark 1.1 and eleven points above Spark 1.0. It is roughly level with GPT-5.5 xhigh and Grok 4.5 high, while remaining behind Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, and Kimi K3.

On GDPval-AA v2, a benchmark for professional work, Spark rose from an Elo score of 1371 to 1631. Artificial Analysis ranks it fifth among the models it tested. Its Terminal-Bench 2.1 result increased from 78 to 80 percent.

Costs are relatively low at the same time. Artificial Analysis calculated about 0.40 dollars per Intelligence Index task. It is not the cheapest model in every comparison, but it is competitive within this performance class.

Not every result improved. Science benchmarks remained mostly flat. The measured hallucination rate also fell partly because the model attempted fewer answers. Inventing less is positive, but excessive caution during development can also leave work unfinished.

The benchmarks need context

Meta reports a score of 82.9 percent on Terminal-Bench 2.1 and 59.3 percent on DeepSWE 1.1. These are strong results, but the comparison is not completely equal.

Each model was tested in the coding agent Meta considered most suitable. Spark used Muse Code, Claude used Claude Code, and OpenAI models used Codex. Spark 1.2 was also co-trained with Muse Code. The outcome therefore measures not only the model, but the combination of model, tools, prompts, and agent environment.

Meta acknowledges that third-party environments may not be optimally configured for every model. The results mainly show how well each complete product works. They do not prove that Spark will perform equally well in every other agent or existing development environment.

Why we want to use Muse Spark 1.2

We deliberately work tool-agnostically. One model may be better at architecture, another at visual work, analysis, debugging, or rapid implementation. The current AI arms race means these differences keep changing.

Muse Spark 1.2 looks particularly useful for analysing large codebases, complex bugs, controlled migrations and refactors, parallel scoped tasks, long-running optimisation, and projects where both cost and performance matter.

We do not intend to make it our default immediately. We want to compare it with our current models on real project tasks. We will measure correct results, review time, introduced defects, speed, and total cost rather than looking only at benchmark scores.

A cheaper model can become more expensive when it requires more human correction. Conversely, a slightly lower benchmark score can be perfectly acceptable when a model behaves predictably and is easy to review.

What are the risks?

Muse Spark 1.2 is available through a closed Meta API and remains an early product. Pricing, availability, terms, and behaviour may change. This creates vendor dependency risk.

Privacy needs special attention. Meta offers different usage options and discounts may come with different data-use terms. We would never process confidential client code through a training or contributor programme. Before professional use, current retention periods, training terms, processing locations, and contractual protections must be clear.

An autonomous coding system also introduces practical security risks. Code, documentation, or dependencies may contain malicious instructions. An agent with excessive permissions could delete files, access secrets, or introduce unsafe changes. Such systems should therefore run in restricted environments without production data or permanent access keys.

Human oversight remains essential. Every change must be reviewed, tested, and subjected to security checks where appropriate.

Meta previously published a safety report for Muse Spark covering areas including cybersecurity and loss of control. Meta considered the residual risks acceptable within its own framework. At the time of writing, however, we could not find a separate and equally detailed safety report for the more capable 1.2 release.

Our conclusion

Muse Spark 1.2 is not a reason to replace every other model. It is a serious and cost-efficient new specialist.

We want to use it selectively for suitable coding tasks, inside a restricted environment and with human review. Real project results will then determine whether it deserves a permanent place in our stack.

That is how we believe rapidly changing AI technology should be approached: not by remaining loyal to one brand, but by choosing the best tool for each assignment.

Sources: Meta on Muse Code and Muse Spark 1.2, Meta's evaluation methodology, Artificial Analysis' independent assessment, and the Muse Spark Safety and Preparedness Report.

Facts and availability checked on 11 August 2026.

All hands on deck?

Tell us what you're building. We'll tell you what we can ship by Friday.