Meta's Muse Code ships with a crash-safe log but Claude Opus 5 stays ahead on benchmarks

Meta released Muse Code (beta), a terminal-based coding agent, on Wednesday. The company’s own charts show its model beating OpenAI’s Codex and Google’s Antigravity on most coding tests but losing to Anthropic’s Claude Opus 5 on every benchmark Meta released.
Muse Code trails Opus 5 but adds a crash-safe log
Muse Code is powered by Muse Spark 1.2, an update to the coding model Meta opened up to U.S. developers in July.
Muse Spark 1.2 lags behind Opus 5 on the coding benchmarks Meta presented but outperforms Codex and Antigravity on most of them. Meta said the model was its “next step toward the frontier, with larger and much more capable models on the way.”
The company says that Version 1.2 is better at code generation, debugging, and understanding large codebases after it scaled up compute spent on coding tasks during training.
In one case study, the model rewrote GPU kernels for NVIDIA Hopper chips over more than 1,000 tool calls and for as long as 24 hours, working from a baseline it was told not to copy from existing libraries.
… Continue reading the full article at the original source below.



