Muse Spark has been officially launched by Meta Superintelligence Labs as a natively multimodal reasoning model designed to support tool-use, visual chain of thought, and multi-agent orchestration.
The new artificial intelligence model represents the initial step in the company’s updated scaling strategy, which includes infrastructure investments such as the Hyperion data center. Notably, the system integrates visual information across multiple domains to assist with entity recognition, localization, and visual science, technology, engineering, and mathematics (STEM) queries.
In addition to the base model, the developer is rolling out Contemplating mode, a feature that coordinates multiple agents to reason in parallel. This mode achieved a 58% score on the Humanity’s Last Exam benchmark and 38% on FrontierScience Research. Meanwhile, the company plans to deploy this feature gradually to public users.
Technical Capabilities of Muse Spark
The model features specialized training for health-related queries, developed in collaboration with more than 1,000 physicians to curate factual training data. Consequently, the system can generate interactive displays explaining nutritional information and muscle activation during physical exercise. This integration of medical data aims to improve factual accuracy in the health technology sector.
Infrastructure and Pretraining Efficiency
Over the past nine months, the laboratory rebuilt its pretraining stack, modifying the model architecture, optimization protocols, and data curation methods. As a result, Muse Spark requires over an order of magnitude less compute to reach the same performance levels as the previous Llama 4 Maverick model.
Furthermore, the reinforcement learning framework utilizes thinking time penalties to optimize token usage during test-time reasoning. This method forces thought compression, allowing the model to solve complex problems using fewer tokens while maintaining performance.
Safety Evaluations and Alignment
Before deployment, the model underwent safety testing under the updated Advanced AI Scaling Framework to evaluate risks in cybersecurity and autonomous behavior. The evaluations indicated that the model remains within safe margins, showing strong refusal behaviors in high-risk domains such as chemical and biological weapons.
Notably, third-party testing by Apollo Research indicated that the model demonstrated a high rate of evaluation awareness. The system frequently identified testing scenarios as alignment traps, though the researchers concluded this did not pose a barrier to the public release.
Future Development Outlook
The development team stated that larger models are currently in production to continue the scaling trajectory. The current release serves as the foundation for future personal superintelligence systems designed to handle long-horizon agentic tasks and complex coding workflows.





