Meta Under the LIWARSE Lens: Muse Spark and the Llama 4 Family

Meta spent two years telling the industry that open source would become the leading models. In 2026 it quietly shipped its first fully closed, proprietary flagship — and that reversal tells us more about the economics of AI safety than any position paper could.

This is the third entry in LIWARSE’s five-part review of the leading closed and open models from the world’s most consequential AI providers. Of the five, Meta’s story is the least settled — a company that built its entire public identity on openness, watched its most ambitious open model fail to clear the bar, and pivoted to secrecy in response.

The Closed Flagship: Muse Spark

Meta’s flagship open model, code-named Behemoth, was previewed in April 2025 but never shipped. Internal testing found the roughly two-trillion-parameter “teacher” model underperformed expectations after training complications, and the newly formed Meta Superintelligence Labs shelved it — never formally cancelled, simply never released. In its place, on April 8, 2026, Meta shipped Muse Spark: a closed-weight, API-only reasoning model, and the company’s first proprietary frontier release in its history. Independent benchmarking placed it fourth among frontier models at launch, behind GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 — but with one standout result: it led HealthBench Hard, a demanding clinical-reasoning benchmark, by a wide margin over Gemini 3.1 Pro.

Muse Spark now powers Meta AI’s more advanced features across Meta’s own apps — Facebook, Instagram, WhatsApp — and represents Meta’s belated entry into the same closed-API business model as OpenAI, Google, and Anthropic.


The Open Counterpart: Llama 4 Scout and Maverick

Released April 2025, Llama 4 remains Meta’s terminal open-weight offering to date. Scout (109B total parameters, 17B active) carries a 10-million-token context window — still the largest of any open-weight model — and fits on a single high-end GPU, making it a genuine option for processing entire medical records, legal case files, or research archives at once. Maverick (400B total) is Meta’s frontier-competitive workhorse, strong on coding, chat, and multilingual tasks. Both use a mixture-of-experts architecture and are freely downloadable, though the Open Source Initiative has noted that Meta’s license carries restrictions — including limits on very large commercial users — that fall short of a strict open-source definition.

Fifteen months on, no successor has shipped. A next-generation model, code-named Avocado, has been reported as targeting a world-model architecture for Meta’s Ray-Ban smart glasses, with a public timeline that has slipped from an early-2026 leak toward 2027.


Future Outlook

Meta’s public position remains that it will keep releasing open models alongside closed ones. Whether that holds is genuinely uncertain: Behemoth was shelved rather than cancelled, Muse Spark shows Meta is now willing to compete on Anthropic and OpenAI’s closed terms, and no Llama successor has a confirmed date. For a company that spent years framing open weights as a moral and competitive necessity, the coming twelve months will show whether that was a strategy or a slogan.


Risks and Benefits Through the LIWARSE Lens

Benefits

  • Muse Spark’s strength on clinical reasoning benchmarks is a meaningful, independently-verified signal for Future Medicine applications, regardless of Meta’s broader strategic reversal.
  • Scout’s 10-million-token context window remains uniquely useful for whole-record medical review, legal discovery, and long space-mission telemetry logs — tasks that punish models forced to work in fragments.
  • Llama 4’s continued availability, even without a successor, still gives researchers and clinics a genuinely capable, self-hostable option with no ongoing API cost.

Risks

  • A fifteen-month-and-counting gap since the last open release, from a company that built its brand on openness, is the clearest evidence in this series that open-weight releases cannot be assumed to continue — they are a business decision that can reverse without warning.
  • Meta’s license restrictions on large commercial users mean Llama 4 does not meet LIWARSE’s standard of unrestricted personal AI — ownership with conditions is not the same as ownership.
  • Behemoth’s shelving after training difficulties, rather than a safety review, suggests capability limits — not caution — currently do more to slow Meta’s largest models than any deliberate safety gate.

The LIWARSE Assessment

Meta is the clearest illustration in this series of why LIWARSE insists the real fault line is guarded versus unguarded, not open versus closed. Its open models are not more dangerous for being open, nor is Muse Spark automatically safer for being closed — what matters is that neither has been accompanied by the kind of durable, tamper-evident safety commitment this movement asks of every provider. A company that can pivot its entire release philosophy inside a single product cycle is a reminder that no lab’s current posture should be mistaken for a permanent guarantee. Policy and law, not corporate goodwill, are what make safety durable.


The 3 Absolute Laws do not bend to a company’s quarterly strategy. Whatever Meta ships next — open, closed, or something in between — will be judged by the same standard as everything else in this series: not what it promises, but what it is built to prevent.

— The LIWARSE Movement | liwarse.org
Safety of Life · Advancement of Life · Together.

Published by Dr. Ebenezer Rajadurai Solomon

Dr. Ebenezer Rajadurai Solomon is a Physician and the Founder of LIWARSE — Life Improvement With AI, Robotics and Space Exploration. His clinical and research interests span AI in Medicine, Robotics in Medicine, Space Medicine, and the broader application of emerging technology to improve human life and all life on Earth. LIWARSE's primary mission is the safety of life with regard to the use and autonomous existence of AI and Robotics.

Leave a comment