0
MODEL SIGNAL · META

Llama 4 Behemoth

Meta's teacher model (still training) — 288B active params, 16 experts, ~2T total params. Outperforms GPT-4.5 and Claude Sonnet 3.7 on STEM benchmarks. Used to distill Scout and Maverick.

CATEGORYReasoning
CONTEXT1M
RELEASEDApril 5, 2025
Key Features
  • 288B active / ~2T total params
  • Best STEM benchmarks in class
  • Distillation teacher for Llama 4
  • MoE 16 experts
  • Not yet publicly released

Provider announcement →

Read the Model Signal report →

MODEL SIGNAL

Llama 4 Behemoth

Meta's ~2T parameter teacher model powering the distillation of the Llama 4 lineup.

Bottom line

Meta has confirmed the development of Llama 4 Behemoth, a massive 16-expert Mixture-of-Experts (MoE) model with approximately 2 trillion total and 288 billion active parameters. Currently still in training and not publicly released, Behemoth serves as the distillation teacher for Llama 4 Scout and Maverick, and reportedly outperforms frontier models like GPT-4.5 and Claude 3.7 Sonnet on STEM benchmarks.

Signal

Based on Meta's primary disclosures, Behemoth represents a significant scale-up in the Llama architecture. The model employs a 16-expert MoE design, activating 288 billion parameters per forward pass out of roughly 2 trillion total parameters. Its primary verified role is acting as an upstream teacher model to distill capabilities down to the smaller Llama 4 Scout and Maverick variants.

Meta states the model is currently yielding best-in-class results on STEM benchmarks, specifically outperforming competing frontier models like OpenAI's GPT-4.5 and Anthropic's Claude 3.7 Sonnet. Behemoth is explicitly noted as still being in the training phase.

Noise: What is not settled

Several circulating specifications have not been verified by primary Meta sources. Speculation regarding a 1 million token context window and a firm public release date of April 5, 2025, remain completely unconfirmed. Furthermore, because Behemoth is heavily emphasized as a "teacher model," its ultimate release posture—specifically whether the ~2T parameter weights will be made available for public deployment or kept strictly internal to Meta—remains unresolved.

Where it fits

The operator read on Behemoth is that it currently serves as a leading indicator of Llama 4's capability ceiling rather than an immediate infrastructure target. If the model remains an internal teacher, operators will not deploy it directly; instead, they will benefit from its distilled reasoning capabilities through the much lighter Scout and Maverick models.

The emerging pattern suggests that even if Behemoth's weights are eventually released, a model of this scale will require massive, specialized compute clusters to serve. This would restrict direct self-hosting to hyperscalers and top-tier enterprise AI environments with highly provisioned, multi-node GPU infrastructure.

Model Signal · Signal + Noise · Isaiah Steinfeld