MODEL SIGNAL
Llama 4 Behemoth
Meta's ~2T parameter teacher model powering the distillation of the Llama 4 lineup.
Bottom line
Meta has confirmed the development of Llama 4 Behemoth, a massive 16-expert Mixture-of-Experts (MoE) model with approximately 2 trillion total and 288 billion active parameters. Currently still in training and not publicly released, Behemoth serves as the distillation teacher for Llama 4 Scout and Maverick, and reportedly outperforms frontier models like GPT-4.5 and Claude 3.7 Sonnet on STEM benchmarks.
Signal
Based on Meta's primary disclosures, Behemoth represents a significant scale-up in the Llama architecture. The model employs a 16-expert MoE design, activating 288 billion parameters per forward pass out of roughly 2 trillion total parameters. Its primary verified role is acting as an upstream teacher model to distill capabilities down to the smaller Llama 4 Scout and Maverick variants.
Meta states the model is currently yielding best-in-class results on STEM benchmarks, specifically outperforming competing frontier models like OpenAI's GPT-4.5 and Anthropic's Claude 3.7 Sonnet. Behemoth is explicitly noted as still being in the training phase.
Noise: What is not settled
Several circulating specifications have not been verified by primary Meta sources. Speculation regarding a 1 million token context window and a firm public release date of April 5, 2025, remain completely unconfirmed. Furthermore, because Behemoth is heavily emphasized as a "teacher model," its ultimate release posture—specifically whether the ~2T parameter weights will be made available for public deployment or kept strictly internal to Meta—remains unresolved.
Where it fits
The operator read on Behemoth is that it currently serves as a leading indicator of Llama 4's capability ceiling rather than an immediate infrastructure target. If the model remains an internal teacher, operators will not deploy it directly; instead, they will benefit from its distilled reasoning capabilities through the much lighter Scout and Maverick models.
The emerging pattern suggests that even if Behemoth's weights are eventually released, a model of this scale will require massive, specialized compute clusters to serve. This would restrict direct self-hosting to hyperscalers and top-tier enterprise AI environments with highly provisioned, multi-node GPU infrastructure.