Model Signal
Model releases, translated into operating decisions.
A recurring Signal + Noise series on frontier, open, multimodal, and specialist models — what changed, where each model fits, what breaks, and what teams should do next.
Qwen3.8-Max
Alibaba's 2.4-trillion-parameter flagship large language model, released in preview to compete directly with leading frontier systems.
Qwen3.8-Max
Alibaba's 2.4-trillion-parameter flagship large language model, released in preview to compete directly with leading frontier systems.
Kimi K3
Moonshot AI has released Kimi K3, a 2.8-trillion parameter open-weight model that achieves frontier-level coding capabilities and outperformed Claude Fable 5 on the Frontend Code Arena.
ChatGPT Work
ChatGPT Work is an OS-level AI agent by OpenAI that integrates across Mac and Windows applications to automate research and execute multi-step document creation tasks for business workflows.
Muse Spark 1.1
Meta's Muse Spark 1.1 is an API-accessible, agentic coding model featuring a 1-million-token context window and optimized for automated tool and computer use.
GPT-5.6 Sol
GPT-5.6 Sol is OpenAI's flagship frontier model, offering top-tier capabilities in complex coding, scientific reasoning, and cybersecurity.
GPT-5.6
OpenAI's GPT-5.6 is a major frontier model family featuring three tiers—Sol, Terra, and Luna—designed to deliver improved intelligence and cost efficiency across enterprise workloads.
GPT-Live
OpenAI's GPT-Live is a family of full-duplex voice models designed for low-latency, real-time human-AI interaction.
Base 1
Base 1 is a proprietary code generation model trained by Base44 to natively power its web application development and 'vibe coding' workflows.
TabFM
Google's TabFM is a zero-shot foundation model capable of performing classification and regression on tabular data without requiring per-dataset training.
Nano Banana 2 Lite
A fast, cost-efficient text-to-image model from Google designed to generate outputs in approximately four seconds.
Claude Sonnet 5
Anthropic's Claude Sonnet 5 introduces advanced agentic capabilities and a one-million token context window at a lower cost profile than its Opus counterparts.
LFM2.5-230M
Liquid AI has released a 230-million parameter open-weights model optimized for local data extraction and on-device agentic workflows.
Sakana Fugu
Sakana Fugu is an orchestration model that functions as a multi-agent system by dynamically routing tasks across a swappable pool of underlying large language models.
VibeThinker-3B
VibeThinker-3B is a 3-billion-parameter open-weights reasoning model developed by WeiboAI.
GLM-5.2
GLM-5.2 is a 753-billion parameter open-weights large language model from Z.ai that is specifically optimized for long-horizon coding tasks.
Varya
An open-weight video generation model optimized for Indian cultural contexts and highly cost-efficient inference.
Claude Mythos 5
Anthropic's frontier Claude Mythos 5 model was launched and subsequently disabled worldwide following a US government security directive.
Claude Fable 5
Anthropic's most powerful publicly released model — its first Mythos-class model available to enterprise and paid users. Leads on SWE-bench, knowledge work, and scientific research, scoring 10%+ above Opus 4.8 on key benchmarks.
Nemotron-3-Ultra-550B-A55B-NVFP4
A frontier-scale 550-billion parameter hybrid model from NVIDIA featuring a 1-million token context window and NVFP4 precision.
Claude Opus 4.8
Builds on Opus 4.7 with stronger agentic reasoning, adaptive thinking, dynamic parallel workflows in Claude Code, mid-conversation system messages, and 2.5x fast mode at 3x lower cost.
Qwen3.7 Max
A frontier-discount agent substrate from Alibaba — strong agentic index, 1M context, and prompt-cache economics that change what becomes economical to automate.
Command A+
Cohere’s open-weight sparse Mixture-of-Experts model built for enterprise agentic workloads, combining text and vision inputs, multilingual support, tool use, and complex reasoning within a 128K context for sovereign, privately deployable AI.
DeepSeek V4 Flash
Reasoning-optimized flash variant of DeepSeek V4 — neck-and-neck with Kimi K2.6 on coding benchmarks at significantly lower latency.
GPT-5.5
OpenAI's latest flagship with noticeably stronger reasoning and autonomy than GPT-5.4. Available in standard and Pro variants. 1M context window, $5/$30 per 1M tokens for standard.
Qwen3.6-27B
A 27.8-billion parameter dense multimodal model optimized for local deployment with strong coding capabilities and a 256K context window.
Qwen3.6-35B-A3B
A mid-sized, open-weight multimodal mixture-of-experts model from Qwen featuring 35 billion total and 3 billion active parameters.
GLM-5.1
Open-weights model from Z.ai (Zhipu AI) with strong web-dev coding performance. Ranks above Kimi K2.6 on Code Arena WebDev leaderboard (1,534 Elo). Competitive on multilingual tasks.
gemma-4-26B-A4B-it
A 26-billion-parameter, instruction-tuned Mixture-of-Experts model from Google supporting multimodal inputs and a 256K-token context window.
gemma-4-31B-it
An instruction-tuned 31-billion parameter open-weights model from Google's Gemma 4 family, featuring native multimodal support and a 256K context window.
Kimi K2.6
1 trillion-parameter vision-language model from Moonshot AI. Highest-ranked open-weights model on Artificial Analysis leaderboard (score 54). Designed for long-horizon agentic coding with plan-write-test-debug loops lasting days.
DeepSeek V4
DeepSeek's latest open-source flagship. Rivals closed frontier models on coding and reasoning benchmarks while maintaining extremely low inference cost. MoE architecture.
Mistral Small 4
Efficient small model with configurable reasoning — set reasoning_effort from none to high on the fly. 5x more params than Small 3 but only 6B active per token, 40% faster end-to-end.
Qwen3.5-9B
An open-weights 9-billion parameter multimodal model from Alibaba featuring native vision support and an extended 262K context window designed for efficient local or edge deployment.
GPT-5.4
OpenAI’s current frontier flagship model for professional work, unifying general reasoning, coding, native computer use, and long‑context workflows across ChatGPT, the API, and Codex.
Mercury 2
Fastest model on the Artificial Analysis leaderboard by output speed. Diffusion-based LLM architecture enabling massively parallel token generation — entirely different inference paradigm from autoregressive models.
Gemini 3.1 Pro
Google's most advanced reasoning model — doubles ARC-AGI-2 performance vs Gemini 3 Pro. Handles text, audio, images, video, and entire code repositories in a 1M context window.
Grok 4.20
xAI frontier 4.x model with up to a 2M‑token context window, industry‑leading low hallucination rate, and an optional 4‑agent collaboration system, released as Grok 4.20 Beta in February 2026 with later GA and multiple API variants (reasoning, non‑reasoning, and multi‑agent).
Claude Sonnet 4.6
Most capable Sonnet yet — approaches Opus-level performance for coding, computer use, and document work at a mid-tier price point. 1M context window in beta.
Qwen3-Coder-Next
An 80B-parameter open-weight Mixture-of-Experts model by Qwen specifically optimized for agentic coding tasks.
gpt-oss-120b
OpenAI's gpt-oss-120b is a 117-billion-parameter, open-weight Mixture-of-Experts model optimized for complex reasoning and agentic workflows.
gpt-oss-20b
An Apache 2.0 licensed open-weight 21B-parameter mixture-of-experts language model from OpenAI optimized for edge and on-device deployments.
Qwen3-Coder-30B-A3B-Instruct
Qwen3-Coder-30B-A3B-Instruct is an open-weights Mixture-of-Experts (MoE) model optimized for software development tasks, designed for high inference efficiency and long-context processing.
Qwen3-Coder-480B
Alibaba's most advanced agentic coding model — 480B total / 35B active MoE. Excels at full software dev pipelines, codebase debugging, and browser interaction. Context extendable to 1M.
Grok 4
xAI's flagship model with 100x training improvement over Grok 3. Leads ARC-AGI benchmarks, includes native tool use, real-time X search, and multi-agent coordination.
Qwen3-0.6B
A highly compact, 0.6-billion parameter language model from the Qwen3 family optimized for on-device reasoning and efficient fine-tuning.
Qwen3-235B-A22B
Alibaba's flagship open-source hybrid reasoning MoE model. 235B total / 22B active params. Seamlessly switches between thinking and non-thinking modes. Trained on 36T tokens, supports 119 languages.
Llama 4 Behemoth
Meta's teacher model (still training) — 288B active params, 16 experts, ~2T total params. Outperforms GPT-4.5 and Claude Sonnet 3.7 on STEM benchmarks. Used to distill Scout and Maverick.
Llama 4 Scout
Meta's efficient open-weight multimodal model with an industry-leading 10M token context. 17B active params with 16 experts — fits on a single H100 GPU.
Llama 4 Maverick
Meta's open-weight MoE model with 17B active / 128 experts. Best multimodal in its class — beats GPT-4o and Gemini 2.0 Flash. ELO of 1417 on LMArena.
Claude Code
An agentic command-line coding tool developed by Anthropic that integrates directly into local developer environments.
Agentic Pretrained Transformer (APT-1)
APT-1 is a specialized AI model designed specifically for executing reliable agentic workflows and automated actions.
Sand
Cursor is developing Sand, a general-purpose AI agent designed to automate everyday knowledge work tasks such as email and document management for non-developers.
Base44 Model
Base44 has introduced a proprietary code-generation model to natively power its app-creation platform, reducing its reliance on third-party frontier models.
OCR 4
A specialized document extraction model providing structured outputs with layout analysis and confidence scoring across 170 languages.
Model Signal tracks AI model releases across every provider. Each model has a compounding Model Signal report — a synthesized brief that auto-updates as the Signal + Noise wire mentions it. It covers releases through an operator lens: who should care, where each model fits, what changes, and what does not.