Model Signal
Model releases, translated into operating decisions.
A recurring Signal + Noise series on frontier, open, multimodal, and specialist models — what changed, where each model fits, what breaks, and what teams should do next.
GLM-5.3
GLM-5.3 is a post-training update from Z.ai focused on long-horizon coding and advanced cybersecurity tasks.
GLM-5.3
GLM-5.3 is a post-training update from Z.ai focused on long-horizon coding and advanced cybersecurity tasks.
Qwen3.8 27B
Qwen3.8-27B is a dense 27-billion-parameter native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B is a sparse Mixture-of-Experts causal language model from Qwen with 2.4 trillion total parameters, 95 billion activated parameters per token, and open-weight availability.
Gemini 3.7 Flash
Gemini 3.7 Flash is a lightweight, high-speed model from Google DeepMind optimized for coding and agentic workflows at a lower price point.
Palmyra X6
Writer released Palmyra X6, an enterprise-focused general model optimized for agentic workflows and reduced operating costs.
DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is the official GA release of DeepSeek V4 Pro, superseding the preview version and available on DeepSeek’s app, web, and API.
BDH-CQ
Pathway's BDH-CQ is a 150-million-parameter reasoning model in a Post-Transformer architecture that scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.0007 per task.
Muse Glimmer 30B
Muse Glimmer is a ~29.6B-parameter dense multimodal causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks and local agent workflows on consumer hardware, supporting multimodal understanding, tool use, and failure recovery without requiring cloud infrastructure or network access.
Muse Glimmer
Muse Glimmer is a 30-billion-parameter open-weight model by Meta designed specifically for local agentic workflows.
Muse Spark 1.2
Muse Spark 1.2 is a natively multimodal reasoning model from Meta optimized for complex agentic tasks and coding workflows.
Muse Code
Meta has released a beta version of Muse Code, a terminal-based AI coding agent designed to autonomously navigate and modify large codebases.
LFM2.5-2.6B
A 2.6B dense model trained for agentic workloads, built for on-device deployment with native tool calling.
Qwen3.8-Max
Qwen3.8-Max is Alibaba's frontier multimodal foundation model, offering a 1-million-token context window for complex reasoning and large-scale visual understanding tasks.
Astra
OpenAI has internally previewed Astra, a next major model family oriented toward deep, long‑horizon reasoning, which it used to generate advances on ten long‑standing problems in mathematics and theoretical computer science.
H3
MiniMax H3 is a general-purpose multimodal video-generation model that jointly understands text, images, video, and audio and generates up to 15‑second 2K video clips with native stereo (dual‑channel) audio.
Gemini Spark
Gemini Spark is an agentic AI model integrated directly into Google Chrome to autonomously execute long-running background web tasks.
Gemini Robotics 2
A vision-language-action model designed to provide whole-body control and physical reasoning capabilities for humanoid and general-purpose robots.
Inkling Small
An open-weights, 276-billion-parameter sparse Mixture-of-Experts multimodal transformer (276B total / 12B active) designed to deliver much of the flagship Inkling model’s reasoning, coding, and multimodal performance at lower compute cost and latency.
Gemini Robotics ER 2
Gemini Robotics ER 2 is Google DeepMind’s most capable embodied reasoning vision-language model, acting as a high-level brain for robots to communicate with humans, understand the physical world, plan multi-step tasks, and orchestrate tools and action models, including multi-robot collaboration.
Qwen3.7 Flash
Qwen3.7 Flash is a high-efficiency vision-language model from Alibaba featuring a 1M-token context window, optimized for multimodal agents, visual coding, and spatial understanding.
Claude Opus 5
Claude Opus 5 is Anthropic's step-change improvement over Claude Opus 4.8, designed for deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.
Gemini 3.6 Flash
Gemini 3.6 Flash is a high-efficiency multimodal model from Google optimized for coding, agentic workflows, and large-scale knowledge tasks.
Gemini 3.6 Flash
Gemini 3.6 Flash is a high-efficiency, natively multimodal model optimized for fast execution, coding, and agentic workflows across large context lengths.
Laguna S 2.1
Poolside has released Laguna S 2.1, a 118-billion-parameter open-weight Mixture-of-Experts model optimized for agentic coding and software development tasks.
Gemini 3.5 Flash Cyber
Gemini 3.5 Flash Cyber is a lightweight, domain-specific model fine-tuned from the Flash architecture to identify, validate, and patch software vulnerabilities.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing.
Gemini 3.6 Flash
Gemini 3.6 Flash is a workhorse, frontier‑level multimodal model that delivers better coding, knowledge work, and token‑efficient performance than Gemini 3.5 Flash, optimized for real‑world agentic and everyday tasks at higher speed and lower cost.
Qwen3.8-Max
Alibaba's 2.4-trillion-parameter flagship model, released as Qwen3.8-Max-Preview.
Kimi K3
Moonshot AI has released Kimi K3, a 2.8-trillion parameter open-weight model that achieves frontier-level coding capabilities and outperformed Claude Fable 5 on the Frontend Code Arena.
GPT-5.6 Sol
GPT-5.6 Sol is OpenAI's flagship frontier model in the 5.6 series, optimized for complex reasoning, multi-step coding, and agentic workflows.
GPT-5.6 Luna
GPT-5.6 Luna is OpenAI’s fastest and most affordable GPT-5.6 tier, positioned for low-latency, cost-sensitive use cases.
ChatGPT Work
ChatGPT Work is an agent inside ChatGPT, powered by GPT‑5.6 and Codex, that can gather context across your apps, files, and workflows on web, mobile, and desktop to turn high‑level goals into finished work artifacts such as documents, spreadsheets, presentations, reports, and web applications, staying with complex multi‑step projects for hours.
Muse Spark 1.1
Meta's Muse Spark 1.1 is an API-accessible, agentic coding model featuring a 1-million-token context window and optimized for automated tool and computer use.
GPT-5.6 Sol
GPT-5.6 Sol is OpenAI's flagship frontier model, offering top-tier capabilities in complex coding, scientific reasoning, and cybersecurity.
GPT-5.6
OpenAI's GPT-5.6 is a major frontier model family featuring three tiers—Sol, Terra, and Luna—designed to deliver improved intelligence and cost efficiency across enterprise workloads.
Grok 4.5
Grok 4.5 is xAI's flagship multimodal model, featuring a 500,000-token context window and advanced capabilities in coding, STEM, and general knowledge work.
GPT-Live
OpenAI's GPT-Live is a family of full-duplex voice models designed for low-latency, real-time human-AI interaction.
Claude Sonnet 5
Anthropic's Claude Sonnet 5 is a Sonnet-tier model with a 1M-token context window, 128k max output tokens, adaptive thinking, and standard pricing of $3 per million input tokens and $15 per million output tokens; it launched with introductory pricing of $2/$10 through August 31, 2026 and became the default model for Free and Pro plans.
TabFM
Google's TabFM is a zero-shot foundation model capable of performing classification and regression on tabular data without requiring per-dataset training.
Nano Banana 2 Lite
A fast, cost-efficient Gemini Image model from Google (Gemini 3.1 Flash-Lite Image / Nano Banana 2 Lite) designed for high-throughput text-to-image generation and editing with latency of roughly four seconds per image.
Base 1
Base 1 is Base44’s first proprietary large language model, fine-tuned on an open‑source base and trained on tens of millions of real user interactions, purpose‑built to create and edit web applications inside the Base44 vibe‑coding platform.
GPT-5.6 Terra
GPT-5.6 Terra is the balanced, mid‑tier model in OpenAI’s GPT‑5.6 family, designed to provide strong reasoning and tool‑use capabilities with lower cost and high scalability, and it supports up to a 1M‑token context window for large‑scale applications.
GPT-5.6 Sol
GPT-5.6 Sol is OpenAI's flagship frontier model in the GPT-5.6 family, built for complex professional work such as deep reasoning, advanced coding, and long‑horizon agentic workflows, with multimodal (text + vision) support.
GPT-5.6 Terra
GPT-5.6 Terra is OpenAI's mid-tier multimodal model, positioned between the flagship Sol and efficient Luna tiers to provide a balance of reasoning performance and cost-efficiency.
LFM2.5-230M
Liquid AI has released LFM2.5-230M, a 230-million-parameter open-weight language model built to run anywhere (CPUs, GPUs, NPUs) for on-device and edge AI workloads with strong performance on tool use and structured data extraction.
OCR 4
A specialized document extraction model providing structured outputs with layout analysis and confidence scoring across 170 languages.
Sakana Fugu
Sakana Fugu is an orchestration model that functions as a multi-agent system by dynamically routing tasks across a swappable pool of underlying large language models.
VibeThinker-3B
VibeThinker-3B is a 3-billion-parameter open-weights reasoning model developed by WeiboAI.
GLM-5.2
GLM-5.2 is a 753-billion parameter open-weights large language model from Z.ai that is specifically optimized for long-horizon coding tasks.
Varya
An open-weight video generation model optimized for Indian cultural contexts and highly cost-efficient inference.
Claude Fable 5
Claude Fable 5 is Anthropic’s most capable generally available Mythos‑class model, designed for long‑horizon agentic knowledge work and coding, with multimodal inputs and a 1‑million‑token context window.
Claude Mythos 5
Claude Mythos 5 is the same model as Claude Fable 5 but with cyber safeguards lifted, and it initially became available to existing Claude Mythos Preview users such as cybersecurity partners in Project Glasswing. Anthropic suspended access to both models on June 12, 2026 after a U.S. government export-control directive, then restored access to Mythos 5 for a set of U.S. organizations after government approval.
Claude Fable 5
Anthropic’s most capable publicly available model and its first Mythos‑class model made safe for general use, released for global access across the Claude API, Claude Platform, Claude Code, Claude Cowork, and major cloud partners.
Nemotron-3-Ultra-550B-A55B-NVFP4
A frontier-scale 550-billion parameter hybrid model from NVIDIA featuring a 1-million token context window and NVFP4 precision.
Qwen3.7-Plus
Qwen3.7-Plus is a mid-tier, multimodal model from Alibaba featuring a one-million token context window, optimized for cost-effective reasoning and agentic workflows.
Claude Opus 4.8
Claude Opus 4.8 is Anthropic's generally available Opus model, announced and released on May 28, 2026.
Gemini 3.1 Flash Image
Gemini 3.1 Flash Image (Nano Banana 2) is a high-efficiency image generation and conversational editing model from Google, optimized for speed, low latency, and high-volume developer use cases.
Claude Opus 4.8
Builds on Opus 4.7 with stronger honesty and reliability, improved long-horizon agentic coding and professional work, dynamic workflows in Claude Code that orchestrate hundreds of parallel subagents, support for mid-conversation system messages, and a 2.5x fast mode at unchanged base pricing but much cheaper than prior fast-mode offerings.
Qwen3.7 Max
A frontier-discount agent substrate from Alibaba — strong agentic index, 1M context, and prompt-cache economics that change what becomes economical to automate.
Grok Build 0.1
Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows and interactive coding agents, supporting text and image inputs with text output and a 256K-token context window.
Command A+
Cohere’s open-weight sparse Mixture-of-Experts model built for enterprise agentic workloads, combining text and vision inputs, multilingual support, tool use, and complex reasoning within a 128K context for sovereign, privately deployable AI.
Gemini 3.5 Flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model optimized for fast, cost-effective coding and parallel agentic execution.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a specific revision of DeepSeek's efficient mixture-of-experts model, updated via re-post-training for enhanced coding, reasoning, and agent workflows.
DeepSeek V4 Flash
Efficiency-focused Mixture-of-Experts variant in the DeepSeek V4 series, optimized for fast, cost-efficient inference while preserving strong reasoning and coding performance with a 1M-token context window.
GPT-5.5
OpenAI's latest flagship with noticeably stronger reasoning and autonomy than GPT-5.4. Available in standard and Pro variants. 1M context window, $5/$30 per 1M tokens for standard.
Qwen3.6-27B
A 27-billion-parameter dense multimodal (image-text-to-text) model with native 262K token context, released as the first open-weight Qwen3.6 variant and designed for strong coding and agentic/repository-level reasoning.
Qwen3.6-35B-A3B
An open-weight sparse Mixture-of-Experts multimodal model from Qwen with 35 billion total parameters and about 3 billion active per token, combining a causal language model with a vision encoder and optimized for long-context, agentic coding and reasoning.
GLM-5.1
Open-weights model from Z.ai (Zhipu AI) with strong web-dev coding performance. Ranks above Kimi K2.6 on Code Arena WebDev leaderboard (1,534 Elo). Competitive on multilingual tasks.
gemma-4-26B-A4B-it
A 26-billion-parameter, instruction-tuned Mixture-of-Experts model from Google supporting multimodal inputs and a 256K-token context window.
gemma-4-31B-it
An instruction-tuned ~31-billion parameter dense open-weights model from Google's Gemma 4 (Gemma 4 31B IT) family, offering multimodal input support for text, image, and video (via frames) and a 256K-token context window.
Kimi K2.6
1 trillion-parameter vision-language model from Moonshot AI. Highest-ranked open-weights model on Artificial Analysis leaderboard (score 54). Designed for long-horizon agentic coding with plan-write-test-debug loops lasting days.
DeepSeek V4
DeepSeek's latest open-source flagship. Rivals closed frontier models on coding and reasoning benchmarks while maintaining extremely low inference cost. MoE architecture.
Mistral Small 4
Efficient small model with configurable reasoning — set reasoning_effort from none to high on the fly. 5x more params than Small 3 but only 6B active per token, 40% faster end-to-end.
GPT-5.4
OpenAI’s current frontier flagship model for professional work, unifying general reasoning, coding, native computer use, and long‑context workflows across ChatGPT, the API, and Codex.
Qwen3.5-9B
An open‑weights 9‑billion‑parameter **dense** multimodal model from Alibaba’s Qwen team, featuring **native multimodal vision support for images and video**, a **262,144‑token native context window extensible to ~1,010,000 tokens**, and a hybrid **Gated DeltaNet + Gated Attention** architecture optimized for efficient local and cloud deployment.
Mercury 2
Extremely fast diffusion-based reasoning language model from Inception Labs that refines tokens in parallel rather than generating them sequentially, achieving around 1,000 tokens per second and ranking near the top of Artificial Analysis speed leaderboards.
Nano Banana 2
A Gemini 3.1 Flash Image–based **image generation and editing model** from Google, positioned as its best fast image model, integrated across Gemini, Search, AI Studio and Cloud, with web-grounded generation, high-fidelity visuals, accurate multilingual text rendering, and advanced editing capabilities.
Gemini 3.1 Pro
Google’s most advanced Gemini reasoning model for complex tasks, significantly improving reasoning and multimodal performance over Gemini 3 Pro and handling text, audio, images, video, PDFs, and entire code repositories within a 1M‑token context window.
Grok 4.20
xAI frontier 4.x model with up to a 2M‑token context window, industry‑leading low hallucination rate, and an optional 4‑agent collaboration system, released as Grok 4.20 Beta in February 2026 with later GA and multiple API variants (reasoning, non‑reasoning, and multi‑agent).
Claude Sonnet 4.6
Most capable Sonnet yet — approaches Opus-level performance for coding, computer use, and document work at a mid-tier price point. 1M context window in beta.
Qwen3-Coder-Next
An 80B-parameter open-weight Mixture-of-Experts model by Qwen specifically optimized for agentic coding tasks.
gpt-oss-120b
OpenAI's gpt-oss-120b is an open-weight Mixture-of-Experts language model with approximately 116.8–117 billion total parameters, designed for powerful reasoning, tool use, and agentic workflows.
gpt-oss-20b
An Apache 2.0 licensed open-weight 20B (21B) Mixture-of-Experts language model from OpenAI, designed as a medium-sized reasoning model optimized for low-latency, local, and consumer-hardware deployment.
Qwen3-Coder-30B-A3B-Instruct
Qwen3-Coder-30B-A3B-Instruct is an open-weights Mixture-of-Experts (MoE) model optimized for software development tasks, designed for high inference efficiency and long-context processing.
Qwen3-Coder-480B
Alibaba’s most advanced open‑source agentic coding model — a 480B total / 35B active Mixture‑of‑Experts code model optimized for multi‑step software engineering, repository‑scale reasoning, tool use, and browser‑style interaction, with a native long context (≈256K–262K tokens) extendable to 1M via YaRN/extrapolation.
Grok 4
xAI’s flagship Grok 4 model, positioned as its most intelligent frontier‑level system, with native tool use and real‑time search integration and a Heavy variant that adds a multi‑agent layer and advanced reasoning capabilities.
Imagen 4
Google's latest text-to-image model with photorealistic output, precise text rendering, and advanced style control. Powers Gemini's image generation and is available via Vertex AI.
Gemini 2.5 Flash
Google’s most efficient **workhorse** model in the Gemini 2.5 family, designed for **speed**, **low latency**, and **cost efficiency**, and serving as Google’s first fully **hybrid reasoning** model with configurable “thinking” budgets, natively **multimodal** (text, images, audio, video) with a **1M‑token context window**, generally available via the Gemini API, Google AI Studio, Vertex AI, and accessible in the Gemini app.
Qwen3-0.6B
A 0.6-billion-parameter dense causal language model in the Qwen3 family, released by Qwen/Alibaba Cloud under Apache 2.0, with strong reasoning, instruction-following, agent, and multilingual capabilities.
Qwen3-235B-A22B
Alibaba's flagship open-source hybrid reasoning MoE model. 235B total / 22B active params. Seamlessly switches between thinking and non-thinking modes. Trained on 36T tokens, supports 119 languages.
Llama 4 Behemoth
Meta's teacher model (still training) — 288B active params, 16 experts, ~2T total params. Outperforms GPT-4.5 and Claude Sonnet 3.7 on STEM benchmarks. Used to distill Scout and Maverick.
Llama 4 Scout
Meta’s open-weight natively multimodal Mixture-of-Experts model with an industry-leading 10M token context window. 17B active parameters with 16 experts (109B total) and designed to fit on a single NVIDIA H100 GPU.
Llama 4 Maverick
Meta’s open‑weight natively multimodal Mixture‑of‑Experts model with ~400B total parameters and 17B active parameters using 128 experts, designed for high‑capacity text‑and‑image reasoning and cost‑efficient performance.
Agentic Pretrained Transformer (APT-1)
APT-1 is a specialist Agentic Pretrained Transformer system designed specifically for agentic applications and customer service workflows, optimized for deterministic execution of actions and strict policy adherence rather than general-purpose text generation.
Claude Code
An agentic coding tool developed by Anthropic that lives in your terminal, reads your codebase, edits files, runs commands, and integrates with your development tools across terminal, IDE, desktop app, and browser.
Sand
Cursor is developing Sand, a general-purpose AI agent designed to automate everyday knowledge work tasks such as email and document management for non-developers.
Base44 Model
Base44 has introduced its first proprietary large language model, **Base 1**, trained on tens of millions of real app-building interactions, to power its vibe-coding app-creation platform and reduce reliance on third-party frontier models.
Model Signal tracks AI model releases across every provider. Each model has a compounding Model Signal report — a synthesized brief that auto-updates as the Signal + Noise wire mentions it. It covers releases through an operator lens: who should care, where each model fits, what changes, and what does not.