0
Applied AI·August 26, 2026·1 min read

Qwen4’s architecture is here early, firing 6B parameters out of 125B

Share

Qwen3.8-Flash-Next previewing a 125B-parameter model that only activates 6B per token shows sparse, mixture-of-experts-style scaling is becoming the default path to efficiency. For teams betting on open weights, pay attention not just to benchmarks but to licensing — EU AI Act open-source exemptions may not apply here, which affects compliance calculus.

Applied AI

Gemini Live gains agentic tasks, voice inbox control and personal intelligence

Gemini Live turning spoken commands into background agents across Docs, Sheets, Drive, and Gmail means the productivity assistant is becoming an OS-level workflow router. The DMA requirement to let rival assistants reach comparable functionality on Android cracks the door for third-party agents—if you build SaaS, assume voice-first, cross-app automation will be a user expectation, not a novelty.