
The next age of LLMs? Dev gets a small LLM running at 10 tokens a second locally on a $10 microcontroller
THE SO WHAT
A 28.9M-parameter model running near 10 tok/s on a sub-$10 microcontroller — with most weights staying in flash — pushes “good enough” language intelligence into the bill-of-materials noise. Hardware teams building devices, sensors, and appliances should be scoping where on-device LLMs can replace cloud calls and unlock offline, low-latency features.
READ THE SOURCE
MORE FROM THE WIRE
Applied AIWhatsApp scam costs Hong Kong man $1.27 million after criminals used AI voice notes to impersonate his father — experts say secret codewords are the best way to stay safe
Voice is no longer an authentication factor when $1.27 million can move on the back of a cloned WhatsApp note. If you run any high-value workflows over consumer messaging — family offices, VIP client services, internal approvals — you need out-of-band codewords and hard limits this week, not just awareness training.
Meta CTO says employees should use AI productivity gains to do more work — not take more time off
Meta is making the implicit explicit — AI gains are being reinvested into output, not work-time reduction. Expect similar cultural pressure elsewhere: if you’re rolling out AI tools, be deliberate about whether the value shows up as headcount avoidance, faster roadmaps, or burnout risk from quietly raised expectations.
Applied AIPalantir’s stock stages best week since 2024, showing it’s no longer an AI loser
Public markets are starting to reward actual AI revenue, not just AI narrative — Palantir’s rebound is tied directly to “booming demand” for deployed AI solutions. If you’re selling into similar buyers, expect procurement to benchmark you against vendors that can show production usage and spend, not pilots.
Applied AIOpenAI pumps the brakes on new Astra model over cybersecurity concerns
Pausing internal work on Astra because internal evals flagged “powerful cybersecurity abilities” is a line in the sand — offensive capability is now a gating factor for model release. If you’re building on frontier models, budget time for capability downgrades, stricter access tiers, and slower rollouts as security standards harden.