Can Independent Testing Make AI Safer?
THE SO WHAT
Independent eval shops like Vals AI are emerging because model capabilities are outpacing in-house testing capacity—third-party red-teaming is becoming part of the deployment bill of materials. If you're shipping on top of frontier models, budget time and money for external evals the same way you do for pen tests and audits.
READ THE SOURCE
MORE FROM THE WIRE
Applied AIAI infrastructure company Cornelis raises $205M to chip away at Nvidia’s dominance
Cornelis is going straight at the interconnect bottleneck—$205M on Active Compute Fabric is a bet that wasted GPU cycles, not raw FLOPs, are the next margin lever. If you’re planning multi-node training, you now have more negotiating power on the network layer, not just on accelerators.
Applied AIGPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
The question is no longer “can AI review code” but “what quality premium are you actually buying per PR.” Teams should be benchmarking cheaper models like GPT-5.6 Luna against top-tier options on their own repos—if a $1.20 run catches 90% of issues, your unit economics on AI-assisted dev change fast.
Applied AIAre We Losing Control of AI? What’s Driving New Fears
“Loss of control” has moved from sci-fi to mainstream political and media framing, which will harden into expectations for technical and institutional brakes. Builders should assume higher scrutiny on autonomy, kill switches, and escalation paths—even for consumer-facing assistants.
Applied AI'Encouraging' to Hear AI Firms Taking Risk Seriously, Says Moynihan
When a major bank CEO publicly blesses AI firms’ safety posture, it’s a signal that large financial institutions are preparing to lean in—under the cover of shared responsibility. If you sell into FS, be ready with concrete answers on model risk, audit trails, and capital impact, not just productivity stories.