
OpenAI releases MentalHealthBench, an open benchmark to evaluate AI responses in realistic mental health conversations, developed with 80+ licensed experts
THE SO WHAT
Mental health is moving from "do not touch" to "govern with benchmarks" for frontier models. If you're deploying AI in any regulated or high-risk domain, expect pressure to anchor claims to open, expert-built eval suites like MentalHealthBench rather than internal red-team anecdotes.
READ THE SOURCE
MORE FROM THE WIRE
Applied AIRingg’s AI agents resolve up to 65% of customer calls with OpenAI
If GPT-5.6 can handle 65% of calls across channels at 90% lower unit cost than GPT-4.1, the economics of front-line support are tilting from augmentation to replacement. CX leaders should be modeling what happens when the marginal cost of an incremental support interaction rounds toward zero and quality becomes the only real constraint.
Applied AIAnthropic: We're not trying to eat startups' lunch
When a model provider says it’s not going up-stack, that’s a positioning move—assume the boundary will keep shifting as they learn where usage concentrates. Founders should design so they can swap underlying models and keep their own data, workflow, and distribution as the defensible layer.
Applied AIAI safety startups mapped: 37 companies building Europe's trust layer
Dozens of AI safety and governance startups in Europe means compliance and assurance are fragmenting into a new vendor layer, not just a feature of the big clouds. If you’re deploying AI at scale, expect procurement pressure—from boards and regulators—to show you have third-party checks, not just internal policies.
Applied AIQualcomm announces the Snapdragon Sound Elite Gen 2 for audio wearables, with up to 2x greater on-device AI performance and up to 40% lower power consumption
2x on-device AI with 40% lower power in audio chips means “smart” earbuds and wearables can run real models locally instead of phoning home for everything. If your product relies on cloud inference for voice or sensing, assume users will expect offline, low-latency behavior within a hardware cycle.