0
Applied AI·September 6, 2026·1 min read

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch

Share

Model evals are now a moving target, even post-launch—which means you cannot outsource risk or performance assumptions to vendor benchmarks. Treat vendor scores as marketing inputs and build your own task-level evals for the workflows that matter to you.