0
Applied AI·August 28, 2026·1 min read

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Share

Agent benchmarks are moving from toy tasks to full-stack scientific workflows — literature review, experiment design, analysis. If you’re building agents for real work, expect procurement and regulators to start asking for performance on domain-specific suites like this, not generic leaderboards.