0
Applied AI·September 15, 2026·1 min read

Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs

Share

If Vera Rubin NVL72 is delivering ~7x better tokens per MW on a 1.6T model than Blackwell, the constraint for frontier inference shifts even harder toward power and systems engineering. For anyone budgeting large-scale agentic workloads, model choice now has to be co-optimized with energy contracts and datacenter design, not just GPU count.

Applied AI

Thoughts on AI labs' safety concerns: a coordinated slowdown may look like an antitrust conspiracy to limit output that would preserve frontier model margins

If labs coordinate on “safety slowdowns,” regulators may read it as output restriction to protect frontier margins rather than altruism. For operators, that means AI supply, pricing, and release cadence are now entangled with antitrust risk—treat long-term model access as a regulatory, not just technical, dependency.