Survey of HPC in US Research Institutions
hpcsupercomputingexascaleuniversity-computinggpu-computingai-infrastructure
Abstraction: Comparative survey of HPC capabilities across US universities, national labs, and industry
Key points:
- University HPC CAGR ~18% vs. national labs ~43% and industrial hyperscalers ~78%, widening the capability gap
- DOE leadership machines: Frontier (ORNL) 1.35 exaflops (AMD EPYC + MI250X), Aurora (Argonne) 1.01 exaflops (Intel Xeon Max + Ponte Vecchio), El Capitan (LLNL) 1.74 exaflops world #1 (AMD MI300A APUs)
- University flagships: Frontera (TACC) 23.5 PFlop/s, HiPerGator-AI (U Florida) 17.2 PFlop/s; NSF Horizon at TACC projected ~400 PFlop/s for 2025-2026
- Google TPU v4 Pod: ~exaflop-class AI throughput, optical-switch fabric, 1.6-1.7× faster than A100 at 1.3-1.9× less power
- AWS Trainium UltraCluster: 40,000 chips on petabit-scale non-blocking fabric; 50% lower training cost than GPU instances
- Emerging strategies for academic democratization: federated computing, idle-GPU harvesting, cost-sharing, and decentralized reinforcement learning
Connections: Ornl · Tacc · Google · Aws · High Performance Computing · Exascale Computing · GPU Computing