Why Sierra the Supercomputer Had to Die
supercomputerhpclawrence-livermorenuclear-securityhardware-lifecycle
Abstraction: Decommissioning lifecycle of Sierra supercomputer at Lawrence Livermore National Lab
Key points:
- Sierra (IBM Power9 + Nvidia Volta V100, 240 racks, ~7,000 sq ft) was once the world's 2nd fastest supercomputer at 94.64 petaflops; decommissioned after 7 years — a typical lifespan for HPC systems
- Replacement El Capitan (AMD Instinct MI300A APU, unified CPU/GPU memory) reached 1.809 exaflops in 2025 — ~19x faster than Sierra, consuming up to 36 MW vs Sierra's 11 MW
- Hardware obsolescence drove retirement: IBM Power9 CPUs and Nvidia Volta V100 GPUs no longer in production; IBM no longer supports the Red Hat Enterprise Linux version Sierra ran
- "Bathtub curve" of hardware failure: early defects, a golden era of stability, then rising failure rate as chips age — Sierra was approaching the high-failure tail
- Decommissioning involved classified data destruction: flash memory ground to fine powder, magnetic drives degaussed, all compute nodes shredded offsite — no simple donation possible
Connections: Lawrence Livermore National Laboratory · Ibm · Nvidia · High Performance Computing · Hardware Lifecycle
Source: https://www.wired.com/story/why-sierra-the-supercomputer-had-to-die/