Shared Everything

The Hard Lessons Behind Computing’s Biggest Bets

Episode Summary

Some of the most valuable lessons in computing never make it into the success stories. In this episode, HPC and AI veteran Glenn Lockwood draws on firsthand experience across scientific computing, national labs and hyperscale AI (NERSC, SDSC and Microsoft on its Fairwater supercomputer) to revisit ambitious systems that broke new ground, while revealing the technologies, assumptions and architectural bets within them that didn’t work quite as expected. From attempts to make distributed computing complexity disappear, to a high-performance storage technology users largely chose not to use, to enormous AI infrastructure bets overtaken by changes in how frontier models are built, these are remarkably candid stories from inside consequential computing projects. The failures expose lessons about complexity, adoption, general-purpose versus specialized technology, and what happens when workloads evolve faster than infrastructure can.