Shared Everything

CoreWeave on Building Resilient AI Infrastructure at Massive Scale

Episode Summary

At CoreWeave, building AI infrastructure at massive scale means rethinking the datacenter as a system. Jacob Yundt, Senior Director of Compute Architecture at CoreWeave, joins us to explore how GPU infrastructure is evolving as racks become tightly integrated computers and power, cooling, networking, storage, firmware and software orchestration become inseparable parts of the architecture. We dig into the demands of inference, heterogeneous GPU fleets, workload placement, observability and resiliency, and why the next generation of AI infrastructure may require much more than simply scaling up the designs we use today.