In this episode, we're joined by Stephen Bates, AMD Fellow in the AI Group, and Alon Horev, Co-founder and CTO of VAST Data, to explore one of the biggest architectural shifts happening in AI today: the move from training to inference. Using AMD and VAST's expanded partnership as a starting point, the conversation examines why KV cache has become a critical part of AI infrastructure, how long-context and agentic AI are changing the way systems are designed, and why storage, networking, and open software are becoming just as important as the GPUs themselves in delivering fast, scalable inference.