The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.
Follow Latent Space to make it a durable For You signal.
Latent Space’s interview with Baseten’s Philip Kiely and Ali Taha examines inference engineering as a distinct discipline focused on making trained models fast, reliable, and affordable in production. The discussion covers cache-aware routing, separating prefill from decode, speculative decoding, quantization, model parallelism, GPU kernels, and hardware-aware serving. The article says Baseten has raised a $13 billion Series F and portrays the company as a major beneficiary of growing demand for inference infrastructure, though the supplied reporting does not provide further details about the financing. The engineers describe reported cases in which quantization preserved benchmark quality while increasing throughput by 20%, and say broader serving optimizations can produce gains of 20%, 100%, or 200% and make models up to 10 times faster. They also distinguish data-center inference—focused on reducing latency and cost—from local inference, where fitting models onto limited hardware and preserving quality are the primary constraints. Beyond language models, the conversation addresses diffusion and autoregressive video generation, the memory and attention limits facing long-form AI,