Beyond One-Size-Fits-All: PIM, HBF and More for the New Spectrum of AI Serving
As LLM services diversify from ultra-low-latency interactions to cost-sensitive, long-context workloads, a homogeneous GPU-HBM architecture is no longer sufficient. This presentation introduces a workload-optimized approach spanning both hardware and software.
Processing-In-Memory (PIM) targets decode-heavy premium workloads by reducing data movement and providing high memory BW. High Bandwidth Flash (HBF) addresses large-batch, long-context deployments by expanding memory capacity with the density of 3D NAND. We also present SALT-KV (Semantic-Aware Lifecycle Tiering for KV Cache), a software solution for managing KV cache across GPU memory, CPU DRAM, and SSD. By considering cache semantics and reuse potential, SALT-KV improves cache placement, scheduling, and offloading efficiency.
The presentation and demonstration show how PIM, HBF, and SALT-KV can improve the performance, scalability, and cost efficiency of next-generation AI serving.
