Arthur Rasmusson
Arthur Rasmusson is Director of AI Architecture at LightBits Labs, where he works on Inferra, the next generation of KV cache technology designed to eliminate stalls in memory-bound GPU inference. He is a kernel programmer, virtualization specialist, and machine learning engineer focused on GPU drivers and AI inference, with experience spanning open- and closed-source development in roles including Chief Technology Officer, Founding Engineer, Principal AI Engineer, and Machine Learning Engineer.
As Founding Engineer of Arc Compute, Arthur authored open-source documentation and GPU virtualization software for the Open-IOV project. At Cohere, his work on Just-In-Time Inference and Just-In-Time Training reduced auto-scaling latencies by orders of magnitude, introducing the GPUDirect Storage API to enable real-time scaling of hardware allocations in hybrid training and inference environments, meeting customer demand and maximizing utilization. He later introduced Paged Attention over RDMA (PAoR) to the open-source AI community, using distributed filesystems in GPU clusters to eliminate wasted compute from redundant cache regeneration in open-source inference servers, delivering orders-of-magnitude performance gains at scale.
At Weka, as Principal AI Engineer, Arthur authored the open-source software behind AI “token warehouses,” contributing the “KV Cache GPUDirect Storage” feature upstream to the TensorRT-LLM project with supporting code later moved to NVIDIA NIXL as part of the Dynamo stack and implementing Python-native support for GPUDirect Storage APIs used in LMCache for the vLLM ecosystem. His work bridges low-level systems concepts with high-performance data paths for AI, a thread that runs directly into his current role architecting Inferra at LightBits Labs.
Sessions
-
NVMe is All You Need15-Sep-2026Expo Theater 1
