Loading
Arthur Rasmusson

Arthur Rasmusson

Director of AI Architecture, Lightbits Labs

Arthur Rasmusson is Director of AI Architecture at LightBits Labs, where he works on Inferra, the next generation of KV cache technology designed to eliminate stalls in memory-bound GPU inference. He is a kernel programmer, virtualization specialist, and machine learning engineer focused on GPU drivers and AI inference, with experience spanning open- and closed-source development in roles including Chief Technology Officer, Founding Engineer, Principal AI Engineer, and Machine Learning Engineer.

As Founding Engineer of Arc Compute, Arthur authored open-source documentation and GPU virtualization software for the Open-IOV project. At Cohere, his work on Just-In-Time Inference and Just-In-Time Training reduced auto-scaling latencies by orders of magnitude, introducing the GPUDirect Storage API to enable real-time scaling of hardware allocations in hybrid training and inference environments, meeting customer demand and maximizing utilization. He later introduced Paged Attention over RDMA (PAoR) to the open-source AI community, using distributed filesystems in GPU clusters to eliminate wasted compute from redundant cache regeneration in open-source inference servers, delivering orders-of-magnitude performance gains at scale.

At Weka, as Principal AI Engineer, Arthur authored the open-source software behind AI “token warehouses,” contributing the “KV Cache GPUDirect Storage” feature upstream to the TensorRT-LLM project with supporting code later moved to NVIDIA NIXL as part of the Dynamo stack and implementing Python-native support for GPUDirect Storage APIs used in LMCache for the vLLM ecosystem. His work bridges low-level systems concepts with high-performance data paths for AI, a thread that runs directly into his current role architecting Inferra at LightBits Labs. 

Sessions