01 Aug 2026

Outperforming vLLM and Llama.cpp on Gemma4-12B

CloudRift
Dmitry Trifonov, Slawomir Strumecki, Ivan Oleynikov
Emmy, an AI-driven compiler, generates CUDA kernels automatically instead of by hand. On Gemma 4 12B in FP16 its kernels reach up to 1.6x over cuBLAS on the RTX 5090, achieved via TMA transport, hybrid FP16/FP32 accumulation, and FlashAttention-2 parity, shipped as a drop-in vLLM plugin. Read online: https://www.cloudrift.ai/blog/optimizing-gemma-4-12b-rtx
Loading