CloudRift

CloudRift is the operating system for sovereign AI deployments.

We give data center operators and enterprises a single control plane to manage GPU fleets, launch customer workloads, and serve LLM inference, all within their own security perimeter.

Inference performance is where we invest heaviest. Our serving stack pairs OpenAI-compatible endpoints with hardware-aware optimization, and Emmy, our open source ML compiler, generates GPU kernels that have outperformed cuBLAS, vLLM, and Llama.cpp in published benchmarks.

The platform handles full virtualization across VMs, containers, and bare metal, runs on both NVIDIA and AMD hardware, and is SOC 2 certified.

Press Releases

1 Results

    Sponsor Editorial

    Tuning Qwen3 Coder from 277 to over 1,200 tokens per second on the RTX PRO 6000 and RTX 5090.

    Outperforming vLLM and Llama.cpp on Gemma4-12B

    01 Aug 2026 CloudRift Dmitry Trifonov, Slawomir Strumecki, Ivan Oleynikov
    How Emmy, our ML compiler, generates CUDA kernels that beat cuBLAS on the RTX 5090.
    RTX PRO 6000 vs H100, H200, and L40S: LLM Inference
    Single-GPU and multi-GPU LLM inference compared across RTX PRO 6000, H100, H200, and L40S.
    The True Cost of GPU Ownership: Computing Run Costs for Self-Hosted AI Infrastructure
    Breaking down the actual cost of owning and operating GPU hardware, from electricity and depreciation to maintenance and colocation.
    Long-context LLM inference benchmarks across NVIDIA B200, H200, H100, and RTX PRO 6000.
    5 Results
      Loading

      Contact Exhibitor