CloudRift

The Operating System for Sovereign AI and Inference Optimization

Datacenter operators and enterprises get one control plane to manage GPU fleets, rent capacity on demand or on reservation as VMs, containers or bare metal, and serve LLM inference through OpenAI-compatible endpoints, all inside their own security perimeter and under their own brand.

Inference performance is where we invest most. Emmy, our open source ML compiler, generates GPU kernels that run up to 1.6x faster than cuBLASLt on published Gemma 4 12B shapes and improves time to first token against stock vLLM.

Runs on NVIDIA and AMD hardware, SOC 2 certified, in production with telcos and datacenter operators in North America and Central Asia.

See it live: Tuesday September 15, 10:45, Expo Theater 2, "From Rack to API: An Enterprise Engine for High-Performance Inference," and at kiosk 1001 in the Start-up Zone.

Press Releases

1 Results

    Sponsor Editorial

    Tuning Qwen3 Coder from 277 to over 1,200 tokens per second on the RTX PRO 6000 and RTX 5090.

    Outperforming vLLM and Llama.cpp on Gemma4-12B

    01 Aug 2026 CloudRift Dmitry Trifonov, Slawomir Strumecki, Ivan Oleynikov
    How Emmy, our ML compiler, generates CUDA kernels that beat cuBLAS on the RTX 5090.
    RTX PRO 6000 vs H100, H200, and L40S: LLM Inference
    Single-GPU and multi-GPU LLM inference compared across RTX PRO 6000, H100, H200, and L40S.
    The True Cost of GPU Ownership: Computing Run Costs for Self-Hosted AI Infrastructure
    Breaking down the actual cost of owning and operating GPU hardware, from electricity and depreciation to maintenance and colocation.
    Long-context LLM inference benchmarks across NVIDIA B200, H200, H100, and RTX PRO 6000.
    5 Results
      Loading

      Contact Exhibitor