Clockwork Fleet-Monitoring: A Layered Architecture for GPU Network Observability

Clockwork.io Stand: Booth 821
Large-scale GPU fabrics don't fail like ordinary networks — they fail partially, directionally, and in ways amplified by synchronized collective communication. A hot spine, a marginal optic, or an asymmetric ECMP path can turn into a job-wide straggler while every traditional health signal (link status, switch reachability, utilization) still looks green.
This whitepaper lays out Clockwork's layered architecture for closing that gap, and the case for why no single telemetry source — edge or switch — can do it alone.
Loading