yasp benchmarked three models with distinct computational profiles — RegNet, DeepSeekV3 MLA and a MiniGPT Block — on Azure’s AMD Radeon PRO V710 GPU against torch.compile on an Nvidia A10. yasp.compil
…
Benchmarking AI inference on Azure, yasp’s agentic compiler delivered a 2.91× speedup for the MiniGPT Block on AMD’s Radeon PRO V710 GPU over torch.compile. This technical deep-dive breaks down exactl
…
Whisper runs end-to-end on Chimera. Encoder and decoder compiled as native GPNPU kernels, INT4 weights, FP16 attention, top-1 token match against the float32 reference. Scales to four cores with a fla
…
3 Results
Share
Website Search
Wishlist
Add some favourites to your wishlist to get started!