Sundara Raman Ramachandran
Sundara Raman Ramachandran is a Lead Engineer on LinkedIn’s LLM Inference team, where he leads the development of large-scale infrastructure powering LLM-based search and recommendation systems. His work focuses on building low-latency, high-throughput inference platforms that serve latency-critical ranking workloads at global scale.
He has led the production deployment of prefill-only LLM scoring systems and is a core contributor to the SGLang open-source project, where he led the development of the Prefill-Only Ranking API based on production requirements from large-scale enterprise deployments. His work spans LLM serving, distributed systems, GPU optimization, and inference performance.
Sundara has co-authored research papers accepted at MLSys 2026 and KDD 2026 and regularly shares his work with the AI infrastructure community through conference talks and technical publications. Prior to LinkedIn, he worked on Microsoft Azure Identity & Authorization and Microsoft Office. He holds a Master’s degree in Computer Science from The University of Texas at Austin.
Sessions
-
Beyond Text Generation: Scaling LLM Ranking Systems at LinkedIn17-Sep-2026Data & Models Track
