Siddhant Ray
I am a fourth year PhD student in Computer Science at the University of Chicago, advised by Junchen Jiang and Nick Feamster. Overall, I am interested in building efficient inference systems for LLM and Physical AI applications, using application-specific feedback.
My current research focuses on two directions. First, I am developing new Vision-Language-Action (VLA) model architectures and serving systems for resource-efficient inference of long-horizon physical AI workloads. Second, I am exploring how to achieve better cost–quality tradeoffs in long-horizon LLM agent serving systems through joint agent-memory and KV Cache management.
Earlier I have worked on joint quality-delay optimizations in Retrieval-Augmented-Generation systems with query level configuration selection and resource management (METIS, SOSP'25). I also worked on Transformer models for per-packet latency prediction enabling tail-latency reduction for latency sensitive applications (SwiftQueue, NINeS'26).
In the past, I have worked on advances in Software Defined Networking, programmable networks and cloud computing. Additionally I have spent some time developing NLP techniques to analyse political corpora.
I'm fortunate to be additionally supported by the Liew Family Graduate Fellowship. Prior to starting my PhD, I earned my MSc in Electrical Engineering and Information Technology at ETH Zurich and my B.Tech in Electronics and Communication Engineering at VIT Vellore.
News
| Sep, 2026 | Argo: Efficient Importance Labeling for Enterprise Email Systems accepted at PACMI’26. |
|---|---|
| Jun, 2026 | Started my research internship at Microsoft, Redmond working on efficient Physical AI inference. |
| Jan, 2026 | SwiftQueue: Optimizing Low-Latency Applications with Swift Packet Queuing accepted at NINeS’26. |
| Oct, 2025 | METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation accepted at ACM SOSP’25. |
| Sep, 2025 | Serving as a Reviewer for AAAI’26 and ICLR’26 . |
Selected publications
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge FusionIn Proceedings of the Twentieth European Conference on Computer Systems 2025