Workshop
vLLM Inference Meetup Bengaluru – Virtual
Join us virtually for a deep technical dive into vLLM and llm-d and explore the engineering behind running LLM inference at scale. Hear from maintainers, core contributors, and…
Join us virtually for a deep technical dive into vLLM and llm-d and explore the engineering behind running LLM inference at scale. Hear from maintainers, core contributors, and industry experts from Red Hat, IBM, AMD, and NxtGen as they cover distributed inference, GPU optimization, ROCm & WideEP, intelligent semantic routing, sovereign and agentic AI workloads, and real-world production deployments.
Tune in from anywhere to learn how the open-source ecosystem is pushing the boundaries of high-performance, scalable AI inference.
Agenda :
| Time | Session | Speaker(s) | Organization |
|---|---|---|---|
| 1:30 PM – 2:45 PM | Keynote1 : Sovereign AI Inference | Sudhir Dharanendraiah | Red Hat |
| 1:45 PM – 02:00 PM | Keynote2: Tokenomics and Private AI Inference | Ramakrishna Yekulla | Red Hat |
| 2:00 PM – 2:30 PM | vLLM & llm-d Updates — Latest updates and developments across vLLM and llm-d. | Prasad Mukhedkar | Red Hat |
| 2:30 PM – 3:00 PM | llm-d for Sovereign & Agentic Workloads — Distributed inference stack, optimization, and benchmarking for sovereign and agentic workloads. | Pravein Govindan Kannan | IBM |
| 3:00 PM – 3:30 PM | Distributed Inference on ROCm with WideEP on vLLM & llm-d — Distributed inference on AMD ROCm using WideEP with vLLM and llm-d. | Chaitanya Sri Krishna Lolla & Sirra Ajith | AMD |
| 3:30 PM – 4:00 PM | Break | — | — |
| 4:00 PM – 4:30 PM | Inside vLLM Semantic Router: Intelligent Routing for LLM Inference — Intelligent, cache-aware routing based on request semantics, model capabilities, workload characteristics, and infrastructure constraints. | Aayush Saini | Red Hat |
| 4:30 PM – 5:00 PM | Scaling Inference at NxtGen Using the vLLM Ecosystem | Abhishek Kumar Singh | NxtGen |