If vLLM already solved LLM serving, why did SGLang appear?
Updated on
Published on
By
Table of Content
About the author
India’s AI conversation has a framing problem. Everyone keeps talking about training – who’s building the best models, who has the most GPUs, who’s winning the parameter race. It’s a reasonable thing to care about, except it’s not actually where the game gets decided. The part that decides things is what happens after the model is built, when it has to answer a million queries a day, then ten million, then a hundred million. In Hindi, Tamil, Bengali, and other regional languages – across banking, agriculture, healthcare, and half a dozen government services running all at once. That’s inference. And it’s expensive in a way that training simply isn’t.
To put a number to it: training ChatGPT-4 cost over $100 million, which sounds like a lot until you realize the annual inference bill to serve it runs to around $250 million annually and keeps climbing as usage grows. Training is a one-time cost, while inference never stops. And for India specifically, with 1.4 billion potential users and a deployment surface that spans more languages, sectors, and connectivity conditions than almost any other market, the inference opportunity is bigger than most people are currently pricing in.
The clearest signal that the inference market in India is real is through the where serious capital is being allocated. In February 2026, Blackstone led a $1.2 billion financing round into Neysa, making it India’s second unicorn of 2026 at a $1.4 billion valuation. That capital is going toward the deployment of over 20,000 GPUs in India, scaling Neysa Velocis as purpose-built AI infrastructure for enterprises, hyperscalers, and global AI labs looking to deploy in India.
And Neysa isn’t the only signal. India’s AI data center capacity is expected to nearly triple from 1.6 GW today to around 5 GW by 2030, with 650,000 to 700,000 GPUs projected to be deployed in Indian data centers over the next five years. Global hyperscalers have committed approximately $80 billion to India between 2026 and 2030. The IndiaAI Mission has built a national shared compute pool that now exceeds 38,000 GPUs, with a target of 100,000 by late 2026.
Taken together, these are a coordinated read on the same underlying thesis that India’s inference demand is going to be enormous, and the infrastructure to serve it doesn’t yet exist at the required scale.
The energy question tends to get underestimated. Once a model is deployed and running, inference ends up accounting for roughly 90% of the total energy consumed by that system over its lifecycle – training is a blip by comparison. AI-grade data centers need hundreds of megawatts of reliable power, and the next generation of facilities is being planned at gigawatt scale. India’s grid is already under significant pressure from industrial and digital growth, and adding AI infrastructure on top of that without a parallel energy strategy creates a real constraint.
Then there’s the memory problem, which is more technical but equally consequential. For the kinds of large models that actually get deployed in production (7 billion parameters and above) – the bottleneck during inference isn’t compute speed, it’s memory bandwidth: specifically, how quickly the GPU can pull model weights from memory and process them. DRAM prices surged approximately 130% in late 2025 and into 2026, which makes this a cost problem as much as a technical one for anyone running inference at meaningful scale right now.
And underneath all of this sits a dependency that doesn’t get talked about enough. NVIDIA controls somewhere between 80 and 90% of the advanced AI accelerator market, and the optimization techniques that matter most for inference – quantization, kernel-level tuning – are largely designed to run on NVIDIA’s CUDA and TensorRT stack.
Most of India’s inference infrastructure at this point runs on foreign silicon inside cloud environments owned by companies headquartered abroad. The risk isn’t that India would fall behind it’s that India becomes a very large AI market while remaining structurally dependent on others to run it.
Take the multilingual requirement. The fact that India needs AI to work across 22+ scheduled languages isn’t just a localization challenge – it’s a forcing function for building more efficient inference systems. Generic large models built for English-dominant environments are expensive to run and don’t perform well enough for Indian language users at the quality thresholds that actually matter for real deployment. That pressure is pushing Indian teams toward smaller, more efficient models and localized inference architectures that happen to be exactly where the broader global AI industry is heading anyway. India is developing real depth in this space ahead of most markets.
UPI, Aadhaar, and ONDC demonstrated something important: India can deploy digital systems at population scale, reliably and cost-effectively, and do it faster than almost anyone expected. That same infrastructure layer is a natural foundation for AI services. The deployment problem – getting AI into the hands of hundreds of millions of users across wildly different contexts, is one India has already solved once.
The practical question for any enterprise or government institution building AI services in India right now is what inference at scale actually requires beyond just buying GPUs and pointing models at them. A few things matter more than most people account for upfront:
The more achievable route to AI leadership for India probably doesn’t run through training the largest models – it runs through mastering deployment at scale. Inference is where that gets tested, and the infrastructure decisions being made right now will determine how that test goes.
Deploy, run, and scale on infrastructure designed to keep your AI moving forward.

While you were planning AI strategy decks, others were shipping products. AI Cloud has already reshaped hiring, infrastructure, and innovation speed. This blog breaks down what’s changed, who’s gained, and why every delay now comes with a cost.
In practice, doctors do not interact with an “AI model.” They interact with a workflow. They open a patient record, review symptoms and, examine scans. They consult the lab results. If AI adoption in healthcare has to succeed, the system must fit within their existing rhythm.