The Next AI Infrastructure Bottleneck Isn’t the Model. It’s the Machine. | nasscom

WorkAI.TV Editorial Desk
4 Min Read

Share with your CTO

India’s reported order of roughly 9,000 NVIDIA Vera Rubin GPUs for a Hyderabad AI factory, with delivery expected in 2027, signals a country moving from GPU clusters toward full AI factory scale. At 72 GPUs per NVL72 rack system, that order represents approximately 125 rack-equivalent configurations, putting it at roughly a quarter of NVIDIA’s own 100MW, 40,000-GPU Vera Rubin reference architecture. The headline is about compute. The actual engineering problem is thermal: every watt of GPU power eventually becomes heat that must be continuously removed.

What this means for your business

The Rubin generation is designed for 100 percent liquid cooling, with coolant temperatures up to 45 degrees Celsius, and that spec isn’t a preference, it’s a constraint that propagates backward into facility design. CTOs evaluating AI infrastructure at this density will find that the cooling architecture determines whether the compute can be deployed at all, not the other way around. If your organization is planning or procuring AI data center capacity in the next 18 to 36 months, the question isn’t whether you can get GPUs. It’s whether your facility or your colocation provider’s facility was designed around rack-level liquid cooling from the start.

The piece, written for Nasscom’s community platform and therefore pitched toward an audience that benefits from Indian infrastructure optimism, is still analytically sound on the physics. Its core claim holds: liquid cooling at this scale isn’t a facilities upgrade bolted onto a compute purchase. It becomes a co-engineered part of the system. NVIDIA’s NVL72 integrates 72 Rubin GPUs, 36 Vera CPUs, sixth-generation NVLink, and 20.7 terabytes of HBM4 memory into a single rack-scale unit generating heat at densities that conventional raised-floor, air-cooled data centers weren’t designed to handle. Operators who treat cooling as a secondary procurement decision after locking in GPU count will find themselves with expensive hardware running at reduced utilization or, worse, thermal throttling that quietly degrades the AI workloads the hardware was bought to run.

India’s climate adds a variable that generic data center playbooks don’t price in. Heat rejection, the process of moving thermal load from coolant to the outside environment, gets harder as ambient temperature and humidity rise. A closed-loop liquid cooling system inside the facility still needs somewhere to dump that heat, whether through cooling towers, dry coolers, or chillers, and each of those options carries water consumption, energy overhead, and seasonal performance variance that differ meaningfully between Hyderabad and, say, a Nordic colocation site. The facilities that win at sustained GPU utilization over a multi-year operating life won’t necessarily be the ones with the most GPUs. They’ll be the ones where the power-to-compute-to-heat-rejection chain was optimized as a single system before the first server was racked.

Concept deep-dive: Thermal throttling

Thermal throttling is when a processor automatically reduces its clock speed to generate less heat when temperatures exceed safe limits, think of it as the chip pulling back to avoid damaging itself. In a single workstation it’s a minor annoyance. Across a 9,000-GPU cluster running continuous AI training jobs, it’s a silent tax on utilization, distorting the effective compute capacity the organization paid for. Poor cooling doesn’t cause visible failures; it causes expensive underperformance that’s easy to misattribute to software or model architecture.

Based on reporting from The Next AI Infrastructure Bottleneck Isn’t the Model. It’s the Machine. | nasscom, originally published 2026-09-15 05:14:00.

TAGGED:
Share This Article