Share with your CTO
HPE is betting that the infrastructure gap between national supercomputing and enterprise AI has effectively closed, and it’s positioning its Cray-heritage stack as the bridge. Trish Damkroger, SVP and GM for HPC and AI infrastructure at HPE, made the case at the AMD Advancing AI event, pointing to the Frontier system at Oak Ridge and the forthcoming Discovery system, both AMD-HPE builds, as proof that HPC and enterprise AI infrastructure now share the same architecture. Liquid cooling, once a supercomputing specialty, is becoming a hard requirement as GPU thermal loads make air cooling economically absurd.
What this means for your business
The practical question for any CTO running or planning a GPU cluster isn’t whether liquid cooling is coming, it’s whether their facility can support it before their next hardware refresh forces the issue. HPE’s GX5000 platform is designed for warm-water cooling up to 45 degrees Celsius, which matters most to organizations building in Europe or anywhere with strict water efficiency standards. If your current data center was designed around air-cooled density assumptions, you’re already operating on borrowed time against vendors who aren’t.
The AMD-HPE partnership deserves more scrutiny than it typically gets in coverage like this, where HPE’s position as the builder of the world’s top supercomputing systems creates a natural incentive to frame national-lab credibility as transferable enterprise proof. It is, to a point. The engineering at Frontier and El Capitan is real, and AMD’s new EPYC 6 and Instinct MI430X give HPE a competitive silicon story against Nvidia-centric alternatives. But supercomputing procurement cycles and enterprise AI refresh cycles run on completely different timelines, and the organizational complexity of deploying Cray-heritage infrastructure in a commercial data center is genuinely different from dropping it into a national lab with a dedicated facilities team.
The claim that agents are already running scientific simulations autonomously is the most consequential and least substantiated point in the piece. If that’s true at scale, it reframes how CTOs should think about the staffing model for simulation workloads, fewer PhDs running jobs manually, more engineers building the agentic pipelines that run them. I’d revise the enthusiasm here if it turns out “agents steering HPC workflows” means workflow scheduling automation rather than genuine autonomous scientific reasoning, which is a much more modest claim wearing more impressive clothes.
Concept deep-dive: Liquid cooling
Liquid cooling routes chilled or warm water directly through server components rather than blowing air across them, the way a car radiator replaces a desk fan for a high-performance engine. It exists because modern AI accelerators generate heat densities that air simply can’t remove fast enough at reasonable cost. The business consequence is physical: liquid cooling requires building infrastructure, pipes, cooling distribution units, facility retrofits, that air-cooled data centers were never designed to support, turning a hardware decision into a capital planning problem.
Based on reporting from HPC AI infrastructure convergence drives HPE strategy, originally published 2026-07-25 13:07:00.

