Enterprises ran their own data centers a few decades ago. They owned the servers, controlled the racks, and managed every aspect of power, cooling, and cost. The public cloud has shaped the model for more than a decade. Now the pendulum is swinging back. Enterprises are investing heavily to ensure their business can securely harness massive computing power without sacrificing control, latency, overall performance, and, importantly, data privacy. Generative AI is edging toward agentic AI, shifting infrastructure demands from general-purpose virtual machines to high-density GPU clusters.
Why are companies thinking about going back to the on-premises data center models? High-density AI workloads and the need to protect data sovereignty and core IP are driving a new generation of bare-metal deployments that feel like the old-school on-prem model, using server racks with Nvidia Blackwell GPUs. The difference, though, is how you run the AI infrastructure with modern density, liquid cooling, and specialized providers handling the heavy lifting.
What does high-density AI mean? – It means packing massive computing power (like specialized GPUs and server racks) into a very small physical space, which demands far more electricity and advanced liquid cooling than standard servers.
Why is high density needed and what are its advantages? High-density architectures became essential because modern large AI models cannot fit or run efficiently on standard, dispersed server setups.
Minimizing latency – AI training and real-time inference require billions of parameters to communicate constantly. Packing GPUs inches apart allows ultra-fast interconnects (like NVLink) to transfer data with near-zero latency, avoiding networking bottlenecks across long cables.
Massive shared memory – Frontier models require terabytes of unified, high-speed memory (HBM). High-density rack architectures (such as NVIDIA NVL72) let dozens of chips pool memory and act as a single, massive GPU. This is crucial – the entire server rack acts as one single brain with multiple GPUs on it.
Space & infrastructure efficiency – Data centers have limited real estate and power hookups. Stacking 40 kW to 130+ kW into a single rack footprint maximizes compute capacity while enabling centralized direct-to-chip liquid cooling systems to remove heat far more efficiently than air cooling.
What does vertically integrated mean in this context? So, when we talk about ‘vertically integrated’ here, it means taking care of everything in the compute system from start to finish. This includes designing and managing all the layers, from specialized silicon like GPUs and rack-level interconnects such as NVLink to power distribution, direct-to-chip liquid cooling, and orchestration software. Instead of separate parts, everything works together as a single, unified system. In an ‘AI Factory,’ the rack itself is considered the entire computer, making the setup more seamless and efficient.
So, how are companies bringing data centers and bare metal back, and what does it actually take to run them? I’ve broken this down into a three-part story,
1. How are companies bringing bare metal back?
2. How are the chips actually used (Inference vs. Training)?
3. How are critical MEP (mechanical, electrical, plumbing) systems handled when someone else hosts the infrastructure?
Story 1 – Bringing the bare metal back – 2 approaches.
Traditional enterprise data centers gave companies full ownership and control of the infrastructure, from owning the real estate to managing the infrastructure on it. Today, two distinct approaches are restoring that opportunity without forcing every organization to build a new facility from scratch. These types of offerings are called neo cloud providers.
Approach A – Own the silicon, but outsource the facility
Enterprises or any system integrators purchase the actual hardware, such as Nvidia Blackwell GPUs, GB200 superchips, or even full server racks, directly from Nvidia or from server manufacturers and OEMs that integrate Nvidia platforms, and then ship those racks to specialized hosting providers.
These providers do not own the GPUs. They supply the physical space, power, cooling, security, and the full MEP infrastructure, while the customer retains ownership of the bare-metal servers. This arrangement is similar to the classical colocation model.
Approach B – List the capacity from specialized AI infrastructure operators
The providers acquire or lease large quantities of Nvidia hardware, install it in their own high-density data centers, and then offer multi-year bare metal leases to enterprises. Customers never take ownership of the servers. Instead, they receive dedicated, non-virtualized access for a fixed term. This model removes the customer’s capital expenditure while still delivering the performance isolation offered by true bare metal servers.
Both approaches restore the ‘my hardware, my performance’ experience that a pure multi-tenant hyperscaler cloud provider cannot provide. The difference is mainly ownership versus long-term lease and who carries the balance sheet risk for the GPUs.
So, what is common in both approaches?
1) Isolated bare metal servers and sovereign private deployments – The intelligence you build stays entirely your property without leaving your machine,
2) Zero ecosystem lock-in – The weight, fine-tuning data, and the intellectual property remain inside your code, completely decoupled from third-party software ecosystems,
3) Cloud management software – Provide a very thin layer of cloud management software and handle the complex networking.
Story 2 – How are the chips used? Inference versus testing and training.
Once the bare-metal servers are in place, the same Nvidia Blackwell GPUs (or application-specific ICs from other vendors) are typically used in two different modes for higher efficiency and throughput.
Mode 1 – Inference
This is production serving – answering real user questions, generating responses, running retrieval-augmented generation for a Gen AI search service, or powering low-latency applications moving toward agentic AI. Inference workloads care about consistent tokens per second, low latency, and high utilization across many concurrent requests. NVIDIA GPUs excel here, but purpose-built inference chips also matter. Examples include Amazon’s Annapurna Labs Inferentia chips and Google’s TPU variants optimized for serving. These ASICs often deliver better performance per watt or lower cost for steady-state inference once a model is trained and tested.
Mode 2 – Training and testing
This is the experimental and development side – model training, hyperparameter search, fine-tuning, evaluation, and spot or burst experimentation. These workloads are bursty, memory-intensive, and often require full floating-point performance and high-speed interconnects that Nvidia platforms provide. Training-oriented chips such as Amazon Trainium or Google’s higher-end TPU pods are also used. However, most enterprises still prefer Nvidia bare metal for flexibility across both training and inference on the same hardware.
Why does splitting the two modes help?
Keeping them separate stops training jobs from slowing down live customer traffic. Companies can also put cheaper specialized chips on steady inference work and save the more expensive NVIDIA hardware for the heavy experimental work. The result is better performance, lower cost, and cleaner operations.
Story 3 – Cost benefits and how MEP is handled when hosting for others.
Bare metal offers several advantages because no hypervisor sits between the virtual machine and the physical hardware. A portion of the system’s performance and capacity is consumed just to run the virtualization layer. In the long run, when utilization is high, long-term ownership or a multi-year lease becomes cheaper than on-demand cloud instances. You can run any software stack, quantization method, or network configuration without the cloud provider’s limitations.
When the provider hosts bare metal on behalf of the enterprise, it becomes responsible for the facility’s mechanical, electrical, plumbing, security, and construction. Let us discuss the three core components of running the AI factory.
Electrical (Power) – This is the primary operational parameter for AI factories. Power is needed not only to run high-density AI racks but also to operate the factory’s coolant systems. In addition to the power grid, factories also rely on UPS systems. Modern AI racks draw 40 kW to 130+ kW – for example, the NVL72 architecture. Hence, the electrical systems must handle very high density and often use 480 V or higher distribution to reduce current and cable size.
Mechanical (Cooling and HVAC) – This removes the heat generated by the AI factories. GPUs generally produce enormous heat. Proper mechanical systems should monitor heat and ensure GPUs stay within allowed temperature limits. Traditional air cooling will be used for comfort, humidity control, and ventilation of support spaces. In addition, the following two cooling methods are used to keep the GPUs within the optimum temperature range.
1. Direct-to-chip (DTC) liquid cooling – In this method, a special liquid, which is a combination of water and propylene glycol, is used to cool the GPUs.
2. Full immersion cooling – Here, the entire server rack is submerged in a dielectric fluid.
The difference between dielectric fluid and the water-plus-propylene-glycol mix is that the former is nonconductive and the latter is conductive. They serve very different purposes: the dielectric fluid contacts the chips directly, while the water + glycol mix is confined to closed plate loops. Thermal transfer in the water and propylene glycol mix is much higher than in the dielectric fluid.
Plumbing (Water and processed fluids) – This is the fluid-transport side that supports both the building and the system’s cooling. Common and vital elements include domestic water, sanitary drainage and venting, storm drainage, and fire protection (sprinklers, standpipes, fire pumps, etc.). The most important part for the AI data center, though, is managing the cooling water and process fluid systems. Let’s see in detail below.
1. Chilled facility water – Clean water circulating in closed loops between chillers and heat exchangers to carry heat outside of the data centers.
2. Water+Glycol mix (50/50) – Conductive liquid that circulates through cold plates on GPUs to absorb heat safely without freezing or bacterial buildup.
3. Dielectric fluid – Electrically non-conductive liquid that directly covers submerged live server racks without causing short circuits.
4. Makeup water – Fresh supply water that replaces liquid lost to evaporation in cooling towers.
5. Chemical treatment additives – Chemicals added to water lines to stop corrosion, mineral scaling, and algae.
6. Secondary containment or catch basins – Safety reservoirs and sensor zones that catch and detect leaks before liquids hit hardware.
In immersion cooling, the dielectric fluid itself is treated as a specialized process fluid, while the secondary loop that cools the dielectric fluid usually uses a water+glycol mix.
How do the 3 systems work together?
1) The electrical system delivers and manages power -> the GPU generates heat.
2) Mechanical cooling systems remove the heat.
3) The plumbing system observes the heat from the mechanical equipment and transfers it to the outdoor heat rejection systems.
In short,
1) Electrical = energy in,
2) Mechanical = heat management,
3) Plumbing = the fluid highways that move heat and serve the building.

What are the pros and cons?
Pros –
1) Full control and no hypervisor overhead.
2) Predictable costs with no token-based pricing.
3) Strong data and IP protection.
4) High performance for long-running agentic AI workloads.
Cons –
1) Higher upfront or long-term commitment.
2) Less elastic than public cloud.
3) Still depends on specialized providers for facility and MEP
4) Hardware can become outdated quickly
The return to bare metal is not a step backward; it’s an evolution driven by the need to maintain the proprietary data Enterprises own and to take advantage of the agentic AI that is defining how customers interact with the organization and how employees interact with enterprise systems. Moving to dedicated GPU instances provides the Enterprises true freedom from token-based pricing constraints, hypervisor overhead, and API rate limits. This helps run long-running agent flows without unpredictable bills. Ultimately, the future of agent AI depends not just on model size, but also on tokens per second per kilowatt of energy consumed, making full stock control of silicon, power, and high-density liquid cooling the true long-term competitive differentiator.
What are your thoughts on shifting to bare-metal infrastructure for high-density AI workloads? Share your perspective and experience in the comments.
Happy learning!