Key Takeaways
- Broadcom reported that 56% of surveyed professionals plan to run production AI workloads in private clouds.
- New Belgium Brewing and US Senate Federal Credit Union described using specialized models and on-premises infrastructure for operational and security use cases.
- Falling model requirements, governance demands and token costs are making hybrid AI architectures more practical.
Enterprise enthusiasm for private AI is beginning to look less theoretical. At VMware Explore, customers described running open-weight and open source models inside their own infrastructure, while Broadcom argued that economics and governance are shifting production workloads away from an overwhelmingly public-cloud model.
The change is not a wholesale retreat from public cloud. It is closer to a recalibration. Sensitive data, latency requirements and the cost of repeatedly moving or processing large data sets can make local infrastructure attractive. Meanwhile, smaller models are becoming capable enough to run on older GPUs and, for some specialized tasks, CPUs.
A VMware product owner at a North American bank reported a shift in perspective since Broadcom leadership promoted an on-premises future during a 2024 keynote. The product owner initially dismissed the idea, but now sees GPU capacity as part of the next generation of workloads the internal service-provider environment will need to accommodate.
That shift mirrors a wider investment pattern. IDC reported that AI infrastructure spending reached $47.4 billion in 2024, up 97% year over year, and projected that it would exceed $200 billion by 2028 (source). Dedicated cloud and private infrastructure services also reached about $20.4 billion in 2024, nearly twice the level recorded in 2021. Those figures suggest private infrastructure is becoming one component of scaled AI programs, rather than an isolated alternative to cloud services.
New Belgium Brewing offered a concrete example. The organization's director of IT operations noted the brewer has used open models developed for Nvidia's Omniverse platform to visualize and optimize centrifuges within its manufacturing pipelines. As New Belgium Brewing expands agentic automation across brewing and business processes, it is selecting different models for controls data, preventive maintenance and other workloads.
"A frontier model is not fit for purpose for everything," the director said. If a specialized model can perform a narrow industrial or administrative task with less compute, enterprises may not need the latest GPU for every inference request. The ability to exchange models also reduces dependence on a single provider, although it introduces work around evaluation, version control and monitoring.
US Senate Federal Credit Union took a similar approach while developing its first cybersecurity AI agent. The credit union's CIO noted that keeping data on-premises provided security, control and infrastructure visibility. VMware Private AI Foundation Deep Learning VMs in VCF 9 gave the team a familiar environment for development and testing, supported by NSX and vDefend.
Infrastructure location does not create trustworthy AI by itself. Enterprises still need controls for data access, model behavior, human oversight and ongoing testing. The NIST AI Risk Management Framework provides one structure for identifying and managing those risks, including generative AI concerns. Private deployment can support that work by giving operations teams more direct control, but governance depends on policy and execution as much as hardware placement.
Broadcom also highlighted the economic factors, noting that token costs are driving market shifts. A recent survey of 1,800 professionals indicated that 56% planned to run production AI workloads in private clouds. Meanwhile, public-cloud preference for AI operations fell by 15%, reaching 41%.
The direction fits the broader market. Gartner estimates worldwide AI spending will reach $2.596 trillion in 2026, as organizations move beyond experiments and confront recurring production costs (source). In many enterprises, workloads are likely to settle into a managed mix: public cloud for elasticity and external services, with private infrastructure for governed, data-intensive and predictable AI deployments.
⬇️