Enterprises trying to scale artificial intelligence (AI) workloads are running into a hardware market where graphics processing units (GPUs) are effectively rationed and colocation space is increasingly reserved for the largest cloud providers.
That scarcity is reshaping infrastructure planning. Some companies are stockpiling capacity before they know what they will use it for, while others discover late in the process that they lack the power or cooling to run what they have already committed to build.
“All capacities, all GPU cards, are running at full scale, so none are wasted,” said Yitian Xu, head of solutions for Alibaba Cloud’s UK and Ireland, Nordics and Cyprus region. “All customers who want to deploy GPU cards have to wait.”
Xu said customers now have to commit to one- or two-year contracts to reserve capacity, as demand from both external customers and Alibaba Cloud’s own internal model use keeps climbing. He linked the strain to a wider constraint: the UK’s limited electricity supply, which he called a national growth problem.
Retika Gupta, a solutions architect at Enterprise Rent-A-Car, said buying GPUs has become difficult for the same reason. Suppliers routinely quote six to 12 months for delivery, she said, and the timeline can shift again after an order is placed.
She said the uncertainty raises a harder question before any purchase: whether it is worth buying high-capacity GPUs at all, given they are likely to be outdated within 18 months, rather than leaning on cloud capacity or existing infrastructure instead.
“When COVID happened, we all went into hoarding essentials,” said Uma Mudigonda, vice president of solutions at Kyndryl. “That has to stop. It is going to bite us pretty soon.”
She described a recurring pattern among her customers: businesses declare themselves ready to move from testing into private AI infrastructure, only to discover during the design process that they lack the power to support it.
“The private AI vendor gives you the kit, but you have to host it,” she said. “When there is a colocation and a service provider, there is an opex (operating expense). A lot of things are playing, and then when the business case stacks up, the question is, do I need this?”
Hybrid by design
The panel took place at DCD Connect London on September 16, organized by DatacenterDynamics (DCD).
Titled “Hybrid cloud in the AI era: best practice or a model under pressure?” and moderated by Dan Swinhoe, editor-in-chief of DCD, the session examined how enterprises are deciding where to place AI workloads as cost, capacity and sovereignty pressures mount.
Kyndryl describes itself as the world’s largest infrastructure services provider, managing mission-critical systems for clients across 60 countries. Alibaba Cloud is both a cloud provider and a model developer, with its own custom AI chips built as an alternative to Nvidia hardware.
Gupta said hybrid infrastructure is not fading away in her experience, despite years of industry rhetoric about moving everything to the cloud.
“It’s not dying. It’s our reality, so it’s not going anywhere,” she said. “You do have a capital expenditure which you have already spent. You’re not going to get rid of it at one moment.”
Enterprise Rent-A-Car runs a mix of owned data centers, colocation facilities and both private and public cloud. The company briefly pursued a cloud-first strategy but abandoned it once AI workloads made the case for hybrid clearer. Gupta now designs each application first, establishes where its data lives and what compute it needs, and only then decides where it should run.
John Bradshaw, field CTO for Akamai Technologies in Europe, the Middle East and Africa, described a telecommunications client in South America that connected WhatsApp to its customer relationship management (CRM) and provisioning systems, letting customers resolve issues without calling a contact center. The workflow reached production and touched live network provisioning and customer data.
“Their agentic workflows were delivering real value to them very quickly, and they loved the outcomes they were getting,” Bradshaw said. “What they stopped loving were the bills from the frontier LLMs (large language models), because they were working out to be more expensive than having an agent able to answer the phone.”
He said organizations are increasingly willing to commit to specific workloads once the value is easy to validate, but large capital outlays remain risky given how fast the field moves. Eighteen months ago nobody used the term “agentic,” he said, and in another 18 months something equally unforeseen is likely to reshape the calculus again.
Akamai provides GPUs as a service along with model routing and gateway infrastructure delivered globally.
Xu said Alibaba Cloud has built a routing gateway that judges each incoming request by task and complexity before sending it to the most cost-appropriate model.
“Normally, customers are okay to have the copilot, so they adapt quite quickly,” he said. “For the autonomous side, they are always thinking about the scenario.”
He said he advises customers to keep a human in the loop for autonomous agents, and to avoid letting them run unsupervised on financial transactions or other mission-critical systems.
Sovereignty enters the calculus
Data sovereignty has become as pressing a question as cost or capacity for many enterprises weighing where to place AI workloads.
“Conversations used to be about which cloud to use. But now it is, what am I getting, and where am I allowed to run,” Mudigonda said. “There is a lot of confusion around sovereignty, around user data, where the data is.”
She described a case in which a US company’s acquisition of a Netherlands-based company was blocked by the Dutch government because the target hosted citizen data in its facilities and the deal was judged a sovereignty risk.
Xu said European customers are increasingly pressing Alibaba Cloud on where their data and models are physically hosted.
“Are you sure that all the data, all the models, are deployed on the EU data centers?” he said. “This is an increasing question to us.”
He said some customers in regulated sectors such as education and finance are choosing to compress large open-weight models to run on smaller, on-premises hardware rather than rely on cloud-hosted large models, partly for governance reasons.
The panel split over how much the open-versus-closed choice still matters.
“Open weight is not the decision factor for me,” Mudigonda said. “Enterprises are picking a small model, not the large model, and the situation is, they have already started long back with AI and ML (machine learning).”
“I don’t agree with that part, because as an architect, we decide that,” Gupta said. “Even if it has been decided by someone that we are going to use that model, we can go and challenge it.”
Gupta said no single model works for every use case, so architects need to keep re-evaluating the choice as needs change; a decision that was right a year ago may not hold after 18 months. Mudigonda said she agreed with the underlying point but that in her own customer base the decision was often already locked in by the time she became involved.
Asked whether colocation providers increasingly prefer to sell entire buildings to hyperscalers rather than retail space to individual enterprises, Bradshaw said the shift toward cloud will continue by necessity wherever local capacity runs short.
“If there isn’t sufficient capacity in the geographic region you’re operating in, the supply chain isn’t there. Then by necessity, it’s going to shift,” he said.
With sovereignty rules tightening and hardware still scarce, the panel’s shared expectation was that hybrid infrastructure, rather than a clean move to any single cloud, will keep defining how enterprises build for AI in the years ahead.



