Artificial intelligence (AI) agents that write and run their own code are driving a compute surge nobody predicted this fast.
Morgan Stanley estimated in January that global semiconductor industry revenue would top $1 trillion for the first time in 2026, a threshold this year’s AI buildout has pulled forward.
These tools behave nothing like human engineers. They do not sleep, and they spawn their own tasks inside data centers.
“The number of new models and new agentic workflows being deployed, from coding agents through to SaaS (software as a service) systems, is exploding, and that is continuing to drive the demand for general-purpose compute that goes with it,” said Paul Williamson, senior vice president of corporate ventures at Arm.
“I don’t think it’s something that many would have predicted only a couple of years ago,” he said.
Williamson pointed to GitHub star counts as early evidence, citing OpenCode, an open-source AI coding agent that broke out in March and kept climbing at what he called a phenomenal scale, against long-established projects such as Linux.
Arm went public with its own forecast the same month: general-purpose compute demand in cloud infrastructure would need to roughly quadruple.
“Since we said four times, other analysts are now saying eight or 10 times as much compute demand in the last few months,” Williamson said. “That is not a blip. As we look forward, projections show it is only scaling further.”
The five biggest hyperscale operators now have to figure out how to commission and architect their data centers from the ground up, fitting several times as much compute into the same power budget they have today, without adding a single new watt.
“It is genuinely mind-blowing to see the scale of compute now delivered,” he said.
Arm, whose architecture underpins the majority of the world’s smartphones, traces its roots to an old turkey barn in Cambridgeshire, where it was founded in 1990. Williamson said the company has spent more than a decade building central processing units (CPUs) capable of running at data center scale.
Not four but two
Williamson made the comments in a keynote at the Semiconductors to Systems Summit 2026, held in London on August 26 and organized by TechWorks in partnership with the UK Semiconductor Centre. The event brought together chip designers, systems engineers and policymakers to discuss the shift from semiconductors to complete systems.
Arm’s response arrived in March, when it launched the AGI CPU, its first direct silicon product built specifically for the data center, co-designed with Meta after more than 10 years of investment in CPU development.
“We’ve not got to four, but we’ve got to two,” Williamson said.
By co-defining the CPU, its input/output (I/O) structure and its rack-level tooling with Meta, Arm said it doubled performance per rack compared with what hyperscale operators had previously installed from Intel, while remaining within the same power envelope.
“If you double performance per rack, you’re looking at up to $10 billion in capex savings alone to roll out compute capacity,” he said.
He pointed to SoftBank Group’s data center buildout in France, where the company has pledged up to $87 billion for roughly five gigawatts of AI capacity, as an example of the scale involved.
“For the largest compute providers in the world, system-level optimization drives really meaningful financial returns,” he said.
That optimization is now as much a physical engineering problem as a chip design one. A standard air-cooled rack draws 36 kilowatts and packs 8,160 CPU cores with more than 180 terabytes of low-latency memory.
“This thing needs a reinforced floor just for the weight of the copper cooling and liquid handling carried in each unit,” he said.
The liquid-cooled racks he described draw 200 kilowatts each, pack 45,696 CPU cores and carry more than a petabyte of low-latency memory apiece, according to his own presentation slides.
“Compute is rapidly looking very different. This is now a system-level challenge to drive density of compute and efficiency at scale,” he said.
Edge AI two-tier split
Arm’s own business model is shifting alongside the hardware. Beyond its traditional licensing of CPU, graphics processing unit (GPU), neural processing unit (NPU) and system intellectual property (IP), the company now also offers Compute Subsystems, preconfigured bundles of that same IP, as well as complete production silicon, such as the AGI CPU.
“Many of you will be familiar with us as an IP company developing CPU IP. Arm has evolved a lot in the last 10 years,” Williamson said.
He said Arm continues to license IP across the industry, but now also delivers finished silicon directly to its biggest hyperscale customers.
Arm’s roadmap runs alongside that shift. The AGI CPU is shipping now, with AGI CPU 2 due in 2027, promising more cores, performance and efficiency, and AGI CPU 3 to follow. Its Compute Subsystems will progress from V3 to V4 and V5 on a similar timeline.
“This is not something you do just once,” he said.
“You’re either going to have lightweight, low-end, cloud-connected laptops, or high-performance AI-capable laptops. Basically two tiers of performance are emerging, and the use cases and workloads are shifting with them,” he said.
His own presentation slides placed Chromebooks, Nvidia’s DGX Spark and Apple’s MacBook among the edge AI devices already splitting along those lines, alongside wearables, smartphones and home hubs moving the same way, under the same two-tier logic he described for laptops.
“In-home devices are moving from cloud-connected voice interfaces to AI-capable systems with personal context and security for your life,” he said.
“I’ve seen a phenomenal diversity of industrial robotics applications, all experimenting in different forms of AI,” he said. “We will see people using AI to engineer and develop systems that look very diverse in their form factor, but all using common and interesting new models to execute.”
He made the point the same week as the World Humanoid Robot Games in Beijing, listing drones, robotic arms, cobots and humanoids as examples.
“We’ve seen a big shift to people optimizing their accelerator for inference token generation or for training, and it isn’t going to be one size fits all in the data center,” he said.
Williamson pointed to Arm’s adoption by Amazon Web Services (AWS), and more recently by Google and Microsoft, alongside Nvidia’s GB200 NVL72 and Vera CPU and Google’s own Cloud TPU racks, as evidence of an increasingly fragmented accelerator landscape now taking shape across the industry.







