Connect with us

AI

Google Cloud Survey: 83% of Firms Need Agentic AI Infrastructure

Google Cloud’s State of AI Infrastructure report found 83% of 1,400 IT leaders need upgrades for agentic AI, with 62% paying an inference tax.

Published

on

Google Cloud surveyed more than 1,400 senior IT leaders for its State of AI Infrastructure report, and the headline number reaches into every enterprise AI roadmap. 83% of organisations now say they need infrastructure upgrades to support production-grade agentic AI. The same survey tracks inference costs, governance pressure, deployment topology, and energy use.

The findings land as the cloud provider’s clearest case yet that compute, not algorithms, has become the binding constraint on rolling out agents that act on their own. Enterprise AI has moved from conversational assistants that answer to systems that act, the report argues, and that shift hits the underlying infrastructure harder than anything before it. Earlier chatbot workloads ran in short bursts and could tolerate round-trip latency; agentic workloads run continuous reasoning loops that hold massive context in memory at once. HCLTech’s June 2026 expansion of its Google Cloud and ServiceNow pact tracks the same shift from the system-integrator side. The survey doubles as the diagnosis and the prescription.

The 83% Who Need a New Foundation

The gap between AI ambition and infrastructure reality is widening, Google Cloud writes, and 83% of organisations in the survey say they require upgrades to support production use of agentic AI. Agentic AI is described in the report as systems that do more than respond to prompts and instead carry out tasks and workflows on their own. Older stacks were built for chatbot traffic, where a prompt came in, a response went out, and the loop closed. Continuous reasoning loops that hold massive context in memory at once stress those stacks at a different level.

The same organisation collecting the answers is also positioning itself to supply the fix. Every section of the survey that follows pairs a finding with a vendor product. The framing is consistent: the compute is the constraint, and only integrated silicon, storage, and orchestration can lift it.

The Inference Tax Nobody Budgeted For

The most painful line item is what Google Cloud calls the “inference tax”. In the survey, 62% of leaders say they are seeing a significant inference tax driven by data egress fees, storage bloat, and idle specialised hardware. That is a per-token cost layered on top of the model price itself, and few finance teams have modelled it. Most procurement conversations still price AI on a per-token basis without the egress and storage overhead that agentic workloads routinely add. The number is the gap between the quoted price and the realised cost.

On top of the inference tax, 81% of leaders cite operational complexity as a hidden cost of scaling AI. Stitching fragmented tools together to keep agents running now has its own labour line on the budget. The fix Google Cloud proposes is fluid compute: matching silicon to workload rather than running everything on a single architecture.

  • 83% of organisations require infrastructure upgrades for agentic AI.
  • 62% are paying a significant inference tax on data egress, storage, and idle silicon.
  • 81% cite operational complexity as a hidden cost of scaling AI.
  • 78% now source generative AI from their primary cloud partner, up 30 points from 2025.

For heavy training, the new TPU 8t is positioned to deliver scale for the world’s most sophisticated models, with the vendor claiming nearly three times the performance of the prior generation. For low-latency inference, the TPU 8i is purpose-built to maximise on-chip memory so agents can think and react in real time. For orchestration, Arm-based Axion CPUs run the AI control plane, reinforcement learning simulations, and agent coordination cost-effectively. The survey treats the misallocation of workloads onto the wrong silicon as the source of the inference tax.

Agent Sprawl and the Governance Rush

When agents read email, query databases, and execute workflows across a business, visibility is the first casualty. 79% of technology leaders cite security, governance, and MLOps as their top challenge to scaling inference, Google Cloud says. Agent sprawl, defined in the report as autonomous systems scattered across tools and platforms with no central oversight, has become the operational problem of the year. The risk is concrete: an untracked agent with broad permissions can act on data the security team has never seen. The fix has to be wired in before agents go live, not bolted on later.

Governance, in the report’s telling, has to mature before innovation can scale, and that means a centralised control plane acting as a single system of record for agent permissions, identity, and workflows. Google Cloud points to Agent Gateway as the reference architecture, with precise read/write scopes, full audit trails on every interaction, and human-in-the-loop approval before an agent takes a sensitive action. The same governance-first pitch shows up outside Google Cloud, as the Gitlab deal that turned governance into a product demonstrates.

The same drive for unified governance is changing vendor behaviour at the buyer level. 78% of organisations now source their generative AI solutions directly from their primary cloud partner, a 30-point increase from 2025, according to a separate Google findings study on agent deployment from September 2025. The implied move is away from bolt-together best-of-breed toward integrated platforms where governance and audit already exist. Buyers now get less vendor choice but more confidence that the controls are wired together.

Where the Workloads Actually Run

The cloud-vs-on-premises fight is over, the report argues; hybrid is the destination. 52% of organisations now use a hybrid multicloud architecture, and 48% of leaders are prioritising infrastructure with strict data residency controls. Sovereignty and data gravity, rather than vendor preference, are driving topology decisions.

The push toward edge deployment is the second leg of that move. 90% of organisations now rank edge deployment as important for AI initiatives, with 72% describing it as extremely or very important. Real-time agents in voice, video, and financial trading cannot absorb round-trip latency to a distant data centre. Edge systems also keep manufacturing plants, retail stores, and hospitals running when connectivity drops. The third motivation is cost: always-on reasoning in the cloud is expensive, and highly optimised models on edge devices shift the compute burden locally.

  • Hybrid multicloud: 52% of organisations have made it the default architecture.
  • Data residency: 48% of leaders are prioritising infrastructure with strict residency controls.
  • Edge deployment: 90% rank it important for AI; 72% call it extremely or very important.

Together the three numbers describe an infrastructure footprint that is layered, not centralised. Public cloud holds the broad compute that needs scale; private stacks keep the workloads that have to stay close to the data; edge devices absorb the latency-sensitive traffic. The trade-off is operational complexity, which the survey already lists as a top cost. The report’s answer is co-designed silicon and orchestration rather than more bespoke integration.

The Energy Wall Hits the Hardware Order Form

Power is moving from sustainability report to operating constraint. 91% of leaders now factor power consumption into their hardware selection, with 61% rating it as a primary or significant factor. Buying more silicon in some regions now means competing for grid capacity; the constraint has moved upstream. Energy budget is now a procurement gate, decided before any benchmark runs, and the buy decisions look different as a result.

Regulation is tightening in parallel. In Germany, new data centres must achieve a Power Usage Effectiveness of 1.2 or lower. Ireland now mandates that large data centres provide 100% on-site dispatchable generation to match their grid draw. For builders of agentic AI, the bottleneck has shifted from chip supply to power purchase agreements, and the timeline is no longer in their hands.

The economics compound. High-power hardware requires larger capital outlays on cooling, specialised rack designs, and facility upgrades before the first token is served. Google Cloud argues that performance per watt is the metric that has to move, and cites its TPU 8t as the proof point. The vendor claims the chip delivers nearly three times the performance of the prior generation while being up to twice as energy-efficient. Buyers are grading silicon on energy alongside throughput.

Google Cloud’s Answer in Its Own Stack

The natural reading of the State of AI Infrastructure report is that it doubles as a buying guide. The unified data layer the report champions is a sales surface for Smart Storage and Cross-Cloud Lakehouse. The governance pitch lands on Agent Gateway. The energy pitch lands on TPU 8t. Each section pairs a finding with a product, and the published overview of the report findings walks through both halves at once.

Fluid compute pulls the silicon story into one frame. Three chip families, three jobs. The TPU 8t handles heavy training, scaling to the most sophisticated models with a claimed threefold performance jump over the previous generation. The TPU 8i handles low-latency inference, purpose-built to maximise on-chip memory so agents can think and react in real time. Axion, Google’s Arm-based CPU, handles orchestration, running the AI control plane, reinforcement learning simulations, and agent coordination at general-purpose economics. The same workload never lives on one chip in the architecture, and the survey treats that mismatch as the source of the inference tax.

Chip family Workload Why it fits
TPU 8t Heavy training Compute accelerator for the largest models; vendor claims nearly 3x prior-generation performance and up to 2x energy efficiency.
TPU 8i Low-latency inference Purpose-built to maximise on-chip memory so agents can think and react in real time.
Axion (Arm) Orchestration and RL General-purpose CPU for the AI control plane, reinforcement learning simulations, and agent coordination.

Splitting training, inference, and orchestration across three silicon families is not unique to Google Cloud. Nvidia’s CPU-GPU-NIC stack and the AMD-Microsoft Azure alliance point at the same problem. The full State of AI Infrastructure report frames integration at the silicon layer as the cost-effective response to that interoperability overhead.

From Pixels to Industrial Robots

One last move in the report connects infrastructure to the physical world. Google Cloud describes a generation of autonomous robots that sense, simulate, and navigate the physical world. Those robots train millions of times in digital twin simulations on Google Cloud before they ever set foot in a real factory. The same fluid compute foundation is what makes those simulations tractable at the scale the survey describes.

From industrial inspections to cinematic videography, the report notes that AI is now solving tangible problems in the real world. The agent stack that runs in a browser is now being aimed at robots, drones, and other real-world hardware. Digital twin simulations on Google Cloud are where the robots learn, before they ever set foot in a real factory. The same fluid compute foundation the survey recommends for data centres is now being used to train those simulations. The constraints that pushed 83% of organisations to plan upgrades apply to physical AI at the same scale.

Frequently Asked Questions

What is Google Cloud’s State of AI Infrastructure report?

The report is built on a survey of more than 1,400 senior IT leaders, published by Google Cloud to characterise the gap between AI ambition and infrastructure readiness. It frames enterprise infrastructure as the binding constraint on rolling out agentic AI in production. Each section pairs a survey finding with a vendor product, from TPU silicon to Smart Storage, Agent Gateway, and Cross-Cloud Lakehouse.

What is the “inference tax”?

It is Google Cloud’s label for the per-token overhead of running AI workloads on stacks that are not purpose-built for agentic AI. In the survey, 62% of leaders said they are seeing a significant inference tax, driven by data egress fees, storage bloat, and idle specialised hardware. The phrase sits alongside a separate finding that 81% of leaders cite operational complexity as a hidden cost of scaling AI.

Why is agentic AI harder on infrastructure than earlier AI?

A single prompt can trigger hundreds of downstream actions and requires massive context windows held in memory, the report argues. Continuous reasoning loops on legacy architecture are financially unsustainable, and 81% of leaders cite operational complexity as a hidden additional cost. The workloads differ from chatbot traffic in duration as much as in compute: an agent that runs for an hour on a single task pulls a different infrastructure profile than one that answers a single question.

How does the report frame governance?

As a precondition to scale. 79% of tech leaders cite security, governance, and MLOps as their top challenge to scaling inference. The push is toward centralised control planes with audit trails and human-in-the-loop approval for sensitive actions, which Google Cloud ties to its Agent Gateway product. The same shift shows up in the buyer-side numbers: 78% of organisations now source generative AI solutions from their primary cloud partner, up 30 points from 2025.

What does the report say about energy?

Power consumption is now a primary concern at the hardware purchasing stage. 91% of leaders factor it into selection, with 61% rating it as primary or significant. Specific regulatory limits cited include Germany requiring new data centres to achieve a Power Usage Effectiveness of 1.2 or lower, and Ireland requiring large data centres to provide 100% on-site dispatchable generation.

Where does Google Cloud say workloads should run?

Across a hybrid multicloud architecture. 52% of organisations now run on hybrid multicloud, and 48% are prioritising infrastructure with strict data residency controls. Edge deployment is also central: 90% of organisations rank edge deployment as important for AI, with 72% calling it extremely or very important.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending