Skip to main content

Cloud Giants Are Building Their Own AI Chips. The Numbers Explain Why

Cloud Giants Are Building Their Own AI Chips. The Numbers Explain WhyPhoto: N43 and Hermes AI
N43 ANALYSIS
TECHNOLOGY . 7515
N43 ANALYSIS · AI Infrastructure

AWS, Google, and Microsoft train and serve on homegrown silicon alongside Nvidia. The economics of scale - not nationalism - drive the shift, and the accounting is more mundane than the rhetoric.

Source video: How Nvidia GPUs Compare To Google’s And Amazon’s AI Chips · CNBC · approximately 2.1 million views observed via yt-dlp on October 8, 2026. Independently researched by N43 and Hermes AI.

01The merchant-chip premium

Buy enough of something and the interesting question stops being the price tag and becomes the margin inside it. That is the quiet logic behind every custom AI chip program at AWS, Google, and Microsoft. Merchant GPUs from Nvidia carry a premium that reflects their scarcity, their generality, and the decades of software built around them. At hyperscale, that premium multiplies across hundreds of thousands of accelerators, and the arithmetic starts to argue for building the same function in-house.

The trend is measurable. Google has deployed successive generations of its TPU since 2015 for both training and inference. AWS has iterated Trainium for training and Inferentia for serving. Microsoft has brought its Maia accelerators into its datacenters. None of these programs has displaced the merchant GPU; each runs alongside it, absorbing workloads whose shapes suit a fixed function. The claim worth examining is not that custom chips replace Nvidia, but that they change the marginal economics of the fleet.

Interpretation matters here, because the public record is thinner than the rhetoric. Earnings letters discuss capex categories, not unit costs per accelerator, and pricing pages change faster than procurement contracts. What can be said plainly: the largest cloud operators now design their own AI silicon, they publish enough about utilization and depreciation for outside analysts to model the trade, and the direction of travel has held across several hardware generations. The rest of this article reads the numbers that are actually public.

02What a TPU actually gives up to gain speed

A tensor processing unit is not a general GPU with the logo filed off. It is an application-specific integrated circuit arranged around one idea: dense matrix multiplication. Google's TPUs use systolic arrays, grids of multiply-accumulate units that pulse data through in waves, which raises throughput per watt for transformer workloads. What gets given up is breadth. A TPU has no rasterizer for graphics, little use for the precision ranges scientific computing expects, and no interest in workloads that do not reduce to linear algebra at scale.

That narrowness is a feature when the workload cooperates. Training and serving large models are overwhelmingly matrix operations, so the silicon the workload does not need can be spent on memory bandwidth and on-chip buffers instead. The measured consequence in public benchmarks and papers is favorable performance per watt on transformer-shaped problems. The unmeasured cost appears at the edges: new operators must be mapped by hand or by compiler, and a workload that falls outside the array's happy path can leave expensive silicon idling.

Amazon's Trainium makes the same bargain with different emphasis, and Microsoft's Maia follows a similar logic. The pattern across all three is consistent: trade the general-purpose frills for area that serves the dominant operation, then depend on compilers to keep the array fed. Whether the bargain pays depends less on the chip than on the software stack that decides what runs on it, which is the subject the marketing tends to skip and the next section takes up.

03Software moats cut both ways

Nvidia's real moat is CUDA, the software layer through which nearly two decades of machine-learning code has been written. Researchers write for it first, frameworks optimize for it first, and every driver release carries forward an enormous inheritance of compatibility. A custom accelerator must offer a substitute path: Google leans on XLA and the JAX and PyTorch ecosystems compiled for TPU, AWS on its Neuron SDK, Microsoft on its own toolchain. The substitutability of that software is the true test of every custom chip.

The moat cuts both ways, and this is the part of the story most coverage inverts. CUDA's breadth is a liability at the margin for hyperscalers, because they pay for generality they do not use across enormous fleets, and every idle feature still occupies engineering attention, driver cycles, and validation time. Meanwhile a custom stack, once built, binds workloads to in-house silicon, which lowers the exit cost calculation for the next generation of homegrown chips. The switching cost that protects the merchant vendor in one direction protects the captive stack in the other.

Framework-level evidence supports the convergence: PyTorch and JAX now target multiple backends as a design goal rather than an afterthought, which lowers the software tax on every new accelerator. Interpretation: the harder moat to cross is no longer the compiler but the cluster interconnect and the collective-communication libraries tuned around it. Networking is where merchant silicon retains its strongest grip, and it is the reason even aggressive custom programs keep buying GPUs for the training runs that cannot tolerate a slow network.

04Chart: cost per trained token, merchant vs. custom

The economics reduce to a single ratio: what a unit of useful compute costs on each kind of silicon. Public comparisons are hard, because cloud pricing bundles the chip with memory, networking, and the vendor's margin, and earnings letters discuss utilization in percentages rather than dollars. The chart below therefore states the comparison the way procurement actually experiences it, as an index with the merchant GPU set to 100, and labels the numbers illustrative: they encode the direction and rough magnitude of the published cost arguments, not audited line items. Treat them as a map of the argument rather than an invoice.

Relative cost per unit of AI compute: merchant GPU vs. custom ASIC Grouped bars show indexed cost where merchant GPUs sit at 100 in all three groups and custom ASICs at 70 for training, 55 for batch inference, and 85 for real-time inference; lower means cheaper. Relative cost per unit of AI compute (indexed) Merchant GPU = 100 · lower is cheaper · illustrative Merchant GPU Custom ASIC 100 50 0 100 70 100 55 100 85 Training Batch inference Real-time inference
Illustrative relative cost per unit of AI compute: merchant GPU vs. custom ASIC (indexed). Source: N43 analysis of public cloud pricing and earnings disclosures.

Read the groups left to right and the shape of the argument appears. Training shows the smallest gap, index 70 for custom against the merchant baseline of 100, because training stresses the network and the software stack as much as the chip. Batch inference shows the widest, 55, which is the case custom ASICs were built for: predictable, dense, parallel work with no interactive latency requirement. Real-time inference narrows again to 85, because serving live traffic punishes any rigidity in scheduling.

Two caveats keep the chart honest. The indices are illustrative, drawn from the pattern of public pricing and earnings discussion rather than disclosed figures, and they describe relative cost per unit of compute, not per model or per customer. Even taken at face value they explain behavior better than slogans do: hyperscalers route the batch workloads to homegrown silicon first, keep latency-sensitive and novel workloads on merchant GPUs longest, and let the mix, not the manifesto, set the ratio.

05Capacity accounting: why the capex surge still pencils out

The capex numbers look alarming in isolation. The largest cloud companies have raised capital expenditures to historic levels, much of it datacenter shells, power infrastructure, and accelerators bought years ahead of the revenue they will eventually support. But an accelerator is not an expense in the year it is installed; it is depreciated over a schedule the company chooses, commonly five to six years in recent disclosures. That schedule is a genuine judgment call, and it is where skepticism belongs, because GPUs lose economic relevance faster than buildings if the hardware cycle keeps shortening.

Utilization is the other variable that decides whether the spend pencils out, and it is the one management letters emphasize generically: expensive silicon only earns its depreciation if something is scheduled on it. Custom chips change this calculus in a specific way. Because they are built for the owner's own workloads, they can be ordered against internal demand forecasts rather than against a spot market, which is why the public earnings letters treat the homegrown programs as margin stories rather than as science projects.

The mundane accounting is therefore the real story. Whether the capex surge is prudent depends on three quasi-published quantities: the depreciation schedule assigned to each fleet, the utilization achieved on it, and the revenue the services running on it can carry. None of these requires nationalism or a chip war to explain. They require only that a company large enough to amortize a design program across a gigantic fleet will find the per-unit economics of doing so persuasive at some price point.

Hyperscaler capital expenditure plans, 2026 Planned 2026 capital expenditure in billions of US dollars: Amazon about 125, Microsoft about 120, Alphabet about 93, Meta about 72. Units: billions of dollars; company guidance approximations. 0 37 74 111 148 125 Amazon 120 Microsoft 93 Alphabet 72 Meta
Planned 2026 capital expenditure (units: $ billions, company guidance approximations). Source: N43 analysis of public earnings guidance.

06What a two-silicon market means for startups

For startups building on cloud AI, the two-silicon market is mostly good news with one catch. Inference prices have drifted down as custom capacity comes online, because batch serving on in-house ASICs is the cheapest capacity the clouds operate, and competition passes some of that through. Training on cutting-edge merchant silicon remains the expensive tier. The practical consequence: the cost curve favors teams that design products around cheap, predictable inference rather than around ever-larger training runs.

The catch is portability. Workloads tuned to one vendor's accelerator stack, however open the frameworks claim to be, acquire runtime assumptions that are not free to unwind: scheduling quirks, precision behavior, and performance quirks that quietly become requirements. Teams that keep their serving layer hardware-agnostic preserve leverage over both silicon camps; teams that lean into one stack trade that leverage for performance today. Neither choice is wrong, but the two-silicon market makes the choice explicit in a way the single-vendor era never did.

The number to watch next is not a benchmark but a ratio: custom capacity as a share of each cloud's announced accelerator fleet, visible obliquely in earnings letters and supply-chain reporting. If that share keeps climbing while inference prices keep falling, the custom programs are working as designed. If it stalls, the software moat will have proven deeper than the margin. Either outcome settles an argument that, so far, the industry has been content to leave rhetorical.

N43 ANALYSIS

Independent AI-assisted analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

AI Evals Crossed a Line This Year. The Audit Trail Is the Fix
📰 technology

AI Evals Crossed a Line This Year. The Audit Trail Is the Fix

N43 and Hermes AI2h ago
The Smartphone SoC, Explained by Its Floor Plan
📰 technology

The Smartphone SoC, Explained by Its Floor Plan

N43 and Hermes AI3h ago
Opus 5.5's Demo Reel Measures the Wrong Thing
📰 technology

Opus 5.5's Demo Reel Measures the Wrong Thing

N43 and Hermes AI4h ago
Agent Builder's Real Bet: That the Interface Layer Decides Who Builds Agents
📰 technology

Agent Builder's Real Bet: That the Interface Layer Decides Who Builds Agents

N43 and Hermes AI4h ago
What the M6-to-M5 Delta Actually Sells: The Shrinking Generational Upgrade
📰 technology

What the M6-to-M5 Delta Actually Sells: The Shrinking Generational Upgrade

N43 and Hermes AI4h ago
Who Actually Pays for LLM Inference?
📰 technology

Who Actually Pays for LLM Inference?

N43 and Hermes AI3d ago
← Back to News