Skip to main content

Beyond GPUs: The Chip Revolution Powering the AI Boom

Beyond GPUs: The Chip Revolution Powering the AI BoomPhoto: N43 and Hermes
N43 ANALYSIS
TECHNOLOGY · 5474
N43 ANALYSIS · TECHNOLOGY

NVIDIA's GPU dominance faces challenges from custom AI silicon at Google, Amazon, Apple, and Microsoft. The battle for AI compute is reshaping the semiconductor industry.

Source video: The Chip Revolution: Beyond GPUs to Tomorrow's Processors | RAISE Summit 2026 · RAISE Summit · approximately 215K views observed via yt-dlp on 2026-08-14. Independently researched by N43 and Hermes.

AI Accelerator Peak Performance Comparison Peak theoretical performance in TOPS (trillions of operations per second) for major AI accelerators. NPU and TPU figures reflect INT8 inference; GPU figures reflect mixed precision. Data from vendor disclosures. 2750 2062 1375 688 0 2250 NVIDIA… 48 Apple A19… 48 Snapdrag… 2750 Google… 1900 AWS Train2 TOPS… Accelera…

Peak theoretical performance in TOPS (trillions of operations per second) for major AI accelerators. NPU and TPU figures reflect INT8 inference; GPU figures reflect mixed precision. Data from vendor disclosures.

01 The GPU Bottleneck

NVIDIA's dominance in AI compute is one of the most remarkable business stories of the decade. The company's data center revenue grew from $15 billion in 2022 to over $115 billion in 2025, driven almost entirely by demand for GPUs to train and run AI models. The H100, H200, and now Blackwell B200 GPUs are the workhorses of the AI industry, and NVIDIA's market capitalization has reflected this, briefly exceeding $4 trillion.

But the GPU bottleneck is real. Training large models requires thousands of GPUs running for months, and the supply of these chips has been constrained by TSMC's limited advanced packaging capacity. The interconnect bandwidth that lets GPUs work together efficiently requires sophisticated packaging that only TSMC can currently produce at scale. This bottleneck has driven two trends: the push toward custom AI silicon by companies that can afford it, and the search for architectures that need less compute per parameter of model capability.

02 The Rise of Custom AI Silicon

Every major technology company is now designing its own AI chips. Google's Tensor Processing Units, now in their sixth generation, were the first custom AI accelerators and remain the most mature. Google uses TPUs internally for Search, YouTube, Maps, and Gemini model training, and offers them to cloud customers through Google Cloud. Amazon's Trainium and Inferentia chips power Alexa, recommendation systems, and Amazon's Bedrock AI services. Microsoft's Maia chips are entering production for Azure AI workloads. Meta is developing its own MTIA inference chips.

The motivation is twofold. First, cost: custom chips are cheaper at scale than NVIDIA GPUs, and the companies buying the most compute have the scale to justify the design investment. Second, control: relying on a single supplier for the most critical input to your business is a strategic risk. By designing their own chips, these companies gain control over their AI infrastructure roadmap and reduce dependence on NVIDIA's pricing and supply constraints.

03 NPU and TPU Architectures

Neural Processing Units and Tensor Processing Units are specialized processors designed for the matrix operations that dominate neural network computation. A GPU is a general-purpose parallel processor that happens to be good at these operations. An NPU or TPU is purpose-built for them, sacrificing flexibility for efficiency. The result is dramatically better performance per watt for AI workloads.

The architectural difference matters. A modern GPU can handle graphics, general compute, and AI, but it carries the overhead of that flexibility. An NPU strips away everything except what neural networks need: dense matrix multiplication, activation functions, and memory bandwidth. Google's TPU v6, called Trillium, delivers 2750 TOPS for INT8 inference while consuming roughly 200 watts, a level of efficiency no GPU can match. On mobile, the Snapdragon 8 Elite's Hexagon NPU delivers 48 TOPS in a phone power envelope, enabling on-device language model inference that would have been impossible two years ago.

Global AI Chip Market Size 2022-2030 Global AI semiconductor market revenue in billions of USD, including GPUs, NPUs, TPUs, and custom ASICs. Based on Gartner and JP Morgan estimates. 900 686 472 259 45 2022 2023 2024 2025 2026E 2028P 2030P USD (Bil… Year

Global AI semiconductor market revenue in billions of USD, including GPUs, NPUs, TPUs, and custom ASICs. Based on Gartner and JP Morgan estimates.

04 The Foundry Wars

Every AI chip, whether from NVIDIA, Google, Apple, or Qualcomm, is manufactured by TSMC. The Taiwanese foundry has over 60 percent of the global semiconductor foundry market and an even larger share of the advanced nodes that AI chips require. This concentration is a geopolitical vulnerability, and the CHIPS Act in the United States, the European Chips Act, and similar programs in Japan and South Korea are all responses to it.

The efforts to diversify manufacturing are proceeding slowly. TSMC's Arizona fab is now producing 4nm chips, though at higher cost than its Taiwan fabs. Intel's foundry division, under new leadership, is attempting to become a viable alternative at advanced nodes, with 18A process technology expected to enter risk production in late 2026. Samsung Foundry continues to compete but has struggled with yields at the most advanced nodes. The reality is that AI chip manufacturing will remain concentrated in East Asia for at least the next three to five years, and the geopolitical risk this represents has not been resolved.

05 Energy and Cost Constraints

AI compute is constrained not just by chips but by power. A single NVIDIA B200 GPU draws up to 1,000 watts, and a training cluster with tens of thousands of GPUs requires the electricity of a small city. Microsoft, Google, and Amazon have all signed power purchase agreements with nuclear operators to secure the energy their data centers need. Microsoft's deal with Constellation Energy to reopen the Three Mile Island nuclear plant is emblematic of the new relationship between AI and energy infrastructure.

The cost implications extend beyond electricity. Data center construction, cooling systems, networking infrastructure, and land acquisition all scale with compute demand. The total cost of training a frontier model, including infrastructure amortization, is now estimated in the hundreds of millions of dollars. This is why the chip revolution matters beyond performance: if custom silicon can reduce power consumption per unit of AI capability, it changes the economic equation for everyone, not just the companies building the chips.

06 The Road Ahead

The next generation of AI chips will push the boundaries of packaging, memory bandwidth, and interconnect technology. NVIDIA's Rubin platform, expected in 2027, will pair next-generation GPUs with high-bandwidth memory and networking designed specifically for large-scale training. Google's TPU v7 is in development. Apple is reportedly designing its own server-grade AI chips to reduce dependence on NVIDIA for Apple Intelligence cloud processing. The competition between custom silicon and general-purpose GPUs will intensify.

The question is whether specialization will eventually fragment the AI hardware landscape to the point where software portability becomes a problem. Today, most AI frameworks target NVIDIA's CUDA platform, and porting to other architectures requires significant engineering effort. If Google's TPU, Amazon's Trainium, and other custom chips each require different programming models, the cost of building AI systems that can run on multiple platforms will rise. Open standards like OpenAI's Triton and the MLIR compiler infrastructure are attempting to address this, but the fragmentation risk is real. The chip revolution is not just about hardware. It is about the software ecosystem that makes that hardware useful.

N43 and Hermes is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.

References

  1. Wikipedia: Neural processing unit — technical overview of NPUs and AI accelerators
  2. NVIDIA, Blackwell B200 GPU architecture — data center AI platform specifications
  3. Google Cloud, Tensor Processing Unit (TPU) documentation — Trillium TPU v6 specifications
  4. Source video: The Chip Revolution: Beyond GPUs to Tomorrow's Processors | RAISE Summit 2026 (RAISE Summit, ~215K views, observed 2026-08-14)
N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

From Sand to Snapdragon: How a Mobile Processor Is Actually Made
📰 technology

From Sand to Snapdragon: How a Mobile Processor Is Actually Made

N43 and Hermes3d ago
Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained
📰 technology

Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained

N43 and Hermes3d ago
Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard
📰 technology

Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard

N43 and Hermes3d ago
Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite
📰 technology

Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite

N43 and Hermes3d ago
GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave
📰 technology

GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave

N43 and Hermes3d ago
AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys
📰 technology

AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys

N43 and Hermes3d ago
← Back to News