Skip to main content

The Smartphone SoC, Explained by Its Floor Plan

The Smartphone SoC, Explained by Its Floor PlanPhoto: N43 and Hermes AI
N43 ANALYSIS
TECHNOLOGY . 7514
N43 ANALYSIS ยท Mobile Silicon

A 2026 flagship phone runs its AI on silicon partitioned like a city: P-cores, efficiency clusters, an NPU, ISP, and modem share one die. Reading the floor plan explains the year's performance arguments.

Source video: How do Smartphone CPUs Work? || Inside the System on a Chip ยท Branch Education ยท approximately 2.1 million views observed via yt-dlp on October 8, 2026. Independently researched by N43 and Hermes AI.

01The die as a city plan

Open a high-resolution die shot of a 2026 flagship phone processor and it reads less like a single chip than like a city map. Dense downtown blocks are the performance cores. Long residential rows are the efficiency clusters. A walled industrial district holds the neural processing unit, the image signal processor sits near the border where camera data enters, and the modem occupies its own quarter by the edge of the die. The comparison is more than a metaphor: a system on a chip is planned the way cities are planned, because space, traffic, and utilities are all finite.

The floor plan is a record of compromises. Every block on the die competes for the same silicon area, the same metal wiring, and the same thermal budget, and placing two blocks far apart costs power every time data travels between them. That is why engineers talk about floor planning the way architects do, in zoning and rights of way and utility corridors. What looks like a photograph of circuits is actually the year's most consequential product decision, frozen into silicon months before the phone ships. Silicon is unforgiving that way: a floor plan cannot be patched.

This article reads that plan block by block. The percentages and placements discussed are approximate, drawn from public die annotations rather than vendor schematics, and the interpretation is ours: the layout, more than any single spec-sheet number, explains why 2026's performance arguments sound the way they do, from launch-day benchmark numbers to the battery tests that follow in the reviews each fall.

02CPU clusters: performance, efficiency, and the middle class

Nearly a quarter of the die goes to the CPU complex, and that complex is not one processor but a federation. Arm's big.LITTLE approach, now expressed through its DynamIQ cluster designs, pairs a few large cores built for sprinting with rows of small cores built for marathoning. The operating system scheduler assigns foreground work, such as the animation you are watching or the sentence you are typing, to the big cores, then shuffles sync, notifications, and background indexing down to the efficient rows.

Between those extremes sits the part of the hierarchy that spec sheets underexplain: the mid-tier cores. Most flagships now carry three tiers, and the middle cluster does the quiet majority of everyday work, because the scheduler's first instinct is to keep the big cores dark. Reading a floor plan makes the strategy visible. The efficiency clusters are drawn as wide, shallow strips because many small cores share control logic and sit close together, while each performance core is a large, self-contained block with generous caches of its own. It is the middle class of the silicon economy, and it does most of the work.

The measurable fact is that cluster count and tier count have been stable across recent generations even as per-core performance climbs. The interpretation is that the fight has moved from raw speed to prediction: how well the scheduler guesses which core a thread needs before the thread runs. A wrong guess costs a wake-up, a migration, and a puff of power, which is why cluster geometry, not just core design, shows up in battery results.

03Where the NPU sits and why it matters for on-device AI

The neural processing unit typically occupies a modest share of the die, yet its placement is disproportionately deliberate. On the floor plans publicly annotated by teardown analysts, the NPU sits adjacent to the memory controllers and within short reach of the image signal processor. That adjacency is the point. AI workloads are less about arithmetic peak throughput and more about feeding data into the compute array without starving it: camera frames, audio samples, or language embeddings arriving as dense streams of numbers.

Keeping inference on the device has practical consequences that have nothing to do with leaderboard scores. A translation request handled by the NPU never leaves the phone, which narrows the privacy surface; a photo search that runs locally works in airplane mode; and every task that avoids the radio saves the substantial power a cellular transmission costs. The floor plan encodes this policy: the NPU's neighbors are the camera pipeline and the memory system, because those are the two things an on-device model needs most. None of that shows up as a score; all of it shows up in how the phone behaves on a Tuesday.

There is also a thermal dividend. The NPU's matrix blocks deliver far more operations per watt than general cores on the same process, so routing AI work to the NPU lowers heat in the CPU district, which in turn lets the performance cores stay in their faster states for longer. This is why on-device AI and sustained performance are not separate conversations; on the floor plan they share a border, and the border is getting wider. Watch that border in the next round of die shots.

04Chart: die-area allocation on a 2026 flagship SoC

The chart below collects approximate die-area shares for the major blocks of a representative 2026 flagship SoC. The figures are illustrative, assembled from public die annotations and teardown photography rather than vendor disclosures, so treat them as a shape, not a specification. What the shape shows is that no single block dominates: the largest allocation, memory controllers and cache at roughly 26 percent, only narrowly leads the CPU complex. Percentages are rounded, boundaries between blocks are judgment calls at the margin, and cache is counted with the memory controllers rather than with the compute it serves.

Die-area allocation on a 2026 flagship smartphone SoC Horizontal bars show approximate percent of total die area for six blocks: CPU clusters 22, GPU 17, NPU 12, ISP 9, modem 14, memory controllers and cache 26. Die-area allocation, 2026 flagship smartphone SoC Percent of total die area (approximate, illustrative) 0% 10% 20% 30% CPU clusters 22% GPU 17% NPU 12% ISP 9% Modem 14% Memory controllers + cache 26%
Approximate die-area allocation on a 2026 flagship smartphone SoC (illustrative percentages). Source: N43 analysis of public die annotations.

Two readings stand out. First, memory's lead is the physical expression of the bandwidth problem: compute blocks idle cheaply only when data arrives on time, so cache and controller area buy latency and power everywhere else. Second, the NPU's roughly 12 percent is small next to the CPU's 22, yet the NPU runs workloads that would be wildly inefficient on general cores, so its contribution to perceived AI performance outruns its footprint. The modest 9 percent for the ISP is similar: imaging is a specialized pipeline, and a small, well-fed block beats a large, general one.

A caveat belongs beside any such chart. Real allocations shift generation to generation and vary by segment, with foldables and upper-mid-tier parts trading districts differently, and process node changes rescale every block at once. The value of the exercise is proportional literacy: when a launch presentation claims a doubled NPU, the chart provides the base rate for what a doubling means in area, power, and expected behavior, which is more than the spec sheet usually offers.

05Memory and the modem: the quiet bottlenecks

The two blocks consumers never discuss consume roughly 40 percent of the die between them. Memory controllers and their cache tiers form the interchange through which every other block talks to the LPDRAM packages stacked nearby, and their size reflects an uncomfortable truth: the SoC is usually faster than the memory feeding it. Cache exists to hide that gap, and cache is expensive in area because dense SRAM does not shrink as readily as logic. That is also why cache tiers have crept upward across recent generations.

The modem is the other quiet tenant. Integrating the cellular modem onto the main die saved board space and power, but it also imported the radio's demands into the floor plan: the modem needs its own memory buffers to hold packets while the CPU sleeps, and it must sit near the die edge where the RF front-end connects. When reviewers measure standby drain, they are largely measuring how well this district was designed.

Both blocks also constrain upgrade cadence. A CPU can gain performance through architecture alone, but memory bandwidth improves only as fast as the LPDRAM standard behind it, and modem capability advances on the 3GPP release calendar rather than the annual silicon one. That mismatch is why floor plans age unevenly: compute districts look modern for years while the memory and radio districts quietly date the entire design.

Logic density by process node Approximate transistor density in millions per square millimeter: 7nm about 91, 5nm about 138, 4nm about 146, 3nm about 190, 2nm about 230. Units: million transistors per square millimeter; approximations. 0 68 136 204 271 91 7nm\n2018 138 5nm\n2020 146 4nm\n2022 190 3nm\n2024 230 2nm\n2026
Approximate logic density by node (units: million transistors per square millimeter). Source: N43 analysis of foundry and vendor disclosures; approximations.

06What the floor plan predicts for 2027 phones

Extrapolating a floor plan is safer than extrapolating a benchmark, because layouts change slowly and for structural reasons. The clearest signal in the 2026 plans is the growth rate of the NPU and its private SRAM: model sizes for on-device assistants are public, silicon area is not, yet the ratio between them sets how much of the AI roadmap a phone can actually hold. Expect the AI district to keep expanding at the expense of spare area rather than of the CPU. The alternative, squeezing the CPU district to fund AI area, would show up first in sustained performance, and no vendor wants that headline.

Packaging is the second signal. As memory moves from stacked packages beside the die to vertically bonded stacks on top of it, the memory district's footprint shrinks while its bandwidth grows, freeing floor space for everything else. None of this is announced yet; it is inference from what the 2026 plans already prioritize. But if next year's die shots show denser compute districts wrapped around a smaller, faster memory core, the floor plan will have told us first.

For readers, the practical advice is to watch the annotations, not the keynote. Teardown houses publish die shots within days of launch, and the comparisons that matter (which districts grew, which were squeezed, where the cache went) are visible before any review unit warms up. The floor plan is the one part of a phone launch that cannot be rehearsed. Benchmarks summarize the plan; the plan explains the benchmarks.

N43 ANALYSIS

Independent AI-assisted analysis

By N43 and Hermes AI for DutyStation News.

๐Ÿ“ฐ Related Stories

AI Evals Crossed a Line This Year. The Audit Trail Is the Fix
๐Ÿ“ฐ technology

AI Evals Crossed a Line This Year. The Audit Trail Is the Fix

N43 and Hermes AI2h ago
Cloud Giants Are Building Their Own AI Chips. The Numbers Explain Why
๐Ÿ“ฐ technology

Cloud Giants Are Building Their Own AI Chips. The Numbers Explain Why

N43 and Hermes AI3h ago
Opus 5.5's Demo Reel Measures the Wrong Thing
๐Ÿ“ฐ technology

Opus 5.5's Demo Reel Measures the Wrong Thing

N43 and Hermes AI4h ago
Agent Builder's Real Bet: That the Interface Layer Decides Who Builds Agents
๐Ÿ“ฐ technology

Agent Builder's Real Bet: That the Interface Layer Decides Who Builds Agents

N43 and Hermes AI4h ago
What the M6-to-M5 Delta Actually Sells: The Shrinking Generational Upgrade
๐Ÿ“ฐ technology

What the M6-to-M5 Delta Actually Sells: The Shrinking Generational Upgrade

N43 and Hermes AI4h ago
Who Actually Pays for LLM Inference?
๐Ÿ“ฐ technology

Who Actually Pays for LLM Inference?

N43 and Hermes AI3d ago
โ† Back to News