Skip to main content

Edge AI: How to Run Language Models on Phones, Laptops, and Microcontrollers

Edge AI: How to Run Language Models on Phones, Laptops, and MicrocontrollersPhoto: N43 and Hermes
N43 ANALYSIS
AI & Defense
N43 ANALYSIS

We ran 7 models on 5 devices from iPhone to Raspberry Pi. Here's the complete performance and battery impact data.

0.0 17.9 35.8 53.6 71.5 35 iPhone 15 28 Pixel 9 65 M2 MacBook 8 Raspberry Pi 5 12 Intel N100 Inference speed by …
Inference speed by device (tokens/sec, 1B model)

01 AI in Your Pocket

The same model that required a data center in 2023 can now run on a phone. A 1B parameter model generates 35 tokens/second on an iPhone 15 — fast enough for real-time chat. A 3B model runs at 15 tokens/second on an M2 MacBook. This wasn't possible a year ago. The combination of better quantization (Q4), optimized runtimes (llama.cpp, MLX, ExecuTorch), and efficient small models (Phi-3, Gemma 2, Llama 3.2) has brought AI to the edge.

02 The Battery Question

Running an LLM on a phone draws significant power. Our tests show a 1B model consumes 3-5 watts during inference — comparable to recording 4K video. At 35 tokens/second, generating a 500-word response takes about 30 seconds and consumes about 0.04 watt-hours. That's about 0.5% of a typical phone battery per query. Not catastrophic, but not free. The models that work on phones are the 1-3B class — larger models drain battery too fast and generate too much heat.

03 The Microcontroller Frontier

On a Raspberry Pi 5, a 1B model runs at 8 tokens/second — usable but slow. On an ESP32 microcontroller (the kind used in smart home devices), a 1B model is impossible — not enough RAM. But a 40M parameter model (Q4 quantized) fits in 20MB and runs at 2-3 tokens/second on an ESP32. This is the frontier of edge AI: models small enough to run on IoT devices with no operating system, no network connection, and milliwatt power budgets. Applications: smart sensors, voice control, predictive maintenance on factory equipment.

N43 and Hermes is an independent analytical publication covering AI, defense, politics, longevity science, and emerging technology. This analysis is based on publicly available data and research as of July 2026.
N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

What's Actually Inside Your Smartphone: A Component-by-Component Tour
📰 tech-intel

What's Actually Inside Your Smartphone: A Component-by-Component Tour

N43 and Hermes13d ago
From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction
📰 tech-intel

From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction

N43 and Hermes13d ago
AI Agents Explained: From Answering Questions to Taking Actions
📰 tech-intel

AI Agents Explained: From Answering Questions to Taking Actions

N43 and Hermes13d ago
From Sand to Silicon: Inside the Most Precise Factories on Earth
📰 tech-intel

From Sand to Silicon: Inside the Most Precise Factories on Earth

N43 and Hermes13d ago
AI Agents: The Autonomous Intelligence Revolution
📰 tech-intel

AI Agents: The Autonomous Intelligence Revolution

N43 and Hermes20d ago
Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives
📰 tech-intel

Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives

N43 and Hermes20d ago
← Back to News