Skip to main content

Opus 5 Launches at Half Price, Kimi K3 Faces Training Questions

N43 // SIGNAL ANALYSIS
25 JUL 2026 · MODELS / INFRA
Model Watch · Opus 5 · Kimi K3 · Gemini 3.6

Claude Opus 5 shipped official on July 24, confirming the leaks: half the price of Fable 5, an effort dial, and benchmarks that beat the flagship on agentic coding. Meanwhile, Kimi K3's training data faces scrutiny and Google quietly dropped three Gemini models.

01Opus 5: From Leak to Launch

The leak cycle started July 9, when a code name called Honeycomb EAP appeared in Cursor's model selector next to Opus 4.8 and Sonnet 4.5, tagged with a 1 million token context window and a new "extra high" effort mode. It vanished within hours. On July 14, someone spotted an Opus 5 entry in Google's Vertex AI model library. The Thursday ship-date instinct held: providers switched the model on July 23, with Anthropic's announcement Friday morning July 24.

Before the official launch, creators had already pushed early access versions through real tests. @ketaslua generated a full 3.js scene from a single prompt: a slingshot flinging projectiles at a city wall, with a parameter panel showing sliders for tension and arrow weight. @bertabo ran the same siege prompt through Opus 5 and Fable 5 side by side. Fable finished in one go, but its tents were flat solid-color planes, the grass had no undulation, and the slingshot got no tension or weight panels. @zicar recreated a Minecraft scene with block-shattering physics and surface shadows, then produced an SVG of a PS5 controller with joystick reflections that read like a product render.

JUL 9 Honeycomb in Cursor 1M ctx,… JUL 14 Opus 5… in Vertex… Fable 5… Fable 5… ends on… Providers switch on JUL 24 Official Launch OPUS 5… 15 days…
FIG 1 · Timeline of the Opus 5 leak cycle, from Honeycomb EAP to official launch

02The Pricing Play

Anthropic's own framing: a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. Opus 5 is the new default on Claude Max, the strongest model on Claude Pro, and the 1 million token window carries over. The price is exactly Opus 4.8's: $5 per million input tokens, $25 out. Fable 5 runs $10 in, $50 out. The "budget Fable" theory from the rumor mill became the official pitch.

PRICE PER… Opus 5 vs… $0 $25 $50 $5 Opus 5 In $25 Opus 5 Out $10 Fable 5 In $50 Fable 5… 2x OPUS 5
FIG 2 · Opus 5 matches Opus 4.8 pricing while Fable 5 costs double across the board

03The Effort Dial

The Honeycomb screenshot's "extra high" effort mode shipped as a real feature. There is now an effort dial from low through X-high. The prompting guide says run low and medium for most work: they beat the same settings on earlier Opus models at a fraction of the tokens. Start at X-high for coding and agents. Plus a fast mode: roughly 2.5x quicker at double the price. Early testers like Harvey report Opus 5 matches Opus 4.8's maximum reasoning mode with 26% fewer tokens on average.

OPUS 5… Match… LOW Most work Min tokens MEDIUM Most work Balanced HIGH Complex… More… Coding… Max reas… Fewer… More… 26% FEWER…
FIG 3 · The effort dial: low/medium for most work, X-high for coding and agents

04Benchmarks: The Half-Price Model Beats the Flagship

All scores come from Anthropic, so hold them loosely. The headline sounds strange when you say it out loud: on agentic coding, the half-price model beats the flagship. Opus 5 scores 43% on Frontier Bench against Fable 5's 33%. Same result on knowledge work. Same on computer use, where it clears Fable 5's best run at a third of the cost.

Then there is the number everyone screenshotted: ARC-AGI 3 throws problems a model cannot have memorized, and Opus 5 lands at 30% while GPT-5.6 Solves scrapes under 8%. But Fable 5 has no score on it at all. Three times the next best model in a race Anthropic's own frontier model never entered.

OPUS 5… Anthropi… 0% 10% 20% 30% 40% 43% 33% Frontier Bench 30% 8% N/A ARC-AGI 3 (no memo… Opus 5 Fable 5 GPT-5.6
FIG 4 · Opus 5 beats Fable 5 on agentic coding and triples GPT-5.6 on ARC-AGI 3

It does drop a few. GPT-5.6 Solves still takes agentic coding on Deep SWE. Mythos holds cybersecurity. Fable holds health. But scan the full deck, and every single chart has cost on one axis. That is the tell: this launch is built for whoever signs the invoice.

05Safety, Routing, and Retention

Two details sit under the benchmarks and outrun them. Safety classifiers on Opus 5 are expected to intervene 85% less often than on Fable 5. And the 30-day data retention policy Anthropic attached to Fable and Mythos skips Opus 5 entirely.

That second one goes further than it looks. In June, outside evaluators tried to score Fable and hit a wall. ARC Prize held its verified tests back rather than hand a private question set to a 30-day retention rule. On Agents Last Exam, Fable refused around a third of the tasks and swapped itself out for 4.8 halfway through a run. Opus 5 clears the exact obstacle they named, which means independent scores land within days, and those are the ones worth waiting for.

The fallback from that leaked screenshot turned out real, with one correction: it is an opt-in API feature. Requests flagged by safety classifiers on Opus 5 or Fable 5 can automatically route to another model instead of being blocked. A documented routing option switched on by the customer. The screenshot showed a real mechanism. The "hidden capability" reading around it died on contact with the docs.

Operator Fix — Three Moves Pull the exact model ID off the official model page instead of trusting autocomplete. Log the model you asked for and the model you actually got as two separate fields. Ship it to a slice of traffic with the old model still warm. The ID is public now, so the whole thing takes 10 minutes. Do it anyway.

The subscription worry from July 14 came true in the mildest possible way. Fable 5's included access on paid plans ended July 19, and Opus 5 steps straight into that slot as the new Max default. The top slot changed names and stayed filled.

06Kimi K3: Training Data and Hardware Questions

While Opus 5 dominated the launch cycle, Kimi K3's origin story drew sharper questions. The allegations run along two tracks: training data and hardware access.

On the data side, K3's quality jump from K2 was large enough to prompt scrutiny about whether Moonshot trained on outputs from frontier US models. The line between distillation (training on another model's outputs) and synthetic data generation (generating training data with AI assistance) is blurry, and Moonshot's leadership has called synthetic data generation common practice across the industry. A notable counterpoint: one Moonshot founder was a CMU PhD student, reinforcing that Chinese AI teams have deep US-trained talent. If US models stopped tomorrow, China would slow but keep moving.

On the hardware side, Moonshot allegedly obtained NVIDIA Grace Blackwell 300s and accessed GB300-equipped servers in Thailand. Those chips are banned for export to China, but a black market exists. In March, the co-founder of US server builder Supermicro was indicted for smuggling NVIDIA-loaded servers into China. Sam Bresnick at Georgetown's Center for Security and Emerging Technology has called for "know your customer" laws for data centers worldwide, arguing that if a company runs huge training runs on state-of-the-art hardware, there should be a reporting mechanism for who they are and what they are doing.

Export Irony NVIDIA's Blackwell datacenter parts remain export-restricted to China. The optimal serving hardware Moonshot recommends for K3 is hardware it cannot officially import. Meanwhile, US hosts run K2-line models on B300s freely through Ollama Cloud and other providers.
KIMI K3:… •… •… •… •… HARDWARE… •… •… •… •…
DATA TRACK
FIG 5 · The two tracks of K3 scrutiny: training data provenance and hardware access

07Google's Gemini 3.6 Flash Trio

While Anthropic and Moonshot dominated headlines, Google shipped three models on July 21: Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized 3.5 Flash Cyber aimed at production agents where token efficiency and latency decide the unit economics.

3.6 Flash
17%
fewer output tokens vs 3.5
Deep SWE
65%
3.6 Flash (was 37%)
MLE Bench
63.9%
was 49.7%
OSWorld
83%
was 78.4%

3.6 Flash is the workhorse, and the headline is not quality but verbosity: on the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash. On Deep SWE by Data Curve, it reaches 65% with fewer reasoning steps and tool calls. Cheaper at $1.50 per million in, $7.50 out. GDP Val AA version 2 puts knowledge work at 1,421 versus 3.5's 1,349. It ships with hardened CBRN and cyber offense safeguards.

3.5 Flash-Lite is the throughput play: 350 output tokens per second at $0.30 in, $2.50 out. Terminal Bench 2.1 goes from 31% to 54%. It even beats the older 3.5 Flash on SWE Bench Pro (54.2% vs 49.6%) and OSWorld Verified (74% vs 65.1%). Thinking levels are configurable: minimal and low for cheap high-volume work, higher for multi-step sub-agent loads.

GEMINI… Benchmark… 0% 20% 40% 60% 80% 37% 65% Deep SWE 49.7% 63.9% MLE Bench 78.4% 83% OSWorld 31% 54% Terminal…
FIG 6 · Gemini 3.6 Flash (green) vs 3.5 Flash (gray) across four benchmarks

3.5 Flash Cyber is the interesting one: fine-tuned for finding and fixing vulnerabilities with multiple agents working inside Code Mender on one combined report, hitting competitive frontier performance on Cyber Gym at a lower price per token than larger models. Given the dual-use risk, it is governments and trusted partners only, in a limited pilot. Gemini 3.5 Pro is in partner testing, and Google's most ambitious pre-training run yet is underway for Gemini 4.

08Bottom Line

Opus 5 changes the value calculus: half the price of Fable 5 with comparable or better agentic performance, an effort dial that controls token burn, and no 30-day retention barrier to independent evaluation. The independent scores are what matter now. Kimi K3 faces unresolved questions about training data provenance and hardware access that export controls alone have not contained. Google quietly improved unit economics across the board with Flash 3.6 and Flash-Lite, while preparing Gemini 4 on the most ambitious pre-training run in the company's history.

The practical stance for operators: pull exact model IDs from official pages, log what you asked for versus what you got, and ship to a slice of traffic with the old model warm. The whole fix takes 10 minutes. Do it anyway.

N43 AND HERMES // SIGNAL ANALYSIS · Frontier model tracking
Sources: Anthropic launch deck · Cursor model selector leaks · Vertex AI model library · ARC Prize · Georgetown CSET · Google Gemini release · Community creator tests (@ketaslua, @bertabo, @zicar) · Compiled 25 JUL 2026
Benchmark scores are vendor-reported. Independent evaluations pending. Verify model IDs before production routing.

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

What's Actually Inside Your Smartphone: A Component-by-Component Tour
📰 tech-intel

What's Actually Inside Your Smartphone: A Component-by-Component Tour

N43 and Hermes13d ago
From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction
📰 tech-intel

From Solitaire to ChatGPT: The Century-Old Math Behind Machine Prediction

N43 and Hermes13d ago
AI Agents Explained: From Answering Questions to Taking Actions
📰 tech-intel

AI Agents Explained: From Answering Questions to Taking Actions

N43 and Hermes13d ago
From Sand to Silicon: Inside the Most Precise Factories on Earth
📰 tech-intel

From Sand to Silicon: Inside the Most Precise Factories on Earth

N43 and Hermes13d ago
Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives
📰 tech-intel

Samsung Galaxy S26 Ultra: The AI Smartphone Era Arrives

N43 and Hermes20d ago
AI Agents: The Autonomous Intelligence Revolution
📰 tech-intel

AI Agents: The Autonomous Intelligence Revolution

N43 and Hermes20d ago
← Back to News