Price series · Issue 1 · August 12, 2026
The Model Price War: What DeepSeek's Low Prices and Claude's Permanent Cut Change
Model capability sets the ceiling; token pricing determines how often products can place that capability inside a user action. This is not a quality-free leaderboard. It tracks how list prices, caching, long-context tiers and promotions reshape AI app economics.
This week's signal: Claude made Sonnet 5's lower launch price permanent
Claude Sonnet 5 costs $2 input / $10 output versus Sonnet 4.6 at $3 / $15, a 33.3% per-token cut. On August 10, 2026, Anthropic updated the launch post to make the price permanent instead of ending it on August 31. The new tokenizer may still produce roughly 1.0–1.35× as many tokens for the same text, so the headline cut is not automatically the same reduction on every workload bill.
Current public API list prices
USD per 1M tokens. Standard input, cache-hit input and output are separated. Batch, tools, search, storage, regional premiums, taxes and enterprise discounts are excluded. The two groups are purchasing scenarios, not claims of equal capability.
| Provider | Model | Standard input | Cache hit | Output | Status | Scope note |
|---|---|---|---|---|---|---|
| DeepSeek | DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 | GA | Cache-miss input; cache-hit price shown separately. |
| OpenAI | GPT-5.6 Luna | $0.2 | $0.02 | $1.20 | GA | Requests above 272K input tokens use higher long-context rates. |
| Gemini 3.5 Flash-Lite | $0.3 | $0.03 | $2.50 | GA | Paid standard tier; output includes thinking tokens. | |
| Anthropic | Claude Haiku 4.5 | $1 | $0.1 | $5 | GA | First-party global Claude API; cache-hit price shown separately. |
| DeepSeek | DeepSeek V4 Pro | $0.435 | $0.003625 | $0.87 | GA | Cache-miss input; official API supports thinking and non-thinking modes. |
| Anthropic | Claude Sonnet 5 | $2 | $0.2 | $10 | new-permanent | Anthropic made the $2 input / $10 output launch price permanent on Aug 10, 2026. |
| OpenAI | GPT-5.6 Terra | $2 | $0.2 | $12 | GA | Requests above 272K input tokens use 2x input and 1.5x output rates. |
| Gemini 3.1 Pro Preview | $2 | $0.2 | $12 | preview | Paid standard tier for prompts up to 200K; higher long-context rates apply. |
DeepSeek is not slightly cheaper; its list price sits in another order of magnitude
DeepSeek lists V4 Pro at $0.435 input and $0.87 output, and V4 Flash at $0.14 input and $0.28 output. Cache-hit prices fall to $0.003625 and $0.0028. On public token prices alone, these are far below the other products in the same purchasing groups on this page.
Low price does not prove the best economics for every task. Buyers still need task success rate, retries, latency, availability, context use and tool calls. This series keeps price-list facts separate from workload performance judgments.
Claude shows the price war has reached frontier work, not only small models
Sonnet 5 permanently cuts per-token standard input and output by one third and cache hits to $0.20. This is not a company-wide cut, but it is a clear signal that frontier workloads now compete on durable unit cost as well as capability.
The real checkpoint is how much actual bills fall under the new tokenizer and whether DeepSeek, OpenAI and Google answer through standard, cache or batch pricing.
Which apps feel the pressure?
- Thin model wrappers: Cheaper inference compresses resale margins and differentiation. Workflow, data, distribution or verified outcomes must carry the product.
- AI gateways and routers: Pure price arbitrage weakens; reliability, governance, evaluation, regional compliance and failover become the durable value.
- Traditional SaaS AI add-ons: Falling model cost weakens the rationale for large AI seat premiums and pushes pricing toward outcomes, allowances or complete workflows.
- Vertical agents and automation apps: Cheaper reasoning makes long task chains, retries and background execution more viable—and easier for more products to embed.
How we compare
We record first-party public prices only. The main table uses USD per 1M text tokens and separates standard input, cache hits and output. Free tiers, subscriptions, batch, media, tools, search and negotiated discounts stay outside it.
Models are not equal commodities. This table tracks cost curves and pricing moves; it does not independently establish quality or total bills. Context bands, tokenizers, reasoning tokens, cache writes and regional premiums can all change effective cost.
What counts as a major change
Official pages are scanned weekly. Any trigger opens a special-edition review, published only after verification:
- At least a 20% change to standard input or output for a tracked model;
- A durable new low-price benchmark in a general/frontier tier, or a promotion expiring or extending;
- Cache, batch or free-tier changes moving a representative workload by at least 30%;
- Tokenizer, long-context or tool-fee changes that reverse the apparent bill direction.
Issue 1 | DeepSeek sets a low-price anchor; Claude makes the lower price permanent
August 12, 2026: we establish a baseline across four providers and eight models. The first event is Anthropic's August 10 decision to make Sonnet 5's 33.3% lower per-token price permanent. DeepSeek V4 Pro and Flash set the lowest public standard and cache-hit prices in this table. The next checkpoint is competitive response and actual workload bills.
Official sources
Prices change. Preserve the verification date when citing this page and confirm against the provider's original page.
- DeepSeek API Models & Pricing — DeepSeek
- Claude Platform model pricing — Anthropic
- Introducing Claude Sonnet 5 — Anthropic
- GPT-5.6 Terra model page — OpenAI
- GPT-5.6 Luna model page — OpenAI
- Gemini Developer API pricing — Google
Model specials series
- Claude Academy: Who Bears the Cost of AI Fluency? — Anthropic has launched a free school for learning to work with AI. Its stated framework reaches beyond prompts into delegation, judgment and disclosure. That is a meaningful public resource—and it raises a harder question: when the maker of the disruption also issues the credentials for adapting to it, where does institutional responsibility end and individual responsibility begin?
- Meituan All-in AI: The Execution Costs — Going all-in on AI is easy to announce and hard to govern. The execution bill arrives where strategic urgency meets source provenance, merchant consent and incentives: the less time a team leaves for verification and reversal, the more expensive its speed becomes. This special separates the verified public record from two weak, single-source signals and treats the pattern as a governance problem—not proof that AI investment itself has failed.
- AI Token Monetization: Token Is the New Dollar — At Stripe Sessions, President of Technology and Business Will Gaybrick changed a demo app from a $2 flat fee to $3 per million tokens, then streamed stablecoin payments as each token was consumed. That sequence is more than a billing demo. It shows software moving from seats and monthly access toward metered intelligence: every unit of model work can carry a price, a margin, a fraud risk and a settlement event. Our thesis is that the token is becoming the dollar of AI software—a unit of account for machine work, not legal tender and not a replacement for the US dollar.
- Stripe × OpenRouter: Where Token Monetization Begins — Put the reported $8 billion price aside. Stripe, the payments leader that became a checkout layer for the internet, chose not to buy a model lab but the switchboard between AI applications and hundreds of models. That is the story. Stripe is betting that the most important layer of the AI economy will not only produce intelligence; it will turn intelligence into exchangeable value. Today an LLM token is a billing unit. Tomorrow a model-agnostic AI token could be held, transferred and settled like the generation of digital assets opened by BTC and ETH — representing not digital scarcity or blockchain gas, but a claim on usable intelligence. Axios reports an agreement above $8 billion, while neither company had publicly confirmed it at this update.
- Grok Bot: xAI Gives Every Agent Its Own Computer — Launched in early beta on August 11, 2026, Grok Bot turns the agent from a chat window into a teammate: each Bot gets a persistent cloud computer, signs into the tools you already use and keeps working while you are away. We have been using it since day one. The product is days old and public information is still thin, so this special leads with tested impressions and keeps mechanism facts second — separating what is verified, what is company-stated and what is our read.
- Kimi: Moonshot AI's Open-Weight Sprint to the Frontier — In four months Moonshot AI shipped an open-weight 1-trillion-parameter workhorse, followed it with the reported 2.8-trillion-parameter Kimi K3, paused new paid consumer subscriptions when demand outran capacity, and set off toward a Hong Kong listing. This file collects what is sourced, what is company-reported, and our read on where it fits the timeline this site tracks.
- Zhipu (Z.ai): The First Listed LLM Company and the GLM Agent Bet — Zhipu AI reached public markets before any other large-model lab — listing in Hong Kong in January 2026 — and spent the following months shipping the GLM-5 line into an open-weight agentic flagship while raising prices twice. This file collects the sourced record: the models, the phone-use agent bet, the economics, and our read on what a listed lab means for tracking AI's real impact.
- Qwen: The Open-Weight Leader Starts Charging for the Crown — Alibaba's Qwen is the most-downloaded open-weight model family in the world — by company count, more than three billion downloads and over half the open-source market. In August 2026 it shipped its biggest flagship yet, priced it far under US frontier rates, and put its open weights under a revenue-share license for the first time. This file collects the sourced record and our read on what the pivot means.
- DeepSeek: The Price Anchor Starts Moving — Eighteen months after the R1 moment made frontier-for-pennies the industry's reference point, DeepSeek promoted V4-Pro to general availability — and in the same week announced API price increases of up to 1,100% plus the sector's first peak/off-peak token pricing. The lab that anchored the price war is repricing. This file collects the sourced record and our read.
- MiniMax: The Multimodal Tiger That Doubled on Debut — MiniMax reached the Hong Kong exchange one day after Zhipu and doubled on its first day. But the reason it closes this series is not the listing — it is the product surface. Where the other five files cover text and agents, MiniMax ships video, speech and music at commodity prices, and that points the AI shockwave at a different cohort: creators. This file collects the sourced record and our read.
- A Tribute to Manus — The independent special that anchors this site's digital-labor storyline.