Teams paying DeepSeek, Qwen, or GLM API bills just watched three moves that look contradictory. DeepSeek raised some API tiers by as much as 1,100%. Alibaba, the same week, open-weighted a 2.4-trillion-parameter flagship it had never released. Zhipu shipped GLM-5.3 and boosted coding benchmarks by roughly 6x on the same base model — no retraining. The takeaway: China's labs are shifting from competing on price alone to competing on pricing power. This piece covers the timeline, the rate card, the three strategies, a head-to-head price table, disputed claims, and a six-step checklist.
SECTION 01 Four traps before you quote the 1,100% DeepSeek hike
- The headline is one billing line, not the invoice: "11x", "over 1,100%", and "350%" are all correct. They map to peak cache-hit input, output, and cache-miss input. Output (350%) and cache-miss input move real bills more than the near-zero cache-hit tier.
- Peak hours are Beijing time: 9am–12pm and 2pm–6pm. Everything else is off-peak at half the peak rate. The announcement is asking you to shift load.
- Open weights is not Apache 2.0: Qwen3.8-Max uses a custom license. A Model-as-a-Service or AI Work Assistant business over $50 million in a 12-month window needs a separate commercial deal. Claims that US, EU, UK, or Korea users cannot download the weights are false.
- GLM-5.3 is not a new base: Same 743B parameters as GLM-5.2. Gains come from scaled post-training RL. The scores are vendor-reported; no independent re-run is public yet. Earlier product notes: DeepSeek V4 GA and time-of-day billing, Qwen3.8-Max launch brief.
SECTION 02 Timeline: what shipped, and when prices actually changed
| Date | Event |
|---|---|
| Jul 16, 2026 | Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny |
| Jul 30, 2026 | OpenAI cuts GPT-5.6 Luna by 80% |
| Aug 2–3, 2026 | Alibaba previews, then launches, Qwen3.8-Max as a hosted API |
| Aug 6–7, 2026 | OpenAI makes Luna the free default with unlimited text chats |
| Aug 10, 2026 | Meta releases Muse Glimmer (30B, Apache 2.0) and teases open weights for flagship Muse Spark 1.2 |
| Aug 12, 2026 | Alibaba publishes Qwen3.8-2.4T-A95B on Hugging Face / ModelScope; xAI ships Grok 4.6 |
| Aug 13, 2026 | DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash |
| Aug 14, 2026 | Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base |
| Aug 17, 2026, 00:00 Beijing | DeepSeek's new pricing takes effect |
Zoom out: while Chinese labs raised prices and opened flagship weights, US labs cut prices and went free at the consumer layer. That is two sides of the same fight. OpenAI's cut is covered in the GPT-5.6 Luna / Terra / Sol price-cut note.
SECTION 03 The numbers: DeepSeek tiers, Qwen specs, GLM post-training only
New DeepSeek rates take effect at 00:00 Beijing time on Aug 17. Peak hours are 9am–12pm and 2pm–6pm Beijing time. Figures below are RMB per 1M tokens, cross-checked against the official announcement and multiple reports.
| Billing item | Old price | New off-peak | New peak | Peak increase |
|---|---|---|---|---|
| V4-Flash cache hit (input) | ¥0.02 | ¥0.05 | ¥0.10 | ~400% |
| V4-Flash cache miss (input) | ¥1.0 | ¥1.5 | ¥3.0 | 200% |
| V4-Flash output | ¥2.0 | ¥4.5 | ¥9.0 | 350% |
| V4-Pro cache hit (input) | ¥0.025 | ¥0.15 | ¥0.30 | ~1,100% |
| V4-Pro cache miss (input) | ¥3.0 | ¥4.5 | ¥9.0 | 200% |
| V4-Pro output | ¥6.0 | ¥13.5 | ¥27.0 | 350% |
The 1,100% figure applies to peak-hour cache-hit input — the tier that started closest to free. Output, which dominates most real bills, rose 350%. Independent modeling of a heavy off-peak, half-cache-hit workload lands closer to 1.8x.
| Spec | Detail |
|---|---|
| Parameters | 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared) |
| Context window | 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max defaults to 1M |
| Release cadence | Preview Aug 2 → API live Aug 3 → open weights Aug 12 |
| API pricing (international) | $2/M input, $6/M output |
| License | Not Apache 2.0 — a custom Qwen3.8-Max License |
| Why it matters | First time Alibaba has open-weighted a Max-tier flagship; Qwen3.5 / 3.6 / 3.7 Max stayed API-only |
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | +23.7 pts |
| DeepSWE v1.1 | 46.2% | 66.9% | +20.7 pts |
| Agents' Last Exam (CLI) | 23.8% | 28.5% | +4.7 pts |
| CyberGym | 77.2% | 84.5% | +7.3 pts |
| AutomationBench | 26.2% | 48.2% | +22.0 pts |
These are Zhipu's own numbers. No independent third-party re-run has been published. GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0. It is a top open-weight result, not an outright frontier win.
SECTION 04 Three strategies: time-of-day pricing, a custom license, post-training scale
DeepSeek: this is a capacity problem, not a strategy pivot
The easy misread is "China's cheapest model finally caved to margin pressure." The structure reads more like the opposite: compute constraints made visible on the price sheet. Flat, always-cheap pricing worked as acquisition while GPU supply kept up. Once usage grew exponentially and capacity did not, something had to become explicit. "Encouraging more flexible workload scheduling" is corporate-speak for "peak-hour compute is now scarce — shift the load yourself."
One detail most international coverage missed: at peak hours, DeepSeek's official API is now higher than several third-party resellers (GMI Cloud, Novita, and others still list V4 Pro below the new official peak). The assumption that the official API is always the cheapest way to run DeepSeek is broken for the first time.
Alibaba: open weights buy mindshare; the license protects the revenue ceiling
Alibaba did two things at once. It published the full 2.4T checkpoint for free download. It also attached a custom license — not the Apache 2.0 used for smaller Qwen models — that requires any Model-as-a-Service or AI Work Assistant business earning over $50 million in any 12-month period to negotiate a separate commercial license, and requires products with 100M+ monthly active users or $20M+ in monthly revenue to display the model name prominently.
Give away the weights to win developers, especially internationally. Keep leverage over the handful of companies that can build a competing inference business on top. That is a different bet than Meta's Muse Glimmer, which ships under unrestricted Apache 2.0. Claims that the license bans downloads from the US, EU, UK, and South Korea are false. The published text has no geographic clause.
GLM-5.3: no new base model — the method is the story
Same 743B-parameter base as GLM-5.2. No retraining. A roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) from scaling reinforcement-learning environments in post-training. As pretraining scaling laws show diminishing returns, post-training RL is becoming an independent lever with a much lower cost floor than retraining a foundation model. Mid-tier labs without OpenAI-scale compute can still close the gap on agentic and coding benchmarks.
V4-Pro per 1M tokens CNY (Beijing peak 09-12 / 14-18)
off-peak cache-hit 0.15 miss 4.5 out 13.5
peak cache-hit 0.30 miss 9.0 out 27.0
headline ~1100% = peak cache-hit input only
SECTION 05 Head-to-head: is DeepSeek still the cheapest frontier-class model?
| Model | Input | Output | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (peak) | ¥9.0 (~$1.26) | ¥27.0 (~$3.78) | No |
| DeepSeek V4-Pro (off-peak) | ¥4.5 (~$0.63) | ¥13.5 (~$1.89) | No |
| Qwen3.8-Max (international API) | $2.00 | $6.00 | Yes (custom license) |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 | No |
| Claude Opus 5 (implied, per Alibaba's ratio) | ~$5.00 | ~$25.00 | No |
RMB-to-USD at about ¥7.15/$1, approximate. Off-peak V4-Pro is still well below Claude Opus 5, but it is no longer the outright cheapest option. Qwen3.8-Max international pricing and OpenAI's Luna both undercut DeepSeek's off-peak rate. "Chinese model = cheapest model" held for most of 2025 and early 2026. It is not a safe assumption anymore.
SECTION 06 What is disputed or still unverified
- The 1,100% headline is accurate and misleading without context. It applies only to peak-hour cache-hit input. Output — the cost that dominates most real bills — rose 350%.
- Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips (and "Pangu AL128" supernodes), reported by several Chinese financial outlets, have not been independently confirmed by Alibaba technical documentation or third-party benchmarks. Treat as vendor-adjacent, unverified reporting.
- GLM-5.3's reported discovery of a "serious vulnerability" in Cursor comes from VentureBeat and Zhipu's own disclosure. Technical details have not been made public. Read it as vendor-sourced, not independently audited.
- Reports that China's Ministry of Commerce may be preparing retaliatory export controls on AI / semiconductor technology are speculative and sourced to unconfirmed media reports, not an official announcement.
SECTION 07 Why this matters: two price wars running in parallel
Over the past month, China's top labs have shipped at a pace domestic financial media calls "three model updates a week." DeepSeek, Alibaba, and Zhipu sit on top of Moonshot's Kimi K3 (open-weighted Jul 16, 2.8T) and MiniMax H3. Chinese coverage frames this as open-weight releases forcing a global repricing of the industry — more assertive than most English-language product write-ups. Background on Kimi K3: Kimi K3 full open-weight release.
US labs are running the opposite play at the consumer layer. OpenAI cut its cheapest tier 80% on Jul 30, then made that model free and unlimited a week later. Google shipped a coding-focused model at half the price of its three-week-old predecessor on Aug 13. Chinese labs open-weight flagships and introduce tiered, higher pricing on the compute-constrained top end. US labs race toward free and cheap at the consumer end. Both are real. They optimize different parts of the funnel.
There is also a geopolitical layer worth naming carefully. Moonshot's Kimi K3 already drew US security scrutiny. Some analysts read Alibaba's choice of this window to open-weight a 2.4T flagship as a move to lock in international mindshare and a "technological parity" narrative before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact — and easy to miss if you only read English-language tech press covering these as isolated product news.
SECTION 08 Six-step checklist: how to read the hike against your own bill
- Split the headline before you multiply the invoice. Cache-hit input, cache-miss input, and output are different lines. Do not apply 1,100% to the whole month.
- Map traffic to Beijing peak windows. 9am–12pm and 2pm–6pm Beijing time bill at peak. Batch evals, offline distillation, and overnight agent replay belong off-peak.
- Estimate with your real cache-hit rate. Heavy users with high cache hits and mostly off-peak traffic have been modeled around 1.8x. Peak plus low hit rate is what approaches the headline.
- Read the Qwen3.8-Max License file, not the announcement thread. Personal and internal use are largely unaffected. MaaS / AI Work Assistant revenue over $50 million in any 12-month period needs a separate commercial license. 100M+ MAU or $20M+ monthly revenue requires prominent model-name display. No geographic ban.
- Keep GLM-5.3 vendor scores separate from third-party re-runs. 4.6% → 28.3% on Terminal-Bench 3.0 is self-reported. Do not write "frontier win" until an independent re-run exists.
- Park unverified claims in a separate ledger. Zhenwu M890 / Pangu AL128, the Cursor vulnerability, and export-control rumors are not confirmed facts. Sensitive evals and off-peak experiments that need a controlled egress and Root belong on an isolated physical node, not on a shared cloud session.
SECTION 09 Citeable figures and primary sources
- V4-Pro peak cache-hit input: ¥0.025 → ¥0.30, ~1,100%; output ¥6.0 → ¥27.0, 350%
- Qwen3.8-2.4T-A95B: 2.4T total / 95B active; 262K native, ~1.01M extendable; international API $2 / $6
- GLM-5.3 Terminal-Bench 3.0: 4.6% → 28.3% (+23.7), same 743B base, no retraining; vendor-reported
- License thresholds: $50M/12 months MaaS or AI Work Assistant needs a separate license; 100M MAU or $20M monthly revenue needs prominent naming
Official and third-party links below. Verify the latest official pricing and license terms before you republish. Details flagged above as unverified (domestic chip claims, the Cursor vulnerability report, and export-control rumors) have not been independently confirmed.
Official / repository sources:
DeepSeek API Docs — DeepSeek-V4-Pro GA Release (pricing update)
Hugging Face — Qwen/Qwen3.8-2.4T-A95B
Z.ai — GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
Third-party reporting:
21jingji — DeepSeek peak/off-peak rates take effect, peak hike up to 1,100%
VentureBeat — GLM-5.3 and the reported Cursor vulnerability
Once a shared API moves to time-of-day pricing, cache-hit rate and egress isolation become bill variables. Stacking evals, distillation, and 24/7 agents on one shared session makes both cost and exit policy harder to control. For production work that needs zero-loss native compute, stable iOS CI/CD, and always-on AI agents, a VPSNIX cloud physical node is usually the better fit — genuine Apple hardware, full Root, no hypervisor tax, billed by the day, week, or month, so you can split off-peak experiments from the production pipeline. See the Mac mini M4 rental breakdown and the pricing page.
SECTION 10 FAQ
Is DeepSeek still cheaper than GPT-5.6 or Claude after the price hike?
Its off-peak rate is still cheaper than Claude Opus 5, but it is no longer the single cheapest option overall. OpenAI's GPT-5.6 Luna ($0.20 / $1.20 per million tokens) and Alibaba's international Qwen3.8-Max pricing ($2 / $6) now undercut DeepSeek's new off-peak rates on at least one dimension. DeepSeek is still relatively cheap for a frontier-class model, just not the outright cheapest anymore.
Can I use Alibaba's Qwen3.8-Max open weights for free in a commercial product?
Yes for most use cases — personal projects and internal enterprise use are unaffected. The catch applies if you run a Model-as-a-Service or AI Work Assistant business that has earned over $50 million in any consecutive 12-month period; that tier requires a separate commercial license from Alibaba.
Is Qwen3.8-Max banned or restricted for US, EU, or UK users?
No. That claim circulated online but is false — the published license contains no geographic restriction of any kind. The restrictions are revenue-based, not tied to where you or your users are located.
What is actually different between GLM-5.3 and GLM-5.2?
Nothing at the base-model level — both use the same 743-billion-parameter foundation model. The performance gains (roughly 6x on Terminal-Bench 3.0) come entirely from scaling up reinforcement learning during post-training, with no retraining of the base model.
Will Meta actually open-source its flagship model, not just the smaller Muse Glimmer?
Not yet. Muse Glimmer is a 30B distilled model, not Meta's real flagship. Mark Zuckerberg has said open weights for the larger, closed Muse Spark 1.2 are coming "soon," which — if it happens — would make it the first US flagship-tier model released openly. As of this writing, that release has not happened; treat it as a stated intention, not a confirmed fact.