Home / Blog / V4-Flash
ENGINEERING_BLOG · 2026.08.05

Is DeepSeek's New Model Really 100x Cheaper Than Claude? Inside the V4-Flash Benchmarks

V4-FLASH-0731 · TERMINAL BENCH 2.0
82.7

Vendor-reported agent score — beats V4-Pro preview (67.9) at 284B/13B, measured via unreleased Harness minimal mode

Short answer: on list price, yes — but the benchmarks need context. DeepSeek's official V4-Flash-0731 went live on July 31, 2026 with the same 284B-parameter architecture as April's preview — only the post-training changed. It now beats DeepSeek's own larger V4-Pro preview on agent benchmarks at roughly 1/36 to 1/179 of Claude Opus 4.8's price. This article covers the release timeline, pricing matrix, harness-dependent benchmark caveats, and the "kill line" pricing dynamic reshaping China's open-weight race. Bottom line: V4-Flash optimizes for "good-enough intelligence at a price nobody else can match," not leaderboard supremacy — and the flagship V4-Pro plus in-house Harness framework remain unreleased.

SECTION 01 Four red flags behind the "100x cheaper than Claude" headline

Teams routing high-volume agent traffic need to separate the pricing story from the capability story. Four transparency gaps matter more than any single Terminal Bench row.

  • Harness-dependent scores: V4-Flash-0731's Terminal Bench 2.0 score of 82.7 was measured using DeepSeek's own unreleased Harness in "minimal mode" at max settings. DeepSeek's changelog explicitly warns that agent scores are "extremely sensitive to harness choice" — treat headline numbers as framework-specific, not portable.
  • API-only GA: The July 31 update ships through the API only. The consumer app and web chat were not updated — production teams must verify which endpoint their integration actually hits.
  • Real-world friction reports: Chinese financial outlet 21st Century Business Herald, citing overseas developer feedback, reported low input cache-hit rates and occasional safety-classifier timeouts on the official build — a gap between benchmark narratives and day-to-day reliability.
  • V4-Pro and Harness dates unconfirmed: Multiple Chinese outlets cite an August 10–20 GA window from unnamed sources. DeepSeek's own changelog says only that the official V4-Pro release "will follow as soon as possible" — no date attached.

Timeline:

  • April 24V4 preview launches with V4-Pro (1.6T/49B) and V4-Flash (284B/13B), both MIT open weights with 1M-token context.
  • July 24 — Legacy aliases deepseek-chat and deepseek-reasoner retired; all traffic routes to the V4 family.
  • July 27Moonshot AI ships Kimi K3 open weights (2.8T total parameters), raising competitive pressure days before DeepSeek's update.
  • July 31deepseek-v4-flash promoted to official API beta (build tag "0731"). Same architecture, fresh post-training. Open weights land on Hugging Face the same day. Changelog names "DeepSeek Harness" for the first time.
  • August 2Qwen3.8-Max goes GA through Alibaba Cloud, adding another trillion-parameter competitor to the same window.

SECTION 02 The numbers DeepSeek published — and what is actually verified

Frontier model pricing and status (early August 2026)
Model Status Total / active Input (miss/hit per 1M) Output (per 1M)
DeepSeek-V4-Flash-0731 Official (Jul 31) 284B / 13B $0.14 / $0.0028 $0.28
DeepSeek-V4-Pro Preview only (Apr 24) 1.6T / 49B $0.435 / $0.003625 $0.87
Kimi K3 Open weights (Jul 27) 2.8T / ~104B (est.) $3.00 / $0.30 $15.00
Qwen3.8-Max API GA (Aug 2); weights pending 2.4T / 95B $2.00 / ~$0.17–0.25 $6.00
Claude Opus 4.8 Closed source Undisclosed Vendor list price Vendor list price

All pricing figures are vendor-published rates. DeepSeek has announced a future 2x peak-hour surcharge (9am–12pm and 2pm–6pm Beijing time) with no confirmed effective date yet. Per 21st Century Business Herald, V4-Flash runs roughly 36x cheaper than Opus 4.8 on cache-miss input, ~179x on cache-hit input, and ~89x on output.

SECTION 03 How DeepSeek squeezed more agent performance out of the same 284B model

The most easily missed detail: V4-Flash-0731 is identical in size and structure to April's preview. DeepSeek says the entire performance jump on agent benchmarks came from re-running post-training, not scaling up — a 284B/13B model now beats a 1.6T/49B sibling on multiple agentic tasks.

  • Hybrid attention (CSA + HCA): DeepSeek's technical report describes Compressed Sparse Attention combined with Heavily Compressed Attention, marketed as "DSA," aimed at cutting compute and memory at long context lengths.
  • Manifold-Constrained Hyper-Connections (mHC): An enhancement over standard residual connections, carried over from the April preview architecture.
  • Muon optimizer: Replaces the traditional optimizer for faster convergence and training stability.
  • Efficiency claims at 1M context: DeepSeek reports V4-Pro needs only 27% of V3.2's per-token inference FLOPs and 10% of KV cache footprint at million-token lengths — vendor-reported figures not yet independently reproduced.
  • Harness debut: July 31 marked the first official mention of DeepSeek Harness, an in-house agent framework for file I/O, tool calls, and multi-step engineering — positioned as DeepSeek's answer to Claude Code. Every agent benchmark published for V4-Flash-0731 was measured in Harness's unreleased "minimal mode."

SECTION 04 DeepSeek V4-Flash vs Kimi K3 vs GLM-5.2 vs Qwen3.8-Max

Independent index vs vendor agent scores (early August 2026)
Model Artificial Analysis Intelligence Index Avg. cost per task (AA) DeepSeek agent benchmarks
DeepSeek-V4-Flash-0731 50 (independent) $0.03 Terminal Bench 2.0: 82.7 (vendor + Harness)
Kimi K3 57 $0.86 Not vendor-reported in this release
GLM-5.2 ~1 point above V4-Flash Not verified for this piece Not vendor-reported in this release
GPT-5.6 Sol 9+ points above V4-Flash $1.86 Closed source
Claude Fable 5 9+ points above V4-Flash $3.15 Closed source

The tension: on Artificial Analysis's independent index, V4-Flash trails both Kimi K3 and GLM-5.2 on raw intelligence — but its per-task cost is roughly 1/29th of K3, 1/62nd of GPT-5.6 Sol, and 1/105th of Claude Fable 5. DeepSeek is not competing for the top of the leaderboard; it is optimizing for high-volume agent and batch workloads. The preview reportedly topped OpenRouter's most-used model ranking for seven consecutive weeks. See also July OpenRouter usage rankings and GPT-5.6 Luna's 80% price cut in the same pricing war window.

SECTION 05 Six-step checklist before you route agent traffic to V4-Flash

  1. Migrate legacy model names: If you still call deepseek-chat or deepseek-reasoner, switch to deepseek-v4-flash immediately — both legacy aliases were retired July 24.
  2. Separate vendor scores from independent indices: Terminal Bench 2.0 (82.7) is Harness-dependent and vendor-reported. Artificial Analysis Intelligence Index (50) is third-party. Do not conflate the two when sizing production workloads.
  3. Benchmark on your own harness: DeepSeek's scores used unreleased Harness minimal mode. Blind-test V4-Flash on your actual agent stack — Claude Code, Cursor, OpenCode, or your in-house runner — before migrating CI pipelines.
  4. Model cache-hit economics: V4-Flash's $0.0028/M cache-hit input is where the 179x Claude discount lives. Monitor cache-hit rates; early reports suggest hit rates may be lower than expected on the official build.
  5. Plan for peak-hour surcharges: DeepSeek has announced 2x pricing during Beijing weekday peak hours (9am–12pm, 2pm–6pm) with no confirmed start date. Build scheduling logic before scaling 7×24 agent loops.
  6. Hold V4-Pro decisions until GA: The official V4-Pro release and public Harness remain "coming soon." Treat any August 10–20 GA rumor as unconfirmed until DeepSeek's changelog or official account confirms it.

SECTION 06 Citeable data and sources

  • Terminal Bench 2.0: 82.7 (V4-Flash-0731) vs 67.9 (V4-Pro preview) — vendor-reported via unreleased Harness minimal mode
  • Artificial Analysis Intelligence Index: V4-Flash 50, Kimi K3 57, GPT-5.6 Sol and Claude Fable 5 both 9+ points higher
  • Per-task cost (Artificial Analysis): V4-Flash $0.03, Kimi K3 $0.86, GPT-5.6 Sol $1.86, Claude Fable 5 $3.15
  • Efficiency at 1M context (vendor-reported): V4-Pro needs 27% of V3.2 FLOPs and 10% of KV cache per token
  • Chinese developer framing: "Kill line" (斩杀线) — DeepSeek's good-enough performance plus rock-bottom price sets a bar competitors must beat on both axes or risk irrelevance
  • Market reaction: Nvidia, Broadcom, and AMD saw no significant stock movement on July 31 — a contrast to the 2025 DeepSeek-R1 selloff, suggesting markets now treat efficiency gains as normal engineering

Verify against official sources before you publish or migrate:

DeepSeek API documentation and changelog

Hugging Face — DeepSeek-V4-Flash model card (MIT license)

Artificial Analysis — independent model intelligence index and cost benchmarks

21st Century Business Herald — V4-Flash release coverage and developer feedback

Cheap API tokens look attractive until a 7×24 agent loop runs up an uncapped bill, harness-dependent benchmarks resist independent reproduction, and cloud LLMs still cannot replace Xcode signing or Metal shader debugging on virtualized macOS. For teams that need to iterate agent prompts locally, A/B route between V4-Flash and Kimi K3, and keep iOS CI/CD plus agent automation on stable bare metal, VPSNIX cloud physical nodes are usually the better production substrate — 100% Apple hardware, full root access, zero hypervisor tax, elastic day/week/month billing. Pair with Mac mini M4 rental pricing analysis and the pricing page for a two-tier stack.

SECTION 07 FAQ

Is DeepSeek V4 open source?

Yes. Both V4-Pro and V4-Flash, including the July 31 official V4-Flash-0731 build, ship as open weights under the MIT license on Hugging Face. You can use, fine-tune, and redistribute them commercially without additional permission.

How much cheaper is DeepSeek V4-Flash than Claude?

Based on figures reported by 21st Century Business Herald, official V4-Flash pricing runs roughly 36x cheaper than Claude Opus 4.8 on cache-miss input, about 179x cheaper on cache-hit input, and about 89x cheaper on output, per million tokens. These are vendor list prices, not an independent audit.

When will DeepSeek V4-Pro's official version be released?

There is no confirmed date. DeepSeek's changelog says only that the official V4-Pro release "will follow as soon as possible." Reports of an August 10–20 general-availability window come from unnamed sources in Chinese media and have not been confirmed by DeepSeek.

Can I trust DeepSeek's benchmark numbers?

Partially. Widely adopted third-party benchmarks like SWE-bench Verified carry more weight. Agent-specific scores (Terminal Bench 2.0, Toolathlon, etc.) were measured with DeepSeek's own unreleased Harness framework, and the company warns these numbers are highly sensitive to harness choice — wait for independent reproduction with other agent tools before treating them as general capability claims.

What is DeepSeek Harness?

It is DeepSeek's first self-developed agent execution framework, positioned as an in-house alternative to Claude Code, for tasks like file editing, tool calls, and multi-step engineering work. It was named for the first time in the July 31, 2026 changelog and is not yet publicly available.

Should I pick V4-Flash or V4-Pro for daily use?

For bulk agent loops, batch processing, or cost-sensitive high-volume calls, V4-Flash-0731 offers better value and now beats the V4-Pro preview on agent benchmarks. For deeper world knowledge or complex reasoning with a larger budget, V4-Pro preview remains available — but wait for the official V4-Pro GA before making long-term architecture decisions.