MModelspectraIndependent AI Model Intelligence
Model profile · Updated 2026-09-08
中文
Model Profile · DeepSeek · CN

DeepSeek-V4.1-Flash

552B MoE with 8B-active prefill / 16B-active decode. CED encoder-decoder roughly halves prefill; CSA2 cross-layer KV sharing plus FP4 shrink global KV to ~1/4 and persistent KV to ~1/8. Price/spec fields are estimated from the V4-Flash tier pending the full model card.

Best forLong-context AI agents, Agentic coding, High-throughput deployment
LicensingOpen weights
Updated2026-09-08
01

Snapshot

Overall
76
rank #12 of 27
Coding
89
rank #5 of 27
Multimodal
78
rank #16 of 27
Input / Output
$0.14 / $0.28
USD per 1M tokens
Context
1M
usable recall 98
Speed
~85 tok/s
TTFT 0.3s

Aggregated from public sources and independently weighted; methodology on the Terms page. Scores within 3 points are treated as statistically tied.

02

Specs & pricing

Every tracked field. List API prices in USD per 1M tokens; verify current vendor pricing before purchase.

VendorDeepSeek (CN)
Released / current2026.09
LicensingOpen weights
Context window1M
Max output128K
Effective-context score98
Input / output price$0.14 / $0.28 per 1M tokens
Cache discountNo published cache discount
Free tierDeepSeek App free; open weights to self-host
Speed / TTFT~85 tok/s / 0.3s
Function calling84
Refusal rate~5%
English / Chinese84 / 86
Modalitiestext, image
Fine-tuningYes
Private deploymentYes (enterprise)
SOC2 / no-trainno / yes
03

Strengths & watch-outs

Strengths

  • 1M context at ~890 bytes/token global KV
  • Top agentic coding (DeepSWE 74.2, Terminal-Bench 2.1 90.6)
  • Persistent KV cut to ~1/8 via bounded replay
  • Open weights, extremely low serving cost

Watch-outs

  • Gap on hardest scientific agents (Terminal-Bench 4.0)
  • Multimodal below top closed models
  • Younger, smaller ecosystem
04

Where it ranks among 27 models

Overall 76/100 (#12), coding 89 (#5), multimodal 78 (#16). See the full boards and side-by-side compare on the leaderboard.
05

Pricing in practice

Illustrative monthly bill for 100M input + 30M output tokens: $22 at list. Your mix and cache-hit ratio change this.

Illustrative model using public list rates before any enterprise agreement; cache discount: No published cache discount.

06

Frequently asked questions

How much does DeepSeek-V4.1-Flash cost?
List API pricing is $0.14 input and $0.28 output per 1M tokens. No published cache discount. DeepSeek App free; open weights to self-host
What is DeepSeek-V4.1-Flash's context window, and is it usable end to end?
It advertises a 1M window with 128K max output; its effective-context score is 98/100, which is the better predictor of whether long-document details are actually retained.
Is DeepSeek-V4.1-Flash good for coding?
Its coding aggregate is 89/100, rank #5 of 27. That puts it in the top tier for agentic, multi-file engineering.
Can I self-host DeepSeek-V4.1-Flash?
Yes — it ships open weights and supports private deployment, so you can run it on your own infrastructure for data control; budget for GPU and operations.
How strong is it in Chinese and on multimodal inputs?
Chinese score is 86 versus English 84; multimodal aggregate is 78/100 (rank #16), accepting text, image.