DeepSeek V4.1 Flash Explained: MoE Architecture, KV Cache, Benchmarks

DeepSeek V4.1 Flash Explained: MoE Architecture, KV Cache, Benchmarks

DeepSeek V4.1 Flash Explained: MoE Architecture, KV Cache, Benchmarks

Most open-weight model releases compete on one axis: a bigger number on a leaderboard. DeepSeek V4.1 Flash competes on a different one. Its model card is titled “Pushing the Limits of KV Cache Compression”, and the headline claim is a global KV cache of 890 bytes per token, roughly a quarter of its predecessor and, by DeepSeek’s own figure, 437 times smaller than DeepSeek V1. The weights are a 552B-parameter multimodal mixture-of-experts (MoE) backbone under the MIT license, and the model touches only 8B parameters per token while reading a prompt and 16B while writing the answer.

That matters now because the expensive part of agentic workloads is no longer generating tokens. It is holding and re-reading very long contexts, again and again, across tool calls. This post walks through what DeepSeek actually published, separates vendor-reported numbers from independent ones, works the memory and cost arithmetic, and shows what it takes to run the model yourself or call it by API.

What this covers: the lineage from V4 Flash, the Causal Encoder-Decoder layout and CSA2 attention, training, the reported benchmark table with caveats, access paths and pricing, hardware reality for self-hosting, failure modes, and a decision matrix against peer models.

Context and Background

DeepSeek has spent three years turning inference cost into a research agenda. DeepSeek-V2 introduced Multi-head Latent Attention (MLA), which compresses keys and values into a low-rank latent so the cache per token shrinks. V3 added fine-grained MoE routing with an auxiliary-loss-free balancing scheme. The V4 generation, covered in our DeepSeek V4 architecture and benchmarks deep dive, extended context to one million tokens and introduced compressed sparse attention so that long prompts stopped scaling cache linearly with every layer.

The V4 family shipped in two sizes. V4-Pro is a roughly 1.6T-parameter MoE with 49B active parameters, and V4-Flash is a 284B-parameter MoE with 13B active, according to the base-model table on the V4.1 card. V4.1 Flash, released on 10 September 2026 according to reporting by third-party trackers and consistent with the Hugging Face listing, is not a small distillation of Pro. It is a larger backbone (552B) with fewer active parameters per token, and a new attention and cache design built around the observation that agent workloads are input-heavy.

Why input-heavy? A coding agent that holds a 300K-token repository context and calls a tool every few seconds re-processes that context on each turn unless a cache saves it. Prefill dominates cost; decode is a small tail. If the architecture can run prefill on a thinner slice of the network than decode, and shrink the cache that persists between turns, the economics change more than a few extra benchmark points could achieve. For readers new to the mechanics, our primer on KV cache optimization for LLM inference explains why cache bytes per token, not FLOPs, usually set the concurrency ceiling.

The primary sources for everything below are the DeepSeek-V4.1-Flash model card on Hugging Face, which links the PDF technical report, and the DeepSeek API pricing page. Where this post cites a figure from either, it is the vendor’s own claim. We have not independently reproduced any benchmark, and the card states that base-model numbers come from DeepSeek’s internal evaluation framework.

What DeepSeek V4.1 Flash Actually Is

DeepSeek V4.1 Flash is an open-weight, natively multimodal MoE language model with a 552B-parameter backbone, a one-million-token context window, and an MIT license. It uses a Causal Encoder-Decoder layout that activates 8B parameters per token during prefill and 16B during decode, and a compressed attention scheme that stores about 890 bytes of global KV cache per token.

DeepSeek V4.1 Flash causal encoder-decoder MoE architecture

Figure 1: DeepSeek V4.1 Flash layout. The decoder reads a global KV cache projected once from the final encoder states, and an Engram memory feeds the MoE layers.

The diagram shows the data path the model card describes. Tokens and image patches enter a 20-layer causal encoder. The encoder’s final hidden states are projected into the global KV cache that the 20-layer decoder attends to, so decoder layers do not each build their own full cache. Each MoE layer holds one shared expert and 384 routed experts and activates 6 routed experts per token. An Engram conditional memory, described as 196B parameters accessed by sparse token-based lookup, supplies additional capacity without being run densely.

Causal Encoder-Decoder: why the split saves compute

Standard decoder-only transformers compute every layer for every token, in prefill and decode alike. The Causal Encoder-Decoder (CED) in V4.1 Flash divides the 40 layers in half. The decoder’s global KV is projected from the final encoder hidden states rather than from each decoder layer’s own hidden states. The model card draws the consequence directly: because of that projection, only 8B parameters need to be active per token in prefill, and 16B in decode.

The mechanism is easy to reason about even without the full report. During prefill the model must produce keys and values for every prompt token. If those can come from the encoder stack alone, the decoder half does not need to process prompt tokens except to the extent needed for the final position. During decode, the full 40-layer path runs for each new token, which is why the decode figure is twice the prefill figure. I read this as a deliberate trade: decode quality pays for twice the active parameters, while the cheaper path serves the phase that is token-heavy in agent loops.

This is my interpretation of the card’s two sentences, and the card does not publish a layer-by-layer FLOP breakdown. Treat the 2x ratio as the vendor’s headline, not a measured serving speedup.

An illustrative calculation helps size the gap. Using the rough rule of two FLOPs per active parameter per token, 8B active parameters is about 16 GFLOPs per prompt token, and 16B is about 32 GFLOPs per generated token. A 100,000-token prompt then costs about 1.6 PFLOPs of dense matmul work at the prefill rate, against 3.2 PFLOPs if it ran through a 16B-active path. This ignores attention cost, routing overhead and memory bandwidth, so it is a floor for the ratio, not a throughput prediction.

The MoE configuration and the Engram memory

Each MoE layer combines one always-on shared expert with 384 routed experts, of which 6 are selected per token. That is a very fine-grained design: the router chooses 6 of 384, about 1.6 percent of routed experts, which keeps active parameters low while total capacity stays high. The shared expert captures common transformations so routed experts can specialize. If you want the serving implications of that routing pattern, our guide to expert-parallel MoE inference serving covers all-to-all communication, load imbalance and placement.

The Engram component deserves a careful reading. The card describes it as “Engram conditional memory (196B parameters, sparsely accessed via token-based lookup)”. That is a table of learned parameters indexed by token context rather than a dense layer. It adds stored knowledge at near-zero compute per token, though it still occupies memory when you host the model. The card does not state whether the 552B headline includes the Engram table. Aggregators have reported a larger total, but the card itself only gives the 552B backbone figure, so this post uses that number and treats a combined total as not confirmed by the primary source.

Multimodality from the first token

V4.1 Flash is natively multimodal. A vision encoder named DeepSeek-ViT, trained from scratch with 2D rotary position embeddings and 3×3 pixel-unshuffle downsampling, plus a two-layer MLP projector, converts images to visual embeddings. Those are processed jointly with text embeddings from the start of language-model pre-training, rather than bolted on after the fact. Output is text only; the model does not generate images.

The reported multimodal base-model scores are MMMU-Pro 56.5, CVBench 77.9, DocVQA 95.6 (LLM-judged) and RefCOCO-average 86.0. There is no V4 base-model comparison for these rows because the card leaves those cells blank for the earlier models, so the multimodal numbers stand alone and cannot show a trend.

How the KV Cache Got 4x Smaller

The KV cache is the memory a transformer keeps so it does not recompute attention keys and values for earlier tokens. Its size per token multiplies by context length and by the number of concurrent sessions, and in long-context serving it is usually what runs out first. DeepSeek’s card claims three stacked reductions that together reach 890 bytes per token for the global cache.

KV cache compression in DeepSeek V4.1 Flash with CSA2 and FP4

Figure 2: The three mechanisms the model card credits for the smaller cache, in the order they stack.

Compressed Sparse Attention 2 and its three layer modes

CSA2 assigns every attention layer one of three static modes: Full, Reindex or Reuse. The purpose is to share the main KV and the indexer keys across layers and to reuse the Top-K sparse-attention indices, so most layers do not pay for their own indexing pass or their own cache. The card does not publish how many layers use each mode in the text we could fetch; that detail lives in the technical report.

In sparse attention, an indexer scores earlier tokens and a Top-K selection decides which ones each query actually attends to. Reuse mode borrows the previous selection, Reindex recomputes the selection against shared keys, and Full does the whole job. Because the modes are static rather than learned per token, the serving stack can plan memory in advance, which is friendlier to kernels than a dynamic scheme.

A Hierarchical Sparse Indexer adds a second bound. In the decoder, later indexing layers are restricted to a candidate pool built by the first Full-mode layer. The card states this “bounds deeper indexer cost independently of context length”. That claim is worth underlining for million-token prompts: the expensive search runs once at full width and deeper layers only choose among its survivors.

FP4 main KV in E2M1

The second lever is precision. The main KV is stored in FP4 using the E2M1 format, with one E4M3 scale shared per 16 channels. E2M1 has one sign bit, two exponent bits and one mantissa bit, so each value is half a byte; the E4M3 scale per 16 channels adds 0.5 bits per value on average. That block-scaled layout is the same family of tricks used for FP4 weights on recent accelerators, applied here to the cache.

The practical point is that 4-bit cache is only safe because the scale is local. A single global scale would let one outlier channel crush the resolution of the rest. Sixteen-channel blocks keep each scale tied to a small neighborhood, at the price of a small metadata overhead. DeepSeek does not publish an ablation in the card showing the accuracy cost of FP4 KV, so whether the quality loss is negligible on your task is something to measure.

SWA Bounded Replay

Sliding-window attention (SWA) layers keep only the most recent window of keys and values. Normally a serving system persists that state, often to SSD, so that a returning session can resume without recomputing. V4.1 Flash instead uses SWA Bounded Replay: it reconstructs missing SWA KV by replaying only the most recent n_win tokens, so SWA KV never has to be written to disk. The card says this reduces the persistent KV footprint to roughly 1/8 of V4-Flash.

Notice that this is a different quantity from the 890-byte figure. The 890 bytes per token is the global KV cache, quoted as about 1/4 of V4-Flash. The 1/8 figure is the persistent footprint, meaning what must survive between turns. Mixing the two up is the most common error I have seen in early summaries, and the model card keeps them separate.

Worked arithmetic (illustrative)

Take the card’s figures at face value and do the sums. At 890 bytes per token, a 1,000,000-token context holds about 890 MB of global KV cache. If V4-Flash was about four times larger, the same context would need roughly 3.6 GB. The card’s 437x claim against V1 implies roughly 389 KB per token for the V1 design, or about 389 GB for a million tokens, which is why nobody served V1-style attention at that length.

Now scale to concurrency. One hundred simultaneous sessions, each holding a full 1M-token context, need about 89 GB of global KV at V4.1 Flash rates, against about 356 GB at the V4-Flash rate. These are illustrative products of vendor-reported ratios, they cover only the global cache, and they exclude weights, the SWA state, activations and framework overhead. Still, the shift from “needs a multi-node cluster for cache alone” to “fits beside the weights on a single large node” is the real story of this release.

Why this matters beyond DeepSeek

Cache compression compounds with the techniques in our AI inference cost optimization guide. Prefix caching only pays off if you can afford to keep prefixes resident; a cache that is 4x smaller keeps 4x more prefixes warm. The same logic applies to prefill-decode disaggregation, where cache transfer between pools is a bandwidth bill. A smaller cache makes both strategies cheaper, independent of the model’s quality.

Training and Post-Training

The card gives a compact training recipe, and it is worth reading what it omits as carefully as what it states.

DeepSeek V4.1 Flash training pipeline from pre-training to on-policy distillation

Figure 3: Reported training stages for DeepSeek V4.1 Flash, from the 45T-token pre-training run to on-policy distillation.

Pre-training

V4.1 Flash is trained from scratch on a multimodal corpus of 45 trillion tokens. Sparse attention is trained at a sequence length of 64K, and the context is extended to 1M tokens at the 34T-token mark. Starting the sparse-attention training at 64K and extending late is a standard way to avoid paying quadratic-ish costs for the whole run, and the 34T figure tells you roughly 76 percent of the corpus had been consumed before the long-context extension began (34 divided by 45).

The card does not disclose the hardware, the GPU-hours, the data mixture, the languages, or the dollar cost of the run. Any figure you see quoted for training cost is an outside estimate. The parts DeepSeek chose to detail are architectural, which fits the title of the release.

Post-training: data pipeline over algorithm

The post-training description is deliberately plain: SFT, then RL, then on-policy distillation (OPD), “without algorithmic modifications”. All the substantive changes lie in the data pipeline: large-scale automated synthesis of agent tasks and environments, with progressive scaling of data, tasks and rollouts.

I find that framing credible and instructive. The benchmark profile of the model, which leads or ties on most agentic rows while trailing on knowledge-heavy base metrics, matches a team that invested in synthetic environments rather than in a novel optimizer. It also means the lesson is transferable. If you build agents, your leverage is the quality and diversity of environments, not the RL algorithm name.

On-policy distillation is the step to understand. In OPD, a student samples its own trajectories and a stronger teacher or a set of specialist models supplies the learning signal on those samples, so the student is corrected on the states it actually visits rather than on teacher-chosen states. The card does not say which teachers were used, and we should not guess.

Reasoning effort as a dial

The model exposes a continuously controllable reasoning effort, an integer from 1 to 100. All instruct results on the card use the maximum, reasoning_effort=100. A dial that fine is a cost-control feature more than a quality one: you can spend 100 on a hard repository bug and 10 on a classification call. Our write-up on reasoning effort control and thinking budgets discusses how to route requests onto such a dial.

One caution. The card publishes results only at effort 100. How accuracy and token usage change at 30 or 60 is not in the material we reviewed, so any production policy needs your own sweep. Do not assume the curve is linear; reasoning budgets usually show steep gains early and a plateau late, but that is a general pattern, not a V4.1 measurement.

Benchmarks: What the Card Reports and What It Means

All of the numbers below come from DeepSeek’s model card. They are vendor-reported, run by the vendor’s own harnesses, and not independently reproduced by us. The card’s comparison columns are Opus-5.0, GPT-5.6 Sol, K3, GLM-5.3, DeepSeek V4-Pro and V4-Flash, all at maximum reasoning effort.

Instruct model: selected results

Benchmark V4.1 Flash V4-Flash V4-Pro Best listed in column set
GPQA Diamond 90.9 89.9 92.4 GPT-5.6 Sol 94.1
Codeforces rating 3471 3289 3348 V4.1 Flash
Terminal-Bench 2.1 90.6 82.7 87.9 V4.1 Flash
DeepSWE v1.1 resolved 74.2 54.4 62.7 V4.1 Flash (Opus-5.0 74.0)
CyberGym 88.1 76.7 83.3 V4.1 Flash
AutomationBench 54.8 37.7 43.2 V4.1 Flash
Terminal-Bench 3.0 30.0 7.6 11.8 Opus-5.0 43.3
Terminal-Bench 4.0 31.2 7.0 12.4 Opus-5.0 51.8
HLE (text-only subset) 39.1 37.8 42.7 Opus-5.0 56.3 (full set)

Three readings stand out. First, the generational jump over V4-Flash is large on agentic rows: DeepSWE moves from 54.4 to 74.2, and Terminal-Bench 2.1 from 82.7 to 90.6. Second, V4.1 Flash beats the 1.6T V4-Pro on most agentic rows despite a fraction of the active parameters, which supports the claim that post-training data, not scale, drove the gain. Third, the hardest agentic benchmarks tell a humbler story. On Terminal-Bench 3.0 and 4.0, Opus-5.0 scores 43.3 and 51.8 against 30.0 and 31.2, so a clear gap to the closed frontier remains where tasks are longest.

Base model: where knowledge still favors scale

The base-model table is more revealing about trade-offs, because it removes post-training. On MMLU-Pro, V4.1 Flash scores 74.1 against 73.5 for V4-Pro and 68.3 for V4-Flash. On SimpleQA-Verified, which measures factual recall, it scores 42.3 against 55.2 for V4-Pro. On LongBench-V2 it scores 45.2 against 51.5 for V4-Pro. On MGSM, a multilingual math test, it scores 80.2, below both V4-Flash at 85.7 and V4-Pro at 84.4.

The pattern is coherent. An 8B to 16B active model with a large sparse memory handles reasoning and code well and recalls facts less reliably than a 49B-active model. If your workload is retrieval-free question answering about obscure facts, or long-document comprehension where LongBench-V2 is a proxy, the Pro model leads on the card’s own numbers. The card notes that scores within 0.3 of each other should be treated as equivalent.

Agent scaffold sensitivity

One of the more honest tables on the card shows the same model under different harnesses. On DeepSWE v1.1, results run from 65.5 (OpenCode) to 74.2 (mini-SWE). On Terminal-Bench 2.1, they run from 84.1 (Codex) to 90.6 (DeepSeek Harness Minimal). A spread of roughly nine points on DeepSWE and six on Terminal-Bench from scaffold choice alone is as large as many published model-to-model gaps.

The implication is blunt. A leaderboard difference of two or three points between models, each measured in its own preferred harness, is inside the scaffold noise. Treat the table as evidence that V4.1 Flash belongs in the frontier cluster on agentic coding, not as proof of a strict ranking.

Methodology caveats

  • All figures are self-reported at maximum reasoning effort with temperature 1.0 and top-p 0.95; production settings may differ.
  • Code-agent benchmarks use a 1M-token context window, which stresses the very cache design the model advertises; short-context workloads will not see the same advantage.
  • The DeepSWE and Terminal-Bench runs use N=8 and N=3 samples per task respectively, so small differences are noisy.
  • Benchmark contamination is a known risk for newer public sets, and the card does not publish decontamination details we could review.
  • Independent evaluations (LMArena-style human preference, third-party agent leaderboards) were not available to us at writing; check them before betting a product on these numbers.

Access and Deployment

DeepSeek V4.1 Flash API and self-hosted serving path with KV cache reuse

Figure 4: Request flow through a hosted API or a self-hosted vLLM or SGLang stack, with prefix-cache reuse skipping the prefill step.

The hosted API

DeepSeek’s pricing page lists V4.1 Flash under the model ID deepseek-flash, with a 1M context and a 384K maximum output. Per million tokens, it lists input at $0.003 for a cache hit and $0.15 for a cache miss, and output at $0.60, all at off-peak rates. Peak rates are double: $0.006, $0.30 and $1.20. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays.

The pricing page also states that the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired and that requests under those names are served by V4.1 Flash. A third-party guide reports that the API is OpenAI-compatible at https://api.deepseek.com with an Anthropic-compatible endpoint at /anthropic; confirm both against the current official docs before wiring production traffic. Prices change, and these are the figures as of 3 October 2026.

A worked example (illustrative): an agent session sends a 200K-token context on each of 50 turns and produces 2K output tokens per turn. That is 10M input tokens and 100K output tokens. If 95 percent of the input hits the cache at off-peak rates, the bill is 9.5M x $0.003 per M ($0.0285) plus 0.5M x $0.15 per M ($0.075) plus 0.1M x $0.60 per M ($0.06), about $0.16. With no cache hits, the input alone is $1.50 and the total about $1.56. Cache hit rate, not list price, is the dominant variable, and it is shaped by how you structure prompts so the stable prefix comes first.

Open weights and the MIT license

The repository and the weights are licensed under the MIT License, per the model card. That is the most permissive mainstream license in use for frontier-scale weights: commercial use, modification and redistribution are allowed with a copyright notice. There is no usage-based carve-out in the card. As always, read the LICENSE file in the repository yourself and have counsel review it if you redistribute fine-tunes.

The card lists tensor types BF16, F32, F8_E4M3 and I8. The release ships in FP8 with expert layers already compressed, which is why a further drop to typical 4-bit quantization saves less than it would on a BF16 checkpoint. A hardware review reports about 510 GB for the original FP8 repository, and that review lists community MLX builds at 167 GB for 2-bit and 427 GB for a mixed 4/8-bit build. Those are third-party figures, not from DeepSeek.

Self-hosting reality

The same review measured a 2-bit MLX build running at 9.46 tokens per second with greedy decoding on a Mac Studio with an M3 Ultra and 256 GB of unified memory, peaking at about 167 GiB of process memory. That is a single reported test, with a heavily quantized build, so quality was not characterized here. It proves the model can run on a desktop; it does not prove it should. For a sense of what unified-memory workstations can do, see our analysis of the AMD Ryzen AI Max Pro 400 local inference workstation.

For a server deployment the card lists vLLM, SGLang and Docker Model Runner as supported inference paths, and points to an inference folder for weight conversion. With about 510 GB of FP8 weights, plan on a multi-GPU node: eight 80 GB GPUs hold 640 GB, which is tight once KV cache and activations are included, so 141 GB or 192 GB class parts are more comfortable. That sizing is my estimate, not a published requirement, and DeepSeek does not list a minimum VRAM figure on the card.

The prompt format also needs attention. The release does not ship a Jinja chat template. Instead it provides a Python reference encoder and a Rust toolkit called deepseek-recipe that converts Messages, Chat Completions and Responses API requests into the model’s prompt format. If your serving stack assumes a Hugging Face chat template, expect to integrate the encoder or use the maintained toolkit, and test tool-calling and thinking-mode round trips before launch.

Recommended sampling is temperature 1.0, top-p 0.95 or 1.0, a 1M-token context window, and a max_tokens of at least 256K for reasoning-heavy requests. Those are unusually large generation budgets; make sure your gateway and client timeouts allow them.

Trade-offs, Gotchas, and What Goes Wrong

Total memory is still enormous. A smaller KV cache does not shrink the weights. At roughly 510 GB in FP8, with the Engram table and 384 experts per layer resident, you need the whole checkpoint in fast memory even though only a sliver is active per token. MoE saves compute, not capacity. Teams that read “8B active” as “runs like an 8B model” are in for an expensive surprise at provisioning time.

Expert routing creates hot spots. With 6 of 384 experts chosen per token, load across experts is rarely uniform in real traffic. Code-heavy and image-heavy batches can concentrate on a few experts, and in an expert-parallel deployment the busiest device sets your step time. Watch per-expert token counts in production, not just average GPU utilization.

FP4 cache and sparse indexing are approximations. Both the 4-bit cache and the Top-K selection throw information away. The card reports strong aggregate scores, but needle-style retrieval in a 900K-token prompt, or exact-quote recall from deep in a repository, is where an approximate cache fails first. Our overview of long-context benchmarks and effective context explains why a 1M advertised window is not a 1M reliable window. Run RULER-style tests on your own data.

Factual recall is the weak spot. On the card’s own base-model table, SimpleQA-Verified is 42.3 against 55.2 for V4-Pro. An agent that must recall facts without retrieval will hallucinate more than a Pro-class model would. Pair it with retrieval or tools and verify claims that carry consequences.

Prefill is cheap, so prompt bloat becomes tempting. Because the input side is cheap, teams stuff 500K tokens of context where 50K would do. Cost scales linearly, latency still grows with length, and models degrade when the relevant fact is surrounded by noise. Cheap tokens are a reason to be disciplined, not permissive.

Peak-hour pricing and routing changes. The API doubles prices during weekday peak windows, and legacy model names have been rerouted to V4.1 Flash. Pin model IDs, monitor the pricing page, and keep a fallback provider. A third-party report says V4-Pro requests were also rerouted on 14 September; the official page lists V4-Pro as a separate model, so verify what you are actually calling.

Security posture needs work. The model scores 88.1 on CyberGym and 62.8 on SEC-Bench Pro, according to the card, which is a dual-use signal. Open weights under MIT mean safety layers are yours to build. Community builds that strip refusals have already appeared on Hugging Face. If you deploy agents with shell access, sandbox them, restrict network egress and log every tool call. Our piece on AI model supply chain security and provenance covers verifying that the weights you load are the weights DeepSeek published.

Data residency. The hosted API is operated by a company in China. If your data governance forbids that, self-hosting the open weights or using a Western inference provider is the compliant path. This is a procurement fact, not a technical flaw, and it decides many enterprise evaluations before benchmarks are read.

How It Compares

The card’s comparison columns give a basis for a rough decision matrix. The ratings below are my judgment from the vendor-reported table, not independent test results.

Use case V4.1 Flash V4-Pro Closed frontier (Opus-5.0, GPT-5.6 Sol)
Agentic coding in long repos Strong, leads several rows, very cheap input Behind on reported agentic rows Leads the hardest long-horizon rows
Factual Q&A without retrieval Weaker recall on card base metrics Stronger, SimpleQA-Verified 55.2 Strong, not directly compared on base metrics
Self-hosting and control MIT weights, about 510 GB FP8 MIT-class openness, far larger at 1.6T Not available
Cost per agent turn Lowest of the three on listed API prices About 3 to 7x higher listed rates Not listed on the DeepSeek page
Vision and documents Native, DocVQA 95.6 on base Not reported on the card Strong, leads Chartography and BabyVision

The Pro-to-Flash price ratio comes from the API page: output at $1.98 against $0.60 off-peak is 3.3x, and cache-miss input at $0.66 against $0.15 is 4.4x, and cache-hit input at $0.022 against $0.003 is about 7.3x. Hence the matrix says roughly 3 to 7x, depending on cache hit rate. Whether any of that matters depends on whether Pro’s extra recall buys you something measurable.

Practical Recommendations

Start by deciding which of three questions you are answering. If you need the cheapest strong agent loop, call the hosted API and engineer for cache hits. If you need control, privacy or customization, plan a multi-GPU node and budget engineering time for the prompt encoder. If you only need a capable local model on a workstation, this is probably not the one; the 2-bit desktop result is a curiosity, not a deployment plan.

Then measure on your own tasks. The card’s benchmark spread across scaffolds shows that harness choice moves scores as much as model choice, so your own harness is the only fair comparison. Build a 50 to 100 task regression set from real tickets, run it at reasoning effort 30, 60 and 100, and record accuracy, tokens and latency. Choose the cheapest setting that clears your quality bar.

Finally, treat the cache as a first-class cost lever. Put stable instructions, tool definitions and repository context at the front of every prompt, append volatile content at the end, and log your cache-hit ratio daily. Because the cache-hit price is 50 times lower than a miss ($0.003 against $0.15), a drop in hit rate from 95 to 60 percent changes your bill more than any model swap.

Checklist

  • Confirm the current price, model ID and peak windows on the official pricing page before launch.
  • Pin deepseek-flash explicitly and keep a fallback provider.
  • Sweep reasoning effort on your own tasks; do not default to 100.
  • Run long-context retrieval tests at 128K, 512K and 900K on your data.
  • Sandbox tool-using agents and log every call.
  • Monitor per-expert load and cache-hit ratio in production.
  • Verify checksums of downloaded weights and read the LICENSE file.

Frequently Asked Questions

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is an open-weight multimodal mixture-of-experts model from DeepSeek-AI with a 552B-parameter backbone and a one-million-token context. It accepts text and images and generates text. According to the model card it activates 8B parameters per token in prefill and 16B in decode, and the weights are released under the MIT License. It was reported as released on 10 September 2026.

How many parameters does DeepSeek V4.1 Flash have?

The model card states a 552B backbone, with 8B active during prefill and 16B during decode, using 1 shared and 384 routed experts per layer with 6 routed experts activated. It also describes a 196B-parameter Engram memory accessed sparsely. The card does not give a single combined total, and aggregator figures that differ from 552B are not confirmed by the primary source.

Is DeepSeek V4.1 Flash really open source?

The weights and repository are MIT-licensed, which permits commercial use and modification. Strictly speaking this is open weights: the training data and full training code are not published. DeepSeek does release inference code, a prompt encoder and the deepseek-recipe toolkit. For most builders MIT weights are what matter, but you cannot reproduce the 45T-token pre-training run from what is public.

How much does the DeepSeek V4.1 Flash API cost?

The official pricing page lists deepseek-flash at $0.15 per million cache-miss input tokens, $0.003 per million cache-hit input tokens and $0.60 per million output tokens off-peak. Peak rates, weekdays 01:00 to 04:00 and 06:00 to 10:00 UTC, are double. Context is 1M with up to 384K output. Prices change, so confirm on the page before budgeting.

What hardware do I need to self-host it?

The original FP8 checkpoint is reported at about 510 GB by a third-party review, so a multi-GPU server node with well over 640 GB of combined memory is the practical starting point. Community 2-bit MLX builds around 167 GB have run on a 256 GB Mac Studio at about 9.5 tokens per second, reported by one reviewer. DeepSeek does not publish an official minimum VRAM figure.

How does the KV cache compare with earlier DeepSeek models?

DeepSeek reports 890 bytes of global KV cache per token, about one quarter of V4-Flash and 437 times smaller than DeepSeek V1. Separately, SWA Bounded Replay cuts the persistent KV footprint to roughly one eighth of V4-Flash. These are vendor figures from the model card. A million-token context therefore needs roughly 890 MB of global cache, before weights and other state.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *