Reviewed by Jonathan West · Updated Sep 21, 2026

DeepSeek V4 Flash vs Pro

Compare benchmarks, pricing, and architecture across both model tiers to build an efficient routing setup.

Reviewed by Jonathan West · Updated Sep 21, 2026

DeepSeek V4.1-Flash is the default choice for almost every production workload because it scores higher than DeepSeek V4 Pro across published benchmarks while costing 4.4 times less on input tokens. DeepSeek released V4.1-Flash on September 10, 2026, featuring 552 billion total parameters, native multimodal vision, and a 2,500-request concurrency ceiling. DeepSeek V4 Pro remains available as a 1.6-trillion parameter model designed for specialized deep reasoning and broader parameter scale.

At Layer3Labs, we build custom agents and workflow integrations for small and mid-sized businesses, where production routing decisions come down to cost per completed run and reliable throughput. DeepSeek originally planned to retire DeepSeek V4 Pro on September 14, 2026, by redirecting traffic to V4.1-Flash. DeepSeek reversed that decision after customer feedback, keeping V4 Pro online under its original pay-as-you-go pricing schedule.

The DeepSeek Application Programming Interface (API) is volatile, with pricing windows and model lifecycles shifting on short notice.

DeepSeek V4.1-Flash vs. DeepSeek V4 Pro: Side-by-Side

DimensionDeepSeek V4.1-FlashDeepSeek V4 Pro
What it isMultimodal Mixture of Experts (MoE) model released September 10, 2026Mixture of Experts (MoE) model, 1.6T total, updated August 13, 2026
Model IDdeepseek-flashdeepseek-v4-pro
Parameters (total / active)552B total; 8B active input, 16B active output, plus 196B Engram memory1.6T total, 49B active
Context window and max output1M context window, 384K max output tokens1M context window, 384K max output tokens
Vision inputSupported with native visual understandingNot supported; text and code only
Input price per 1M (cache miss, peak / off-peak)$0.30 peak / $0.15 off-peak$1.32 peak / $0.66 off-peak
Output price per 1M (peak / off-peak)$1.20 peak / $0.60 off-peak$3.96 peak / $1.98 off-peak
Cache-hit input price$0.006 peak / $0.003 off-peak per 1M tokens$0.044 peak / $0.022 off-peak per 1M tokens
Concurrency limit2,500 concurrent requests500 concurrent requests
Published benchmarks (MMLU-Pro, HumanEval, GSM8K, DeepSWE, Terminal-Bench)74.1, 79.4, 93.0, 74.2, 90.673.5, 76.8, 92.6, 62.7, 87.9
Best forHigh-throughput agent pipelines, coding automation, visual parsing, and general production routingDeep multi-domain reasoning, Think Max algorithmic workloads, and tasks tested on 1.6T weights

Are you one of these vendors? Update your listing


The Two Tiers and the September 2026 Shake-Up

DeepSeek restructured its model catalog on September 10, 2026, by launching DeepSeek V4.1-Flash and reversing a planned retirement of DeepSeek V4 Pro four days later. The DeepSeek V4 family initially debuted in preview on April 24, 2026, offering DeepSeek-V4-Pro at 1.6 trillion total parameters with 49 billion active parameters, alongside DeepSeek-V4-Flash at 284 billion total parameters with 13 billion active parameters. On August 13, 2026, DeepSeek shipped an update for Pro under version DeepSeek-V4-Pro-0813, introducing enhanced agent performance and native Responses API support for Codex integration.

The release of DeepSeek V4.1-Flash on September 10, 2026, changed the line-up. Built on a 552 billion parameter Mixture of Experts (MoE) architecture, V4.1-Flash introduced a Causal Encoder-Decoder structure that activates only 8 billion parameters on input and 16 billion on output. It incorporates Engram conditional memory with 196 billion sparsely accessed parameters, 384 routed experts per MoE layer with 6 activated per token, and 1 shared expert. Weights for V4.1-Flash are open under the Massachusetts Institute of Technology (MIT) license on Hugging Face.

DeepSeek originally announced that requests to the deepseek-v4-pro endpoint would automatically route to V4.1-Flash starting September 14, 2026. However, DeepSeek posted an update to its change log stating that in response to user demand, API services for DeepSeek V4 Pro will continue after September 14, 2026, with unchanged billing. The legacy V4-Flash and experimental V4-Flash-Vision-Exp endpoints were retired and, for now, route to V4.1-Flash.

  • April 24, 2026: DeepSeek-V4 preview launched with Pro (1.6T) and Flash (284B).
  • August 13, 2026: DeepSeek-V4-Pro-0813 update added Codex Responses API support.
  • September 10, 2026: DeepSeek V4.1-Flash released; V4 Pro retirement announced then canceled on user demand.

Capability: Where the Published Scores Put Each Model

Published benchmarks place DeepSeek V4.1-Flash ahead of DeepSeek V4 Pro across software engineering, coding syntax, and standard reasoning evaluations. On the Massive Multitask Language Understanding Pro (MMLU-Pro) benchmark, V4.1-Flash scored 74.1 compared to 73.5 for V4 Pro. On HumanEval coding evaluation, V4.1-Flash reached 79.4 while V4 Pro scored 76.8. On Grade School Math 8K (GSM8K), V4.1-Flash reached 93.0 versus 92.6 for V4 Pro.

The gap widens on specialized developer benchmarks. On DeepSWE version 1.1, which measures real-world software engineering task resolution, V4.1-Flash scored 74.2 while V4 Pro scored 62.7. On Terminal-Bench version 2.1, which tests command-line interaction and bash tool use, V4.1-Flash scored 90.6 against 87.9 for V4 Pro. DeepSeek stated that tests by multiple external parties confirm V4.1-Flash leads V4 Pro in runtime speed, task completion, and execution cost.

DeepSeek V4 Pro retains distinct benchmark strengths derived from its larger 1.6-trillion parameter base. DeepSeek reported a 5-shot MMLU score of 90.1, a 4-shot MATH score of 64.5, and a 1-shot LongBench-V2 score of 51.5 for V4 Pro. In its specialized Think Max reasoning mode, DeepSeek reported V4 Pro achieving 93.5 on LiveCodeBench compared to 91.7 for Google Gemini-3.1-Pro High, and a Codeforces competitive programming rating of 3206 compared to 3168 for OpenAI GPT-5.4 xHigh. On SimpleQA-Verified, V4 Pro reached 57.9 compared to 75.6 for Gemini-3.1-Pro High.

V4.1-Flash outperforms V4 Pro on coding and terminal tasks, but V4 Pro retains higher world-knowledge capacity from its 1.6-trillion parameter pretraining base.

Cost: The Rate Cards Side by Side with a Worked Example

DeepSeek V4.1-Flash costs 4.4 times less for uncached input tokens and 3.3 times less for output tokens compared to DeepSeek V4 Pro during peak operating hours. The DeepSeek API operates on a prepaid balance system with no monthly subscriptions or seat licenses. Prices differ based on peak and off-peak windows. DeepSeek defines peak hours as 01:00 to 04:00 and 06:00 to 10:00 Coordinated Universal Time (UTC), Monday through Friday. All remaining hours bill at an off-peak discount of 50 percent.

For cache-miss input tokens, deepseek-flash bills at $0.30 per million tokens during peak hours and $0.15 during off-peak hours. In contrast, deepseek-v4-pro bills at $1.32 per million tokens during peak hours and $0.66 off-peak. For output tokens, Flash bills at $1.20 peak and $0.60 off-peak per million tokens, whereas Pro bills at $3.96 peak and $1.98 off-peak per million tokens. Cached input tokens bill at $0.006 peak and $0.003 off-peak for Flash, compared to $0.044 peak and $0.022 off-peak for Pro. Confirm current schedules on the DeepSeek pricing documentation because DeepSeek adjusts windows periodically. Full rate-card context is on our DeepSeek pricing guide.

Consider a production workload processing 10 million input tokens and generating 2 million output tokens during peak hours with no cache hits. On DeepSeek V4.1-Flash, 10 million input tokens cost $3.00 (10 multiplied by $0.30) and 2 million output tokens cost $2.40 (2 multiplied by $1.20), totaling $5.40. On DeepSeek V4 Pro, 10 million input tokens cost $13.20 (10 multiplied by $1.32) and 2 million output tokens cost $7.92 (2 multiplied by $3.96), totaling $21.12. Running this batch on Flash saves $15.72, reducing the total API invoice by 74.4 percent.

  • Flash input token price: $0.30 peak / $0.15 off-peak per million tokens.
  • Pro input token price: $1.32 peak / $0.66 off-peak per million tokens.
  • Flash output token price: $1.20 peak / $0.60 off-peak per million tokens.
  • Pro output token price: $3.96 peak / $1.98 off-peak per million tokens.
Even with an 80 percent cache-hit rate, Flash costs roughly 3.6 times less overall than Pro on identical input and output volumes.

Speed, Throughput, and Concurrency Limits

DeepSeek V4.1-Flash supports a concurrency limit of 2,500 simultaneous requests, which is five times higher than the 500 concurrent request limit assigned to DeepSeek V4 Pro. DeepSeek does not publish specific per-minute request caps or per-minute token volume limits on its official rate cards. The concurrency limit represents the only hard quota DeepSeek documents for API accounts on platform.deepseek.com.

The architecture of V4.1-Flash explains this throughput advantage. By using only 8 billion active parameters during prompt evaluation and 16 billion active parameters during token generation, the model places substantially lower load on GPU compute clusters. DeepSeek documents a Key-Value (KV) cache memory footprint of 890 bytes per token for V4.1-Flash. Compared to previous generation checkpoints, DeepSeek reports that V4.1-Flash uses one-fourth the High Bandwidth Memory (HBM) and one-eighth the Solid State Drive (SSD) storage.

Teams deploying autonomous background agents or high-volume scrapers will hit the 500-concurrency ceiling on DeepSeek V4 Pro quickly. Setting up automated pipelines on Flash allows 2,500 parallel workers to query the model without exceeding the published concurrency ceiling.

  • Concurrency limits: 2,500 requests on Flash versus 500 requests on Pro.
  • KV cache footprint: 890 bytes per token on V4.1-Flash.
  • Active parameters: 8B input and 16B output on Flash versus 49B active on Pro.

Vision and Multimodal Work: Flash Only

DeepSeek V4.1-Flash includes native multimodal image understanding, while DeepSeek V4 Pro processes text input only. DeepSeek previously offered image processing through an experimental checkpoint named deepseek-v4-flash-vision-exp, released on August 21, 2026. On September 10, 2026, DeepSeek retired that experimental model and integrated its visual capabilities directly into the core V4.1-Flash checkpoint.

DeepSeek's rate card lists no vision input for deepseek-v4-pro, so image payloads cannot be sent to that model. Organizations that extract tabular figures from invoices, analyze chart graphics, or parse mobile interface screenshots must use V4.1-Flash. Attempting to use V4 Pro for multimodal pipelines requires routing images through an external optical character recognition tool before passing extracted plain text to the model.

Do not attempt to pass image data to DeepSeek V4 Pro; multimodal visual processing is supported exclusively on DeepSeek V4.1-Flash.

A Routing Table: Task to Model with Operational Reasoning

Routing production tasks between tiers depends on whether a workload requires visual inputs, coding benchmarks, or the 1.6-trillion parameter base of DeepSeek V4 Pro. For standard software development, code generation, refactoring, and terminal commands, DeepSeek V4.1-Flash provides better performance. Its 74.2 score on DeepSWE and 90.6 on Terminal-Bench exceed V4 Pro while operating at lower latency and lower token rates.

Document ingestion, customer ticket triage, and conversational chatbots should route to V4.1-Flash by default. The lower pricing ($0.30 per million input tokens peak) prevents runaway costs when users send long conversation threads into the 1-million token context window. In contrast, DeepSeek V4 Pro serves workloads that demand extensive world knowledge, edge-case trivia, or high-stakes reasoning where an engineering team has verified that the 1.6-trillion parameter base produces better answers.

DeepSeek V4 Pro is also appropriate when using Think Max mode for competitive coding competitions or complex mathematical theorem proofs. For Think Max requests, DeepSeek recommends maintaining a context window allocation of at least 384,000 tokens, using default sampling settings of temperature 1.0 and top_p 1.0.

  • Route software engineering, code generation, and terminal scripts to Flash.
  • Route document processing, visual charts, and OCR tasks to Flash.
  • Route high-volume data classification and customer support bots to Flash.
  • Route complex algorithmic proofs and deep domain reasoning to Pro with Think Max.

What to Pin in Your Configuration and Watch in the Change Log

Engineering teams should pin explicit model strings in their code and monitor DeepSeek release announcements weekly to safeguard against sudden API routing changes. Avoid using generic aliases if your system depends on specific model behaviors. For V4.1-Flash, specify the model ID deepseek-flash. For V4 Pro, specify deepseek-v4-pro.

DeepSeek V4.1-Flash introduces a continuously controllable reasoning-effort parameter configured as an integer from 1 to 100. This parameter allows developers to throttle the volume of internal thinking tokens generated before the model produces its final response. Teams building latency-sensitive tools can set reasoning effort lower to reduce turnaround time, or raise it toward 100 for difficult analytical queries.

Because DeepSeek reversed its deprecation notice for V4 Pro within four days in September 2026, teams should treat the model catalog as subject to rapid revisions. Track the official DeepSeek change log to catch any future deprecation dates or pricing window shifts before they affect live production traffic.

  • Pin deepseek-flash for V4.1-Flash and deepseek-v4-pro for V4 Pro.
  • Use the reasoning-effort setting (integer 1 to 100) on Flash to control thinking latency.
  • Audit the DeepSeek API change log regularly to track model deprecation updates.

The Verdict

Route new production workloads to DeepSeek V4.1-Flash by default. DeepSeek V4.1-Flash delivers higher published benchmark scores in software engineering and terminal interaction, native visual understanding, five times higher concurrency, and a 4.4 times lower input cost during peak and off-peak periods. For engineering teams operating autonomous agents or high-volume data pipelines, Flash offers a compelling balance of cost and performance.

Retain DeepSeek V4 Pro selectively for tasks that demand its 1.6-trillion parameter knowledge base or specialized Think Max reasoning mode. Applications in competitive coding, deep multi-domain question answering, or complex mathematical derivation may benefit from its scale, provided your team validates that gain through internal empirical testing. Neither tier is suited for teams that require strict United States data residency guarantees or dedicated enterprise support contracts; those organizations should explore managed deployments on third-party cloud platforms.

What would change this recommendation is if DeepSeek reduces V4 Pro pricing to parity with Flash or releases an updated V4.2 Pro checkpoint that establishes a commanding lead on agentic coding benchmarks. Until then, audit your current token logs, pin the deepseek-flash model identifier for your automated routines, and benchmark your most complex reasoning prompts against deepseek-v4-pro to measure if the extra parameter scale justifies the 4.4 times input price difference.

Sources & Disclaimer

Researched from primary DeepSeek documentation and public regulator sources. Pricing and availability are accurate as of Sep 21, 2026 and can change — confirm current terms with each vendor before you buy.

Frequently Asked Questions

  • DeepSeek V4.1-Flash costs 4.4 times less on cache-miss input and 3.3 times less on output than DeepSeek V4 Pro. For uncached input tokens, Flash costs $0.30 per million tokens during peak hours and $0.15 during off-peak hours, while Pro costs $1.32 peak and $0.66 off-peak. For output tokens, Flash costs $1.20 peak and $0.60 off-peak per million tokens, compared to $3.96 peak and $1.98 off-peak for Pro. Cached input tokens cost $0.006 peak on Flash versus $0.044 peak on Pro. Peak hours run from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, with off-peak providing a 50 percent discount. Confirm rates on the DeepSeek pricing documentation.
  • Yes, on published benchmark scores, DeepSeek V4.1-Flash performs better than DeepSeek V4 Pro on software engineering and reasoning tasks while operating at lower cost. On the MMLU-Pro benchmark, V4.1-Flash scored 74.1 compared to 73.5 for V4 Pro. On the DeepSWE software engineering benchmark, V4.1-Flash achieved 74.2 while V4 Pro scored 62.7. V4.1-Flash also includes native multimodal vision support, which V4 Pro lacks. However, V4 Pro retains a 1.6-trillion parameter base, providing broader general world knowledge for non-coding queries.
  • Yes, DeepSeek V4.1-Flash is purpose-built for agentic coding tasks. On the Terminal-Bench 2.1 benchmark, which tests bash execution and command-line tool use in developer agents, V4.1-Flash scored 90.6, surpassing the 87.9 score achieved by V4 Pro. On the DeepSWE v1.1 real-world coding benchmark, Flash scored 74.2 versus 62.7 for Pro. Its 2,500-request concurrency ceiling and low KV cache footprint of 890 bytes per token allow engineering teams to run parallel developer agents cost-effectively.
  • The most powerful model depends on whether raw parameter scale or task-specific performance is evaluated. In terms of parameter count and world-knowledge breadth, DeepSeek V4 Pro is the largest model, containing 1.6 trillion total parameters and 49 billion active parameters, with an MMLU 5-shot score of 90.1. In terms of coding benchmark performance, mathematical problem solving, execution speed, and multimodal capabilities, DeepSeek V4.1-Flash is the more effective model, outperforming V4 Pro on MMLU-Pro, HumanEval, GSM8K, and DeepSWE.
  • DeepSeek did not publish comparative benchmark evaluations between DeepSeek V4 Pro and Claude Opus 4.8. DeepSeek only published cross-lab comparisons for V4 Pro in its Think Max reasoning mode against Google Gemini-3.1-Pro and OpenAI GPT-5.4. In those reports, DeepSeek reported V4-Pro-Max scoring 93.5 on LiveCodeBench compared to 91.7 for Gemini-3.1-Pro High, and achieving a 3206 Codeforces rating compared to 3168 for GPT-5.4 xHigh. Teams comparing DeepSeek against Anthropic models should run internal evals on their specific prompts.

Need help optimizing your model routing and token costs?

Book a free 30-minute AI workflow audit with Layer3Labs. We will review your prompt architecture, benchmark your production tasks across DeepSeek tiers, and set up reliable routing to reduce API spend.

Book Your Free Audit