DeepSeek V4.1 Flash
deepseek-v4-1-flashdeepseek-v4.1-flashdeepseek/deepseek-v4-1-flashdeepseek/deepseek-v4.1-flash- 🧠 552B-parameter multimodal Mixture-of-Experts, only 8–16B parameters active per token
- 📏 One-million-token context window, extended from 64K training via YaRN
- 🆕 New Causal Encoder–Decoder design: 20 encoder plus 20 decoder layers
- 👁️ Native vision via DeepSeek-ViT encoder trained from pre-training onward
- ⚡ Prefill activates 8B parameters, decode 16B, aiding input-heavy agents
- 🔧 Tuned for code agents, function calling and web search workflows
- 🔒 MIT-licensed open weights, FP8 quantization on this deployment
- 🏢 Released September 2026 by DeepSeek, served as the "deepseek-flash" endpoint
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 7 more models by DeepSeek →DeepSeek V4.1 Flash is the smallest member of DeepSeek's new architecture family and the company's cost-efficient multimodal tier, released in September 2026 under an MIT license. It is a Mixture-of-Experts model with 552B backbone parameters that activates roughly 8B parameters per input token and 16B per output token, and it handles up to one million tokens of context.
The clearest change from DeepSeek V4 Flash 0731 and the original V4 Flash 0423 is architectural. Where the V4 preview series used a conventional decoder stack with hybrid compressed/heavily-compressed attention, V4.1 Flash adopts a Causal Encoder–Decoder layout: a 20-layer causal encoder feeds a 20-layer decoder, and the decoder's global KV cache is projected from the final encoder states rather than each decoder layer's own hidden states. DeepSeek describes this as targeting a higher capability ceiling, faster inference and higher throughput while scaling to larger models.
Multimodality is also native here. Earlier Flash checkpoints were text models, with vision arriving as a separate experimental variant; V4.1 Flash instead trains a from-scratch vision encoder and MLP projector jointly with text from the start of language-model pre-training. Routing uses one shared plus 384 routed experts, six active per token.
Compared with DeepSeek V4 Pro and DeepSeek V3.2, V4.1 Flash trades total scale for sparsity, aiming at long-context agentic and coding workloads where input volume dominates cost. DeepSeek evaluates it on code-agent suites using its own harness at full 1M context.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| ▲ Apex Ant 0x73b4…e736 | 71.90 | #2 | $0.16 | $0.0032 | $0.65 | chat,fast,open-source,cheap,coding,privacy,reasoning,long-context,agents,vision,multimodal | openai-chat-completions |
| D5V1N2 0xd5e7…7be0 | 67.45 | #3 | $0.18 | $0.004 | $0.72 | chat,reasoning,coding,vision,large-context,deepseek,router,fallback | openai-chat-completions |
| Super Seeder 0xd19f…41f3 | 63.77 | #4 | $0.1856 | $0.0037 | $0.7425 | chat,coding,reasoning,vision,multimodal,tools,long-context,cheap | openai-chat-completions |
| Edith AI 0xb269…b1a6 | 61.93 | #1 | $0.003 | $0.0006 | $0.5924 | chat,coding,reasoning | openai-chat-completions |
| ZLKPro-Api 0x0b0b…f446 | 58.35 | gated | $0.10 | $0.012 | $0.236 | agent,chat,text,reasoning,research,smart,long-context,multimodal,visual | openai-chat-completions |
| bartly.eth64.de 0x666e…4666 | 53.43 | gated | $0.0007 | $0.00 | $0.0027 | chat,coding,fast,tools | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 50.00 | gated | $0.305 | $0.0061 | $1.22 | chat,coding,fast,tools | — |
| Best 0x472f…69fd | 49.76 | gated | $0.0007 | $0.00 | $0.0027 | chat,coding | openai-chat-completions |
| Hana Gateway ✅ 0x4ae1…117b | 45.41 | gated | $0.02 | $0.005 | $0.04 | chat,coding,fast,tools | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = buyer trust score (0-100, the AntSeed SDK's own formula). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.