DeepSeekDeepSeek·text

DeepSeek V4.1 Flash

CodeVisionReasoningWeb searchFunction calling
Advertised as deepseek-v4-1-flashdeepseek-v4.1-flashdeepseek/deepseek-v4-1-flashdeepseek/deepseek-v4.1-flash
Quick reference
DeepSeek V4.1 Flash — TLDR
  • 🧠 552B-parameter multimodal Mixture-of-Experts, only 8–16B parameters active per token
  • 📏 One-million-token context window, extended from 64K training via YaRN
  • 🆕 New Causal Encoder–Decoder design: 20 encoder plus 20 decoder layers
  • 👁️ Native vision via DeepSeek-ViT encoder trained from pre-training onward
  • ⚡ Prefill activates 8B parameters, decode 16B, aiding input-heavy agents
  • 🔧 Tuned for code agents, function calling and web search workflows
  • 🔒 MIT-licensed open weights, FP8 quantization on this deployment
  • 🏢 Released September 2026 by DeepSeek, served as the "deepseek-flash" endpoint
💰 Best price on AntSeed
$0.0007 / $0.0027−98%
per 1M · cheapest in / out
📏 Context
1M tokens
🐜 Sellers
9
advertising on AntSeed
Provider

DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…

Explore 7 more models by DeepSeek →
About this model

DeepSeek V4.1 Flash is the smallest member of DeepSeek's new architecture family and the company's cost-efficient multimodal tier, released in September 2026 under an MIT license. It is a Mixture-of-Experts model with 552B backbone parameters that activates roughly 8B parameters per input token and 16B per output token, and it handles up to one million tokens of context.

The clearest change from DeepSeek V4 Flash 0731 and the original V4 Flash 0423 is architectural. Where the V4 preview series used a conventional decoder stack with hybrid compressed/heavily-compressed attention, V4.1 Flash adopts a Causal Encoder–Decoder layout: a 20-layer causal encoder feeds a 20-layer decoder, and the decoder's global KV cache is projected from the final encoder states rather than each decoder layer's own hidden states. DeepSeek describes this as targeting a higher capability ceiling, faster inference and higher throughput while scaling to larger models.

Multimodality is also native here. Earlier Flash checkpoints were text models, with vision arriving as a separate experimental variant; V4.1 Flash instead trains a from-scratch vision encoder and MLP projector jointly with text from the start of language-model pre-training. Routing uses one shared plus 384 routed experts, six active per token.

Compared with DeepSeek V4 Pro and DeepSeek V3.2, V4.1 Flash trades total scale for sparsity, aiming at long-context agentic and coding workloads where input volume dominates cost. DeepSeek evaluates it on code-agent suites using its own harness at full 1M context.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
deepseek.comDeepSeek | Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.· deepseek.comapi-docs.deepseek.comChange Log | DeepSeek API Docs· api-docs.deepseek.comhuggingface.codeepseek-ai/DeepSeek-V4.1-Flash · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Usage on AntSeed
Tokens served
2.70B
input + output
Requests
24,385
settled calls
Buyers
41
distinct, on this model
Sellers used
9
of 9 advertising
Settled
$23.71
gross USDC, this model
Sellers serving DeepSeek V4.1 Flash (9)compare on the network explorer →
SellerReputation↓RoutingInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = buyer trust score (0-100, the AntSeed SDK's own formula). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.