E2EE DeepSeek V4 Flash
e2ee-deepseek-v4-flash- 🔒 Runs in a Trusted Execution Environment with hardware attestation
- 🧠 Mixture-of-Experts: 284B total, ~13B activated per token
- 📏 One-million-token context window for long documents
- 🔧 Function calling, code optimization, and integrated web search
- ⚡ FP8 quantization tuned for fast, low-latency inference
- 🆕 Hybrid attention (CSA + HCA) cuts long-context cost
- 📚 MIT licensed with open weights on Hugging Face
- 💬 Supports both thinking and non-thinking modes
DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by Liang. DeepSeek…
Explore 7 more models by DeepSeek →DeepSeek V4 Flash is the efficiency-focused member of DeepSeek's V4 series, a Mixture-of-Experts language model with 284 billion total parameters and roughly 13 billion activated per token, paired with a one-million-token context window. This catalog entry wraps the model in a Trusted Execution Environment, exposing hardware attestation evidence so the enclave's identity and configuration can be independently verified — a privacy layer on top of the standard weights.
Within the family, it sits below DeepSeek V4 Pro, the larger flagship that shares the same 1M context but runs at higher cost. Both are distinct from the earlier DeepSeek V3.2 generation. It also mirrors the non-enclave DeepSeek V4 Flash checkpoint.
The generational gains are architectural. DeepSeek introduced a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention to reduce long-context memory and compute. Post-training used a two-stage paradigm: cultivating domain-specific experts via supervised fine-tuning and GRPO reinforcement learning, then consolidating them through on-policy distillation.
The model supports both thinking and non-thinking modes and is released under the MIT license with open weights. It targets coding, reasoning, and agentic workflows through function calling and integrated web search, served in FP8 for lower-latency inference.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 50.00 | gated | $0.1536 | $0.0321 | $0.3148 | chat,coding,fast,tools | — |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = buyer trust score (0-100, the AntSeed SDK's own formula). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.