E2EE Kimi K3
e2ee-kimi-k3-p- - 🧠 2.8T-parameter Mixture-of-Experts with 104B active parameters
- - 📏 One-million-token context window for whole-repo and long-document work
- - 👁️ Native multimodality: text, images and video in one model
- - 🔧 Kimi Delta Attention, Attention Residuals and Stable LatentMoE architecture
- - 🎯 Built for long-horizon agentic coding and knowledge work
- - ⚡ Configurable reasoning effort; results reported at max effort
- - 🔒 Served end-to-end encrypted, with function calling and web search
- - 📚 Open weights released under Moonshot's own Kimi K3 License
Moonshot is an AI research lab known for developing the Kimi family of large language models. The organization has gained recognition for building capable reasoning-oriented models, with the Kimi line representing its flagship series of text generation systems.
Explore 6 more models by Moonshot →Kimi K3 is positioned as Moonshot AI's flagship open-weight release — a 2.8-trillion-parameter Mixture-of-Experts system that Moonshot describes on its Hugging Face model card as the first open "3T-class" model. NVIDIA's hosted model documentation lists 104B activated parameters, a 160K vocabulary, and a 1,048,576-token input context. This catalog entry is the end-to-end-encrypted deployment of that model, alongside the standard Kimi K3 endpoint and the latency-tuned Kimi K3 Fast.
Architecturally it is a clear break from the K2 line represented by Kimi K2.6, Kimi K2.6 and Kimi K2.5. K3 is built on Kimi Delta Attention and Attention Residuals, with Stable LatentMoE routing and a 401M-parameter MoonViT-V2 vision encoder that makes image and video understanding native rather than a bolted-on adapter.
The stated design target is long-horizon work: navigating large repositories, iterating against logs, tests and screenshots, and producing research artifacts such as interactive dashboards and visualizations. Compared with the code-specialized Kimi K2.7 Code, K3 is a general frontier model rather than a coding-focused checkpoint.
Practical notes: the model is trained with preserved thinking history, so multi-turn and tool-calling applications must return prior assistant messages including reasoning content and tool calls. Reasoning effort is configurable, and Moonshot's published evaluations use the maximum setting.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 50.00 | gated | $3.0928 | $0.3093 | $15.464 | chat,coding,reasoning,tools | — |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = buyer trust score (0-100, the AntSeed SDK's own formula). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.