E2EE Qwen 3.6 35B A3B FP8
e2ee-qwen3-6-35b-a3b- 🆕 First open-weight model in Alibaba's Qwen3.6 series
- 🧠 Mixture-of-experts: 35B total, ~3B active per token
- 🔧 Gated-delta-networks MoE, 256 experts, 8 routed plus 1 shared
- 💬 Switchable thinking and non-thinking inference modes
- 🎯 Tuned for agentic coding and repository-level reasoning
- 🔒 Served inside a Trusted Execution Environment with hardware attestation
- 📏 32K context window in this deployment; FP8 quantization
- ⚡ Compact active footprint enables fast, lower-cost inference
Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research, developing the Qwen…
Explore 34 more models by Alibaba Group →Qwen 3.6 35B A3B FP8 is Alibaba's first open-weight release in the Qwen3.6 line, a sparse mixture-of-experts model with 35 billion total parameters but only about 3 billion active per token. Architecturally it shares the gated-delta-networks MoE design of the 3.5 generation, routing 8 of 256 experts plus one shared expert each step, with switchable thinking and non-thinking modes. This Venice deployment runs the official FP8 checkpoint inside a Trusted Execution Environment, exposing hardware attestation so enclave identity and configuration can be independently verified.
Compared with its direct predecessor Qwen 3.5 35B A3B, Alibaba reports that Qwen3.6 improves agentic coding and reasoning, and adds a thinking-preservation capability for steadier long agent runs. Alibaba also states the model rivals its larger dense Qwen 3.6 27B sibling on several coding benchmarks despite the smaller active count. These figures are vendor self-reported and not yet independently reproduced.
Within the same Qwen family on this catalog, it sits alongside the earlier Qwen3 30B A3B enclave model and an uncensored 3.6 variant. It is licensed Apache 2.0, with capabilities spanning reasoning, code, function calling, and web search.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| Fire Ant 🔥🐜 0xbe05…bc5d | 50.00 | gated | $0.113 | $0.0358 | $0.77 | chat,coding,math | — |
| antseed-neon-puma-944e 0x6650…944e | 35.09 | gated | $0.0621 | $0.06 | $0.4024 | chat,coding,math | openai-chat-completions |
| Leftermute 0x388b…5389 | 28.39 | gated | $0.1162 | $0.1162 | $0.3485 | chat,coding,json | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = buyer trust score (0-100, the AntSeed SDK's own formula). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.