Long Context · Efficient AI · Language Models · Mixture of Experts
Kimi K3 vs DeepSeek V4 vs Qwen3.8-Next: which 1M-context architecture wins?
Four 2026 frontier open models target 1M context four ways: Kimi K3 runs 3:1 delta attention, DeepSeek V4 compresses sparse KV, Qwen3.8-Next runs 3:1 GDN, Nemotron 3 Ultra leans on Mamba-2.
Quick answer
The four big open-model reports of mid-2026 (Kimi K3, DeepSeek V4, Qwen3.8-Next and Nemotron 3 Ultra) are four answers to the same question: how do you keep per-token cost flat when the context window goes to one million tokens? They disagree about where to pay. Kimi K3 replaces most attention layers with a gated delta-rule linear attention and spends its savings on more experts and agentic post-training. DeepSeek V4 keeps attention but compresses the KV cache into grouped entries and attends sparsely on top. Qwen3.8-Next is the most radical on paper: 6B activated parameters plus 51B host-memory n-gram tables, matched against a 397B predecessor. Nemotron 3 Ultra goes furthest toward linear-time decoding with a Mamba-2-heavy stack and accepts a few accuracy points for up to 5.9x throughput. All four train most weight matrices with Muon, which tells you the optimizer argument is settled at the frontier.
Key numbers
Every row below comes from the model’s own technical report, run in its own harness. Treat cross-model rows as “same benchmark name, different harness” unless stated otherwise.
| Dimension | Kimi K3 | DeepSeek V4 | Qwen3.8-Flash-Next | Nemotron 3 Ultra |
|---|---|---|---|---|
| Total / activated params | 2.78T / 104.2B | Pro 1.6T/49B, Flash 284B/13B | 125B / 6B (+51B host-memory n-gram tables) | 550B / 55B |
| Token-mixing design | 3 Kimi Delta Attention : 1 gated MLA per block | Compressed Sparse Attention + Heavily Compressed Attention interleaved, with sliding window and attention sink | 3 Gated DeltaNet : 1 attention per block; Qwen Sparse Attention at continued pretraining | Mostly Mamba-2 state-space layers + few attention layers (64 Q heads, 2 KV heads) |
| Efficiency headline | ~2.5x scaling efficiency over K2 (fitted scaling law) | V4-Pro: 27% of V3.2’s FLOPs, 10% of its KV cache at 1M; Flash 10% / 7% | ~1/9 the training FLOPs of Qwen3.7-Plus, 1/3 activated params, 1/3 tokens | 5.9x / 4.8x / 1.6x decode throughput vs GLM-5.1 / Kimi-K2.6 / Qwen-3.5 at 8K-in/64K-out |
| Long-context score | 1M training context; AA-LCR 74.7 (best in its table) | MRCR 1M 83.5, CorpusQA 1M 62.0 (trails Opus-4.6 by 9–10 pts) | RULER 512K–1M 93.00 with QSA vs 90.08 full attention | RULER 1M 94.7 (best in its table) |
| Coding / agentic headline | Terminal-Bench 2.1 88.3, BrowseComp 91.2 (both best in its table) | V4-Pro-Max: SWE Verified 80.6, LiveCodeBench 93.5, Codeforces 3206 | No post-trained results; base MMLU-Pro 73.23 vs Qwen3.7-Plus 70.90 | SWE-Bench Verified 70.7, Terminal Bench 2.1 56.4 (trails open peers by 5–17 pts) |
| Headroom honesty | Max-effort scoring vs proprietary models at their max settings; several agentic numbers run in Kimi’s own harness | Report calls itself a preview; closed-model baselines in DeepSeek’s harness | Architecture ablation, not a model card; component claims from 25B-A3B runs | NVIDIA’s own max-throughput measurements on NVIDIA hardware; pretraining diverged twice, run cut to 20T tokens |
Two numbers deserve a second look because they are the most quoted and the least comparable. Qwen’s “1/9 the training FLOPs” is derived from the parameter and token ratios, not a reported compute total for either run. Kimi’s “2.5x scaling efficiency” is a fitted scaling-law comparison against Moonshot’s own K2, not a controlled ablation; the report itself credits data and recipe jointly. DeepSeek’s 27% and 10% figures are FP8-equivalent FLOP estimates and cache sizes, not measured latency. Nemotron’s 5.9x is a serving measurement with speculative decoding enabled and the best number chosen per baseline.
How the four designs differ
What gets compressed. The three linear-attention models compress the prefix into a fixed-size state: K3’s Kimi Delta Attention adds a channel-wise forget gate to the delta-rule recurrence so the state can overwrite stale content instead of accumulating it; Qwen’s Gated DeltaNet does the same job in a 3:1 interleave with one full-attention layer kept for exact retrieval; Nemotron’s Mamba-2 layers have constant per-step decode cost in sequence length. DeepSeek instead compresses the cache: groups of consecutive tokens are merged into single KV entries (CSA), some layers compress much larger groups (HCA), and a Lightning Indexer lets each query attend only to top-k compressed entries. The philosophical split is “forget through a gate” versus “keep everything, but smaller and indexed.”
What the saved budget buys. K3 reinvests in scale and depth: Attention Residuals let each block read summaries of all previous blocks, and a 896-expert Stable LatentMoE with 16 active experts runs at 2.8T total parameters. DeepSeek reinvests in training: 32–33T tokens and manifold-constrained hyper-connections to keep a deep MoE stable. Qwen reinvests in the opposite direction: the 51B n-gram embedding tables live in host memory and are prefetched, so a 125B model with 6B active parameters reports matching a 397B-A17B predecessor on 8 of 14 pretraining benchmarks. Nemotron reinvests in serving: NVFP4 weights serve both W4A16 and W4A4 from one checkpoint, and two shared-weight multi-token-prediction heads double as speculative-decoding drafters and RL-rollout accelerators.
Where each one gives something up. K3’s AttnRes only runs over block-level summaries (blocks of 12 layers), which the report says recovers most of the benefit; “most” is doing work there. DeepSeek’s V4-Pro still trails Opus-4.6 by 9–10 points on 1M-context recall, so its efficiency claim is about cost per token, not parity of recall. Qwen3.8-Next’s paper contains no post-trained or agentic numbers at all, so the design’s end-to-end value is unproven in the report. Nemotron’s pretraining diverged twice and the token horizon was cut from plan to 20T, with imbalanced experts and residual-stream norms documented but not diagnosed; its headline suites trail the strongest open models by 5–17 points.
When to use which
- Kimi K3 when the workload is agentic and browsing-heavy: it posts the best numbers in its own tables on BrowseComp (91.2), MCPMark-Verified (94.5) and Terminal-Bench 2.1 (88.3), and was the first open model to top WebDev Arena. The cost is a 2.8T-parameter serving footprint.
- DeepSeek V4-Pro when the workload is coding and reasoning at 1M context: SWE Verified 80.6 and LiveCodeBench 93.5 at 27% of V3.2’s per-token FLOPs. V4-Flash is the efficiency pick: 13B activated, GPQA 88.1, SWE Verified 79.0, 10% of V3.2’s FLOPs.
- Qwen3.8-Next as an architectural reference when training budget dominates: the GDN hybrid beat full attention 53.81 to 49.87 and a sliding-window hybrid 51.15 in a controlled 25B-A3B ablation on identical tokens, and Qwen Sparse Attention raised long-context scores (RULER 512K–1M 93.00 vs 90.08), evidence that sparse attention does not have to trade recall for cost. As a model choice it is unproven until post-trained numbers exist.
- Nemotron 3 Ultra when serving economics dominate: at 8K-in/64K-out decode-heavy settings it reports 5.9x the throughput of GLM-5.1 and 4.8x of Kimi-K2.6, with RULER 1M 94.7 the best in its table. Accept the 5–17 point gaps on SWE-Bench Verified and BrowseComp if tokens-per-dollar is the metric that matters.
Limits and open questions
No third party has run these four models in one harness, and each report measures efficiency differently (FLOP estimates, fitted scaling laws, derived ratios, serving throughput), and these are not convertible into each other. Cross-model benchmark rows above share a benchmark name, not a harness; even within a report, harness choices matter (Kimi’s agentic numbers run with Kimi Code, DeepSeek’s Table 6 uses 500 interaction steps at 512K context, NVIDIA tunes each baseline’s serving configuration). The Muon convergence is real and worth noting (all four reports use it for most matrices), but each pairs it with different stability machinery (per-head orthogonalization, gated residuals, mHC, NVFP4 rollback recipes), so “everyone uses Muon” does not mean the training-stability problem is solved.
FAQ
Which is better, Kimi K3 or DeepSeek V4?
On each report’s own tables: K3 leads agentic and browsing suites (BrowseComp 91.2 vs V4-Pro-Max 83.4; Terminal-Bench 88.3 on version 2.1 against V4’s closest comparable row of 67.9 on version 2.0), while V4-Pro-Max leads coding (SWE Verified 80.6, LiveCodeBench 93.5) and K3 leads on hard knowledge (GPQA Diamond 93.5 vs 90.1). These rows come from different harnesses, so treat the gap direction, not the exact points, as the signal.
Is Qwen3.8-Next’s sparse attention better than DeepSeek V4’s compressed attention at 1M tokens?
The reports point in different directions. DeepSeek’s compressed approach keeps full retrieval in principle but its V4-Pro still scores MRCR 1M 83.5, trailing Opus-4.6’s 92.9. Qwen’s data argues sparsity can beat full attention on recall (RULER 512K–1M 93.00 vs 90.08). Nemotron’s Mamba-heavy stack posts the best RULER 1M number in its table (94.7) while its reasoning scores lag. The honest answer: at 1M tokens every design loses something, and what it loses depends on which layers you compress.
Why does Nemotron 3 Ultra use Mamba-2 when Kimi K3 and Qwen3.8-Next use delta attention?
All three are fixed-state token mixers with constant per-step cost, but they make different bets. KDA’s channel-wise forget gate lets the state overwrite stale content, Gated DeltaNet does the same inside a 3:1 interleave, and Mamba-2’s SSM cache is far smaller than an FP8 KV cache at the same batch and length, which is exactly where Nemotron’s 5.9x throughput claim comes from. The trade shows up in the scores: the two delta-attention models keep their attention layers for exact retrieval, while Nemotron’s table trails the strongest open models on reasoning suites by 5–17 points.
Why do all four models use the Muon optimizer?
Each report independently found Muon worth the engineering: DeepSeek calls its run the largest disclosed Muon pretraining, K3 orthogonalizes per head rather than per matrix, Qwen kept the MoE router on AdamW because Muon destabilized it, and Nemotron paired it with an NVFP4 schedule. The consensus is convergence speed and stability at scale, but none of the four claims the optimizer alone explains their results.
Which of the four is cheapest to train and run?
By their own accounting: Qwen3.8-Flash-Next claims roughly 1/9 the training FLOPs of its predecessor class at 6B activated parameters; Nemotron 3 Ultra claims 5.9x the decode throughput of GLM-5.1 on NVIDIA hardware; DeepSeek V4-Pro claims 27% of V3.2’s per-token FLOPs at 1M context; K3 is the largest at 104.2B activated. These are four different notions of “cheap,” and no common measurement exists.