Kimi K3: an open 2.8T-parameter MoE model with a 1M-token context and native vision
Kimi K3 is an open Mixture-of-Experts model with 2.8T parameters (104B activated per token), native vision input and a 1-million-token context window, built on Kimi Delta Attention and post-trained with reinforcement learning across general, agentic and coding tasks.
- 68 stars/7d
- 16 citations
- 519 upvotes
- #1 mathvision