Institution
Moonshot AI
Beijing AI lab behind the Kimi models, Kimi Delta Attention and the scaled-up Muon optimizer.
Language Models · Moonshot AI
Kimi K3 is a 2.8T-parameter MoE with 104B active, native vision and 1M context, built on 3:1 Kimi Delta Attention and gated MLA, attention residuals and a 896-expert LatentMoE, at 2.5x K2's scaling efficiency.
Language Models · Moonshot AI
Moonshot AI shows Muon, a matrix-orthogonalizing optimizer, reaches AdamW's loss with about half the compute once you add weight decay and rescale its updates; Moonlight, a 16B MoE trained on 5.7T tokens, is the proof.