Kimi just dropped the model weights and technical report for K3, a 2.8 trillion parameter open-weight model with native vision and a 1M token context window.
2.5x the intelligence per unit of compute. Not just bigger. More efficient.
They also open-sourced the infrastructure stack behind it: attention kernels, MoE communication library, and agent environment tooling.

2.5x the intelligence per unit of compute. Not just bigger. More efficient.
They also open-sourced the infrastructure stack behind it: attention kernels, MoE communication library, and agent environment tooling.

42❤️3❤️1