Moonshot AI’s Kimi K3 open-weight model has been read almost entirely through its parameter count since it launchedon July 16. At 2.8 trillion parameters, it is the largest open-weight model released to date. Model sizes are usually grouped into rough brackets, and 2.8 trillion rounds into what the industry calls the 3T class. A tier no openly available model had entered before.

The natural conclusion is that Moonshot has engineered its way around US compute restrictions. The company’s own technical blog suggests something more specific: K3 does not avoid the constraint so much as relocate it, trading compute for memory at almost every layer of the design.
That trade is worth understanding, because compute and memory are not interchangeable constraints, and they are not equally available to a Chinese lab.
Why the Kimi K3 open-weight model is a memory problem
Two different things determine what it costs to run a large model. One is how much calculation the machine does to produce each word. The other is how much of the model has to be held ready and instantly reachable the entire time it is working. The first is compute. The second is memory. Chip export controls have squeezed China hard at first, and Moonshot’s design reads as a sustained attempt to spend less of it.
The main move is a technique called mixture-of-experts. Rather than run the whole model for every word, K3 splits itself into 896 specialised sections and calls on just 16 of them at a time, about 1.8% of the total. The calculation per word drops sharply. The memory bill does not move at all, because all 2.8 trillion parameters still have to sit loaded and ready in case they are the ones called next.
So Moonshot went after that bill directly. It trained K3 to work at four bits of precision per parameter instead of the usual sixteen, a method known as quantisation-aware training, which the company applied from the fine-tuning stage onward and says it chose “for broad hardware compatibility”, a phrase worth pausing on, since it reads as a hedge against running on silicon that is not Nvidia’s. The savings are substantial. Independent analysis of the release puts the model at roughly 1.4TB in that format, against the 5.6TB it would need at full precision.
The second change, Kimi Delta Attention, targets a different memory cost. As a model works through a very long document,…
Source link
Disclaimer
We strive to uphold the highest ethical standards in all of our reporting and coverage. We blogs.grocliq.com want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.
Website Upgradation is going on for any glitch kindly connect at [email protected]