Once again, a large-scale AI model from China is drawing global attention. This time, it comes from the Beijing-based start-up Moonshot AI and is called Kimi K3.
Kimi K3 can fundamentally be self-hosted. However, due to its size, it is not a model for just any company: even in a compressed form, extensive storage and computing resources are required.
(Image: Moonshot AI)
In the early morning hours of July 17, Moonshot released the new model Kimi K3, reports the Chinese financial portal Damo Caijing. It features 2.8 trillion parameters, processes a context window of one million tokens, and is primarily designed for long programming tasks, knowledge work, and complex reasoning. The often intense reactions worldwide are reminiscent of January 2025, when the Chinese startup Deepseek unveiled its open model R1, marking the so-called "Deepseek Moment."
At that time, chipmaker Nvidia lost almost 600 billion US dollars in market value in a single trading day. Investors feared that artificial intelligence would require much less computing power than previously thought, Bloomberg reported. This time as well, the stock prices of chip manufacturers around the globe have dropped again. Just like during the Deepseek moment, it appears that China's AI developers have come very close to their American competitors. There is talk of a reduction in the gap to just a few weeks instead of several months as before.
In the Intelligence Index of the independent benchmarking service Artificial Analysis, Kimi K3 currently scores 57 points. This places the model behind Claude Fable 5 and GPT-5.6 Sol, but ahead of Claude Opus 4.8 and GPT-5.5. However, these rankings are a snapshot and depend on the specific benchmark, the chosen reasoning task, and the agent environment used. In the Frontend Code Arena, Kimi K3 ranks first. This demonstrates the model's exceptional performance, particularly in programming tasks, but cannot be easily extrapolated to its overall performance.
"For open weights, this is a milestone that changes the game," wrote Cline, a platform for open-source programming. Open weights are trained model parameters that are freely published. In other words, they are the numerical values that a AI model learns during training and where its entire capability resides. Anyone who can download and thus possess them can run and modify the model on their own hardware.
In contrast to open-source software, training code and training data remain secret. However, open weights are significantly more transparent than, for example, the closed models of Anthropic and OpenAI, which are only accessible via the providers' cloud. Their weights remain a trade secret. As a result, users must also entrust their company data or private information to the cloud.
A standardized reasoning task with Kimi K3 costs an average of $0.95. Fable 5 charges $2.75 for the same, and GPT-5.6 Sol costs $1.06, according to Artificial Analysis calculations. This makes Kimi K3 significantly cheaper than its American competitors but more expensive than leading Chinese models like GLM-5.2. Kimi K3 is thus not a particularly inexpensive model. Artificial Analysis even classifies it as relatively expensive compared to other open-weight models. Its distinctiveness lies rather in the fact that a model of this performance class is available with open weights.
Programming is the specialty of the new model. "With minimal human supervision, it can endure long engineering sessions, navigate massive repositories, and orchestrate terminal tools," says Moonshot's press release. The model can autonomously work through the branched source code archives of large software projects, thereby simplifying complex tasks for programmers.
Technically speaking, Kimi K3 shifts the bottleneck of AI infrastructure from computing power to storage. The model pushes the so-called "sparsity ratio" to a record high, Bloomberg reports. The higher this ratio, the smaller the fraction of parameters the model activates for a single task. For each query, Kimi K3 essentially only "wakes up" the specialists that are needed at the moment, while the rest of the network remains dormant. This makes the inference, i.e., the model's operational runtime, significantly more efficient.
These dormant specialists still need to be stored. All 2.8 trillion parameters must remain continuously in the memory of the compute accelerators, even if only a fraction of them are used for each task. Even after compression with reduced numerical precision, Kimi K3 takes up about 1.4 terabytes. To fully utilize the model, clusters of memory-rich AI processors, such as Nvidia's GB300 systems, are required. According to Damo Caijing, Moonshot itself recommends a network of at least 64 accelerator cards.
Headwind from the USA
Unlike after the Deepseek shock, the demand for high-bandwidth memory from manufacturers like SK Hynix and for state-of-the-art chip production from TSMC is expected to remain high this time, Bloomberg analyzes. In Washington, the response is hostile. "We have information that Moonshot AI distilled Anthropics Fable to develop its K3 model," wrote Michael Kratsios, the Director of the White House Office of Science and Technology Policy, on X.
Date: 08.12.2025
Naturally, we always handle your personal data responsibly. Any personal data we receive from you is processed in accordance with applicable data protection legislation. For detailed information please see our privacy policy.
Consent to the use of data for promotional purposes
I hereby consent to Vogel Communications Group GmbH & Co. KG, Max-Planck-Str. 7-9, 97082 Würzburg including any affiliated companies according to §§ 15 et seq. AktG (hereafter: Vogel Communications Group) using my e-mail address to send editorial newsletters. A list of all affiliated companies can be found here
Newsletter content may include all products and services of any companies mentioned above, including for example specialist journals and books, events and fairs as well as event-related products and services, print and digital media offers and services such as additional (editorial) newsletters, raffles, lead campaigns, market research both online and offline, specialist webportals and e-learning offers. In case my personal telephone number has also been collected, it may be used for offers of aforementioned products, for services of the companies mentioned above, and market research purposes.
Additionally, my consent also includes the processing of my email address and telephone number for data matching for marketing purposes with select advertising partners such as LinkedIn, Google, and Meta. For this, Vogel Communications Group may transmit said data in hashed form to the advertising partners who then use said data to determine whether I am also a member of the mentioned advertising partner portals. Vogel Communications Group uses this feature for the purposes of re-targeting (up-selling, cross-selling, and customer loyalty), generating so-called look-alike audiences for acquisition of new customers, and as basis for exclusion for on-going advertising campaigns. Further information can be found in section “data matching for marketing purposes”.
In case I access protected data on Internet portals of Vogel Communications Group including any affiliated companies according to §§ 15 et seq. AktG, I need to provide further data in order to register for the access to such content. In return for this free access to editorial content, my data may be used in accordance with this consent for the purposes stated here. This does not apply to data matching for marketing purposes.
Right of revocation
I understand that I can revoke my consent at will. My revocation does not change the lawfulness of data processing that was conducted based on my consent leading up to my revocation. One option to declare my revocation is to use the contact form found at https://contact.vogel.de. In case I no longer wish to receive certain newsletters, I have subscribed to, I can also click on the unsubscribe link included at the end of a newsletter. Further information regarding my right of revocation and the implementation of it as well as the consequences of my revocation can be found in the data protection declaration, section editorial newsletter.
In distillation, a smaller model is trained with the answers of a larger or less specialized one. Anthropic had already accused Moonshot and other Chinese providers of large-scale distillation campaigns back in February. According to Reuters, the company cited more than 3.4 million interactions and hundreds of allegedly fake accounts.
Kratsios also accused Moonshot of accessing servers in Thailand with Nvidia GB300 chips, which are subject to American export controls. Moonshot denies the allegations. U.S. Treasury Secretary Scott Bessent stated that financial sanctions or the inclusion of Chinese companies on the Entity List could be considered in the event of confirmed violations.
The model outperforms exactly those American models from which it theoretically could have learned, including Opus 4.8 and GPT-5.5. Kimi K3 is "a very good model," whose performance cannot be explained by distillation alone, even admits Dean Ball, the futurist at OpenAI.
Particularly threatening to the business model of the American market leaders is the openness of the model, continuing the Chinese open-weight strategy. The announced release of the Kimi K3 weights has now taken place. Moonshot published the complete model around July 27, 2026, on Hugging Face. However, it is not under a completely free standard license but rather under the proprietary Kimi K3 license. Additional conditions apply for certain commercial model-as-a-service offerings.
Kimi K3 can fundamentally be self-hosted. However, due to its size, it is not a model for just any company: even in compressed form, extensive storage and computing resources are required. While an in-house installation allows requests to remain within the company’s own infrastructure, the operation incurs significant costs.
Programming services are ironically among the most lucrative offerings with which Anthropic and OpenAI are preparing their stock market plans, writes Bloomberg. A downloadable top-tier model puts this business model under pressure—not because its use would be free, but because companies gain more control over the model, data, and operation.