First, GLM-5.3-Flash competed anonymously as “Ox Alpha” against other AI models; then Zhipu AI laid its cards on the table: All data traffic is said to have run on more than 100,000 Chinese AI chips. A proprietary inference engine is designed to help circumvent their memory bottlenecks.
Z.ai told ChatGPT how it would like to be portrayed.
(Image: Dall-E / AI-generated)
For a whole week, no one knew where this new AI model came from. It was simply called “Ox Alpha,” a code name. It was pretty good. On the Openrouter marketplace, it stormed the charts. After five days, Ox Alpha’s data traffic was more than double that of DeepSeek. It had surpassed 50 trillion tokens, reported the Chinese financial portal Hua’erjie Jianwen.
On August 26, 2026, the Beijing-based company Zhipu AI, known internationally as Z.ai, revealed itself as the developer of the model. This was six days after its anonymous launch on August 20. The company explained that the model is now called GLM-5.3-Flash and that its weights were available for download starting that same evening. GLM-5.3-Flash made it to 10th place on the Intelligence Index of the independent benchmarking service Artificial Analysis. This placed it ahead of DeepSeek V4 Pro Max, reported the U.S. broadcaster CNBC. It scored 57 points there—the same score as Anthropic’s Claude Opus 4.8—according to Zhipu spokespeople.
The new top-of-the-line model from China wasn't quite as good as the best U.S. models. It still lags slightly behind them when it comes to complex tasks and longer agent operations, according to Tengxun Keju, a Chinese tech portal. However, according to the results of independent tests, “an open-source model has entered a performance range that was previously dominated mainly by high-priced, proprietary flagship models,” the report also noted. So, competition for Anthropic and OpenAI—with open weights and at a fraction of the cost.
Cost-Effectiveness
A user who calls GLM-5.3-Flash via the Zhipu interface pays just $0.15 per million input tokens and $0.50 per million output tokens. That’s on par with DeepSeek and only about one-fortieth of what Anthropic’s Opus 4.8 costs. “The model’s top performance, open weights, large-scale operation on domestic chips, and extremely low API prices all come together here,” wrote the Chinese tech portal.
Another revelation from Zhipu AI made headlines worldwide. The company wrote in the model’s technical documentation that all of Ox Alpha’s data traffic ran on Chinese AI chips. There were more than 100,000 of them. “Hardware efficiency and cost per token have reached a level comparable to that of standard Nvidia GPUs,” the company itself wrote. However, it did not disclose the chip manufacturers. Analysts speculate that the chips are a mix of Huawei’s Ascend processors and components from other Chinese vendors.
The bottleneck for Chinese AI chips is typically their memory. Neither capacity nor bandwidth is sufficient when a model needs to keep track of a context of up to one million tokens. Zhipu AI has therefore developed its own inference engine. This engine compresses the caches, distributes the computational load across multiple chips on a server, and separates the processing of the request from the generation of the response. As a result, performance has tripled on the same hardware, Zhipu announced.
It's Hard to Make a Fair Comparison
This information cannot be independently verified, and an accurate price comparison remains difficult because many details are unknown. What is certain is that the chip cluster handled real-world workloads for days and performed well. U.S. export controls were actually intended to prevent this outcome. Washington prohibits the sale of Nvidia’s most powerful AI chips to China. The goal is to prevent the People’s Republic from catching up to the U.S. in the race for artificial intelligence.
But it is precisely this logic that has now been called into question. If a Beijing-based company can run a model on par with Opus 4.8 using Chinese chips—and at a fraction of the price, no less—then one might conclude that the U.S. sanctions have failed.
“All data traffic runs on in-house chips, with hardware efficiency and cost per token that are now comparable to those of Nvidia GPUs,” commented the semiconductor research firm SemiAnalysis on X.
In July of this year, Zhipu is expected to have completed a large data center in China with a capacity of one gigawatt (GW). According to the Bloomberg news agency, citing an insider, the facility is said to have already been partially commissioned. The facility is so large that, at full capacity, it consumes as much electricity as approximately 750,000 households. At least 10,000 Chinese chips have already been installed. However, to fully equip the facility, Zhipu would have to install hundreds of thousands of chips, the report stated.
Date: 08.12.2025
Naturally, we always handle your personal data responsibly. Any personal data we receive from you is processed in accordance with applicable data protection legislation. For detailed information please see our privacy policy.
Consent to the use of data for promotional purposes
I hereby consent to Vogel Communications Group GmbH & Co. KG, Max-Planck-Str. 7-9, 97082 Würzburg including any affiliated companies according to §§ 15 et seq. AktG (hereafter: Vogel Communications Group) using my e-mail address to send editorial newsletters. A list of all affiliated companies can be found here
Newsletter content may include all products and services of any companies mentioned above, including for example specialist journals and books, events and fairs as well as event-related products and services, print and digital media offers and services such as additional (editorial) newsletters, raffles, lead campaigns, market research both online and offline, specialist webportals and e-learning offers. In case my personal telephone number has also been collected, it may be used for offers of aforementioned products, for services of the companies mentioned above, and market research purposes.
Additionally, my consent also includes the processing of my email address and telephone number for data matching for marketing purposes with select advertising partners such as LinkedIn, Google, and Meta. For this, Vogel Communications Group may transmit said data in hashed form to the advertising partners who then use said data to determine whether I am also a member of the mentioned advertising partner portals. Vogel Communications Group uses this feature for the purposes of re-targeting (up-selling, cross-selling, and customer loyalty), generating so-called look-alike audiences for acquisition of new customers, and as basis for exclusion for on-going advertising campaigns. Further information can be found in section “data matching for marketing purposes”.
In case I access protected data on Internet portals of Vogel Communications Group including any affiliated companies according to §§ 15 et seq. AktG, I need to provide further data in order to register for the access to such content. In return for this free access to editorial content, my data may be used in accordance with this consent for the purposes stated here. This does not apply to data matching for marketing purposes.
Right of revocation
I understand that I can revoke my consent at will. My revocation does not change the lawfulness of data processing that was conducted based on my consent leading up to my revocation. One option to declare my revocation is to use the contact form found at https://contact.vogel.de. In case I no longer wish to receive certain newsletters, I have subscribed to, I can also click on the unsubscribe link included at the end of a newsletter. Further information regarding my right of revocation and the implementation of it as well as the consequences of my revocation can be found in the data protection declaration, section editorial newsletter.
“What GLM-5.3-Flash confirms is a pattern that comes as no surprise. Chinese labs are delivering open-source models that rival the best in the industry at a fraction of the Western price,” Bloomberg quoted Dermot McGrath, founder of the Shanghai-based consulting firm ZenGen Labs, as saying.