On September 18, Zhipu officially launched GLM-5.3-FlashX, delivering inference speeds of up to 200 tokens/s for a faster, smoother experience for enterprises and developers. The API is now live with the model key GLM-5.3-FlashX.

From Ox Alpha to FlashX: From Anonymous Testing to Speed Upgrade
GLM-5.3-Flash previously debuted to global developers under the anonymous model name Ox Alpha, earning widespread recognition and steadily climbing usage. Facing rapidly growing demand, Zhipu further increased investment in infrastructure and inference optimization on top of the inference compute provided by 100,000 domestic chips.
Zhipu states that GLM-5.3-Flash has long maintained the strongest intelligence at its size, and this speed boost further establishes comprehensive competitiveness across intelligence, price, and speed.
Performance: Intelligence Index Score of 57, Matching Claude Opus 4.8
GLM-5.3-Flash has 320B total parameters with only 18B activated, making it the world's first open-source frontier model to adopt a hybrid sparse attention and linear attention architecture. It scored 57 on the Artificial Analysis Intelligence Index, placing it within the global frontier model range and matching Anthropic's Claude Opus 4.8.
Its coding and agent capabilities are also comparable to Claude Opus 4.8, with native support for up to 1M context, images, and video input.
Pricing Strategy: FlashX Positions as a "Speed Premium"
GLM-5.3-FlashX is priced higher than GLM-5.3-Flash, forming a clear product tier:
| Model | Context | Input (¥/M tokens) | Output (¥/M tokens) | Cache Hit (¥/M tokens) | Input Modalities |
|---|---|---|---|---|---|
| GLM-5.3 | 1M | 8 | 28 | 2 | Text |
| GLM-5.3-Flash | 1M | 0.8 | 2.8 | 0.23 | Image, Video, File, Text |
| GLM-5.3-FlashX | 1M | 2 | 7 | 0.57 | Image, Video, File, Text |
FlashX is priced at roughly 2.5x that of Flash in exchange for faster inference. For high-frequency, low-latency scenarios—such as real-time agents and interactive coding—this "speed premium" offers clear value.
Domestic Chips Under the Hood: 100,000-Card Cluster Powers Global Traffic
Behind FlashX's speed is Zhipu's sustained investment in a cluster of over 100,000 domestic chips. During the earlier anonymous testing period, all request traffic was served on domestic chips, peaking at 100 trillion tokens per day. The cluster is connected via proprietary high-bandwidth interconnects, with a custom inference engine built on SGLang, achieving 3x end-to-end service performance over baseline.
This demonstrates that domestic compute can now serve frontier model inference at global scale, with cost and efficiency gradually approaching mainstream international solutions.
GLM-5.3-FlashX marks Zhipu's evolution from "extreme cost-performance" toward "tiered productization." While Flash set a value benchmark at 1/20 the price of GLM-5.3, FlashX now addresses differentiated developer needs through a "speed premium"—those pursuing ultimate cost can stick with Flash, while those needing maximum responsiveness can choose FlashX. This clear product matrix is a sign of a maturing foundation model vendor. And beneath it all is the stable support of 100,000 domestic chips.