Zhipu Open-Sources GLM-5.3-Flash: Ox Alpha Mystery Solved, Entirely Running on Domestic Chips at 1/40 the Price of Opus 4.8

8.27 Over the past week, an anonymous model named Ox Alpha on OpenRouter and OpenCode sparked intense global developer interest, setting new traffic records on both platforms. On August 26, the mystery was solved—Zhipu officially launched and open-sourced GLM-5.3-Flash (320B-A18B) , the first native multimodal model in the GLM-5 series. Even more striking, all traffic during the anonymous testing period was served entirely on domestic AI chips, with peak daily capacity reaching 100 trillion toke

Over the past week, an anonymous model named Ox Alpha on OpenRouter and OpenCode sparked intense global developer interest, setting new traffic records on both platforms. On August 26, the mystery was solved—Zhipu officially launched and open-sourced GLM-5.3-Flash (320B-A18B) , the first native multimodal model in the GLM-5 series. Even more striking, all traffic during the anonymous testing period was served entirely on domestic AI chips, with peak daily capacity reaching 100 trillion tokens. The deployment involved a cluster of over 100,000 domestic chips.

2026-08-27_095457_157

I. Performance Comparable to Claude Opus 4.8: Intelligence Index Score of 57

GLM-5.3-Flash has 320B total parameters with only 18B activated. It scored 57 on the Artificial Analysis Intelligence Index, placing it within the frontier model capability range—tying Anthropic's Claude Opus 4.8. On Zhipu's proprietary Z.ai Code Bench, its coding performance also matched Opus 4.8.

Compared to the previous flagship GLM-5.2, GLM-5.3-Flash delivers superior performance across all benchmarks at just 1/10 the price.

2026-08-27_095547_652

II. Architectural Breakthrough: First Open-Source Frontier Model with Hybrid Sparse+Linear Attention

This breakthrough stems from fundamental architectural innovation. GLM-5.3-Flash is the first open-source frontier model to adopt a hybrid sparse attention and linear attention architecture.

Compared to GLM-4.5, total parameters are similar (355B vs 320B), but activated parameters dropped from 32B to 18B, with layers reduced from 92 to 45. Relative to GLM-5.3, attention computation decreased by 3.01x and KV cache size by 4.44x—significantly lowering long-context serving costs while maintaining precision.

The model also features mHC for improved scaling and is trained on 30T tokens of multimodal pretraining data, achieving stronger performance with less compute. It natively supports up to 1M context windows, image and video input, and integrates visual capabilities into the coding loop.

III. Domestic Chips Under Load: 100,000-Card Cluster Serves Global "Stress Test"

The compute mystery behind Ox Alpha's anonymous testing has also been solved. All traffic was served on domestic AI chips.

Zhipu deployed over 100,000 domestic chips connected via high-bandwidth proprietary interconnects. To overcome memory and bandwidth limitations, Zhipu built a custom inference engine on SGLang, employing W8A8 quantization, mixed cache quantization, tensor parallelism, and EPD separation architecture.

Zhipu stated that end-to-end service performance improved 3x compared to baseline, with hardware efficiency and per-token costs now comparable to mainstream NVIDIA GPUs.

According to sources cited by LatePost, the chip suppliers are likely Huawei, Moore Threads, and Hygon. SemiAnalysis commented on X: "Following Jalapeño, the CUDA moat is being tested once again." 

IV. Pricing: 1/40 of Claude Opus 4.8

GLM-5.3-Flash is priced at 1/10 of GLM-5.3, discounted to 1/20, and 1/40 of Claude Opus 4.8. API pricing is RMB 0.8 per million input tokens, RMB 2.8 per million output tokens, and RMB 0.23 for cache hits.

Model weights are open-sourced on Hugging Face, freely available for download, modification, and commercial use. The model is now available on platforms including ZCode, with API access open through GLM Coding Plan.

GLM-5.3-Flash sends at least three signals: First, domestic chips are now capable of serving frontier AI models at global scale—100 trillion tokens/day of real traffic on a 100,000-card cluster proves domestic compute has moved from "usable" to "production-ready"; Second, the combination of architectural innovation and domestic compute optimization is redefining the cost curve—matching Claude Opus 4.8's intelligence at 1/40 the cost demonstrates engineering capability, not just a pricing strategy; Third, open source remains central to Zhipu's strategy—the combination of anonymous beta testing, open weights, and integration with coding platforms is building a developer ecosystem moat.

Related links:

BigModel:https://docs.bigmodel.cn/cn/guide/models/vlm/glm-5.3-flash

Z.ai:https://docs.z.ai/guides/vlm/glm-5.3-flash

GLM Coding Plan:https://bigmodel.cn/glm-coding

HuggingFace:https://huggingface.co/zai-org/GLM-5.3-Flash