GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
- What it is
- Z.ai revealed that OpenRouter's anonymous “Ox Alpha” was GLM-5.3-Flash: a 320B MoE with 18B active, natively multimodal, trained on a 30-trillion-token corpus with hybrid sparse and linear attention. Weights are MIT-licensed. The company says it serves the model on domestic accelerators, at about 100 trillion tokens per day.
- What was demonstrated
- API price: US$ 0.15 per million input tokens and US$ 0.50 output. Self-reported: 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1 (GLM-5.2: 46.2). Artificial Analysis intelligence index (third party): 57. Z.ai says the new attention cuts compute ~3× and KV cache ~4.4× versus GLM-5.3.
- What was not
- The “Chinese chips only” claim is Z.ai's own; the accelerator and count are undocumented. Benchmark comparisons on the model card are an image, not a table, and evals cap context at 300k tokens despite the 1M window. The linked paper is February's GLM-5 report, not this model's.
- Why it matters
- It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.
Same theme
Infra and cost→The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October
A large context window does not mean the context is used: models lose what sits in the middle
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.