Skip to content
GenAI BR
Infra and costFeatured in Cinco no. 1 · In the archive since 08 Sept 2026

GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026
Download card
What it is
Z.ai revealed that OpenRouter's anonymous “Ox Alpha” was GLM-5.3-Flash: a 320B MoE with 18B active, natively multimodal, trained on a 30-trillion-token corpus with hybrid sparse and linear attention. Weights are MIT-licensed. The company says it serves the model on domestic accelerators, at about 100 trillion tokens per day.
What was demonstrated
API price: US$ 0.15 per million input tokens and US$ 0.50 output. Self-reported: 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1 (GLM-5.2: 46.2). Artificial Analysis intelligence index (third party): 57. Z.ai says the new attention cuts compute ~3× and KV cache ~4.4× versus GLM-5.3.
What was not
The “Chinese chips only” claim is Z.ai's own; the accelerator and count are undocumented. Benchmark comparisons on the model card are an image, not a table, and evals cap context at 300k tokens despite the 1M window. The linked paper is February's GLM-5 report, not this model's.
Why it matters
It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.
  1. The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October

    Regulation and Brazilintroductorytechnical, decision-maker, beginnerannouncement · Governo Federal · 21 Aug 2026
  2. A large context window does not mean the context is used: models lose what sits in the middle

    Evaluationintroductorytechnical, beginnerpaper · Stanford / UC Berkeley / Samaya · TACL 2023 · 06 Jul 2023
  3. DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike

    Infra and costtechnicaltechnical, decision-makerpaper · DeepSeek · 27 Dec 2024
  4. PagedAttention: treating the KV cache like virtual memory doubled inference throughput — and gave birth to vLLM

    Infra and costtechnicaltechnicalpaper · UC Berkeley · SOSP 2023 · 12 Sept 2023

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.