Open models just got too cheap to ignore
This is the first edition of Cinco. The format is what the name says: five things a week, chosen from a couple of hundred I read, each with what it is, what was demonstrated, what was not, and why it matters to people who build. The original source is always one click away.
The week had a theme without anyone agreeing on it: cost. Two Chinese labs released open-weight models under truly permissive licenses — one of them served, the company says, only on domestic chips, at fifteen cents per million tokens. In the same window, OpenAI launched GPT-6 Astra without touching the price, and admitted in its own report that reading the model's reasoning has become less reliable. A Fudan benchmark reminded us that, on real scientific software, the best agent solves fewer than half the bugs. And Brazil published, with a number and a deadline, the tender for the RN supercomputer.
If you only have time for one: the second.
Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
- What it is
- Hy4-preview is a 770-billion-parameter mixture-of-experts with 49 billion active per token, 78 layers, 256 routed experts plus one shared, sparse attention and a 1-million-token window. Tencent published weights, GGUF quantizations and an API, and calls it an “early version” with headroom left in both pre- and post-training.
- What was demonstrated
- Self-reported on the model card: 65.7% on SWE-bench Pro (Kimi K3: 63.3%), 82.9% on SWE-bench Multilingual, 85.4% on Terminal-Bench 2.1, 92.3% on GPQA Diamond. In blind human evaluation against GLM-5.3: 46.8% wins, 12.8% ties, 40.4% losses.
- What was not
- All numbers were measured by Tencent itself; there is no independent replication and no technical report. Training data and hardware were not disclosed. The company admits the model “over-verifies” and reasons longer than needed. Full-precision inference needs about 1.5 TB of VRAM — “open” in practice for labs and clouds.
- Why it matters
- It is the largest Apache-2.0 open-weight release of the period, and it confirms the pattern: Chinese labs shipping frontier-adjacent agentic coders at a fraction of proprietary cost.
GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
- What it is
- Z.ai revealed that OpenRouter's anonymous “Ox Alpha” was GLM-5.3-Flash: a 320B MoE with 18B active, natively multimodal, trained on a 30-trillion-token corpus with hybrid sparse and linear attention. Weights are MIT-licensed. The company says it serves the model on domestic accelerators, at about 100 trillion tokens per day.
- What was demonstrated
- API price: US$ 0.15 per million input tokens and US$ 0.50 output. Self-reported: 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE v1.1 (GLM-5.2: 46.2). Artificial Analysis intelligence index (third party): 57. Z.ai says the new attention cuts compute ~3× and KV cache ~4.4× versus GLM-5.3.
- What was not
- The “Chinese chips only” claim is Z.ai's own; the accelerator and count are undocumented. Benchmark comparisons on the model card are an image, not a table, and evals cap context at 300k tokens despite the 1M window. The linked paper is February's GLM-5 report, not this model's.
- Why it matters
- It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.
- What it is
- SWE-bench Science, from Fudan University's OpenMOSS group, collects 119 real issues from 98 repositories across 20 scientific domains — imaging, spectroscopy, simulation. Failures involve units, coordinate systems, numerical invariants and physical assumptions, not plain logic. The authors propose a four-mechanism failure taxonomy and test whether injecting scientific guidance helps.
- What was demonstrated
- Leaderboard on 1 September: Claude Opus 5 (max effort) 47.9%; DeepSeek-V4-Pro 42.0%; GPT-5.6 Sol 40.3%; Kimi K3 35.3%; GLM-5.2 31.9%. The same models score above 95% on SWE-bench Verified. Well-grounded guidance helps; poorly calibrated guidance causes anchoring and hurts.
- What was not
- Small sample (119 tasks): differences of a few points between models are within noise. Academic group with no vendor ties, but run cost limits replication. Preprint v2, not yet peer-reviewed. Not to be confused with “Terminal-Bench-Science 0.1”, which appears in this week's launch decks.
- Why it matters
- It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.
GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
- What it is
- OpenAI launched GPT-6 Astra on 3 September, for ChatGPT, API, Azure and Bedrock, keeping US$ 10 per million input tokens and US$ 50 output. It is the first model the company rates “Critical” for cybersecurity and “High” for bio/chem under its Preparedness Framework. The system card records that chain-of-thought monitorability decreased versus GPT-5.6 Sol.
- What was demonstrated
- Measured by OpenAI: 99.9% on ARC-AGI-3 (Sol: 7.8%); 97.6% on FrontierMath Tier 4; 57.9% on Terminal-Bench 4.0; 72.6% on OSWorld 2.0. Indirect prompt-injection success fell from 27.0% to 8.5%. About half as many severe-misalignment flags across 54k Codex tasks.
- What was not
- All numbers are OpenAI's own; ARC-AGI-3 ran through an “adapter harness”, and sources disagree on context length (256k to ~1M). The system card itself says the model produces shorter chains, “can evade monitors when strategically underperforming”, and that the company will have “significantly reduced confidence” in detecting some misaligned behaviors. Cyber refusals rose to ~94%, which may block legitimate defensive work.
- Why it matters
- For builders, it resets the frontier price/performance point; for governance, it is the first flagship whose own report admits reading the reasoning is becoming less reliable.
The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October
- What it is
- On 20 August the federal government announced about R$ 2.5 billion in PBIA-linked AI infrastructure, and LNCC, through FACC, published the tender for the Macaíba (RN) supercomputer: 7,200 PFLOPS FP16, over 50k cores, ~4 MW, operations expected by end-2027, with local-content and technology-transfer clauses. The package includes the RNP–Huawei partnership in Rio for Portuguese-language models (R$ 1.276 bn over five years), R$ 250 million for open-architecture chips in Campinas and R$ 50 million for an algorithmic-transparency center at UFMG.
- What was demonstrated
- Tender value: R$ 959,040,959.04. Deadline: 8 October 2026, 10:00. Commitment to train 5,000 professionals over five years. Minister Luciana Santos said 60% of Brazilian AI processing currently happens abroad.
- What was not
- It is a procurement announcement, not a delivered machine: the tender had been promised for 2025 and then early 2026. Supplier, GPU type and FLOPS methodology (dense or sparse) are unspecified; operating costs sit outside the R$ 1 billion. The announcement coincided with the president's first Northeast trip after launching his candidacy, which local press noted. Totals vary between R$ 2.3 bn (LNCC) and R$ 2.5 bn (press).
- Why it matters
- It is the first concrete, dated and funded step toward sovereign AI compute in Brazil — and the tender terms will shape which stacks Brazilian engineers get to use.
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.