Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
- What it is
- Hy4-preview is a 770-billion-parameter mixture-of-experts with 49 billion active per token, 78 layers, 256 routed experts plus one shared, sparse attention and a 1-million-token window. Tencent published weights, GGUF quantizations and an API, and calls it an “early version” with headroom left in both pre- and post-training.
- What was demonstrated
- Self-reported on the model card: 65.7% on SWE-bench Pro (Kimi K3: 63.3%), 82.9% on SWE-bench Multilingual, 85.4% on Terminal-Bench 2.1, 92.3% on GPQA Diamond. In blind human evaluation against GLM-5.3: 46.8% wins, 12.8% ties, 40.4% losses.
- What was not
- All numbers were measured by Tencent itself; there is no independent replication and no technical report. Training data and hardware were not disclosed. The company admits the model “over-verifies” and reasons longer than needed. Full-precision inference needs about 1.5 TB of VRAM — “open” in practice for labs and clouds.
- Why it matters
- It is the largest Apache-2.0 open-weight release of the period, and it confirms the pattern: Chinese labs shipping frontier-adjacent agentic coders at a fraction of proprietary cost.
Same theme
Open models→GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
DeepSeek-V3: a 671B model trained in 2.79 million H800 hours, without a single loss spike
DeepSeek-R1: pure reinforcement learning, with no human reasoning examples, taught the model to verify itself
Llama 3: the 92-page report that showed how a 405B dense model is trained — and what “open” meant in 2024
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.