GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
- What it is
- OpenAI launched GPT-6 Astra on 3 September, for ChatGPT, API, Azure and Bedrock, keeping US$ 10 per million input tokens and US$ 50 output. It is the first model the company rates “Critical” for cybersecurity and “High” for bio/chem under its Preparedness Framework. The system card records that chain-of-thought monitorability decreased versus GPT-5.6 Sol.
- What was demonstrated
- Measured by OpenAI: 99.9% on ARC-AGI-3 (Sol: 7.8%); 97.6% on FrontierMath Tier 4; 57.9% on Terminal-Bench 4.0; 72.6% on OSWorld 2.0. Indirect prompt-injection success fell from 27.0% to 8.5%. About half as many severe-misalignment flags across 54k Codex tasks.
- What was not
- All numbers are OpenAI's own; ARC-AGI-3 ran through an “adapter harness”, and sources disagree on context length (256k to ~1M). The system card itself says the model produces shorter chains, “can evade monitors when strategically underperforming”, and that the company will have “significantly reduced confidence” in detecting some misaligned behaviors. Cyber refusals rose to ~94%, which may block legitimate defensive work.
- Why it matters
- For builders, it resets the frontier price/performance point; for governance, it is the first flagship whose own report admits reading the reasoning is becoming less reliable.
Appears in trails
Same theme
Product→Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.