Archive
Trails
Reading sequences built from archive entries. In order, each reading prepares the next.
Understand agents in six readings
From the 2022 think-act-observe loop to the benchmark showing where 2026 agents still fail. In order, each reading prepares the next.
- ReAct: interleaving reasoning and action is what made the first language agents work
- Workflows are not agents: the distinction that prevents most failed agent projects
- MCP became the tool standard for agents — and since December 2025 it no longer belongs to Anthropic
- SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
- Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
- On real scientific software, the best coding agent solves fewer than half the bugs
How to tell if a model is good
Benchmarks saturate, leaderboards get gamed, long contexts go unused. Six readings to read a launch announcement without being fooled.
- HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number
- A large context window does not mean the context is used: models lose what sits in the middle
- The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
- SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
- Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
- GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
AI in Brazil: what is at stake
The law that will not get voted, the authority testing before regulating, and the public money that finally became a tender. Three readings to understand the board.
- PL 2338: approved by the Senate in December 2024, Brazil's AI bill has been stalled in the Chamber for sixteen months
- ANPD's AI sandbox has three companies testing explainability under the LGPD — and a second cycle announced for late 2026
- The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October