Understand agents in six readings
From the 2022 think-act-observe loop to the benchmark showing where 2026 agents still fail. In order, each reading prepares the next.
ReAct: interleaving reasoning and action is what made the first language agents work
If you are going to build an agent, this is the conceptual starting point: the think-act-observe loop.
Workflows are not agents: the distinction that prevents most failed agent projects
It is the vocabulary product and engineering teams need to share before discussing “putting an agent” on anything.
MCP became the tool standard for agents — and since December 2025 it no longer belongs to Anthropic
Anyone building agent integrations today builds on MCP; knowing who governs the standard is knowing who your integration depends on.
SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.
Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.
On real scientific software, the best coding agent solves fewer than half the bugs
It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.