ReAct: interleaving reasoning and action is what made the first language agents work
- What it is
- The paper by Yao and colleagues (Princeton and Google Brain) proposes that the model alternate between a natural-language reasoning step and an action step (a tool call, a search), using the action's result to update its reasoning. It is the pattern nearly every agent framework reproduced afterwards.
- What was demonstrated
- Absolute gains of +34 success points on ALFWorld and +10 on WebShop over imitation and RL baselines; on HotpotQA and FEVER, it reduced hallucination relative to pure chain-of-thought.
- What was not
- Tested in 2022 with PaLM-540B in simulated environments; absolute numbers are low by today's standards. What survived was the structure, not the results.
- Why it matters
- If you are going to build an agent, this is the conceptual starting point: the think-act-observe loop.
Appears in trails
Same theme
Agents→Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters
On real scientific software, the best coding agent solves fewer than half the bugs
Workflows are not agents: the distinction that prevents most failed agent projects
MCP became the tool standard for agents — and since December 2025 it no longer belongs to Anthropic
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.