Skip to content
GenAI BR
Archive3 trails

Trails

Reading sequences built from archive entries. In order, each reading prepares the next.

  1. Understand agents in six readings

    From the 2022 think-act-observe loop to the benchmark showing where 2026 agents still fail. In order, each reading prepares the next.

    1. ReAct: interleaving reasoning and action is what made the first language agents work
    2. Workflows are not agents: the distinction that prevents most failed agent projects
    3. MCP became the tool standard for agents — and since December 2025 it no longer belongs to Anthropic
    4. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
    5. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
    6. On real scientific software, the best coding agent solves fewer than half the bugs
  2. How to tell if a model is good

    Benchmarks saturate, leaderboards get gamed, long contexts go unused. Six readings to read a launch announcement without being fooled.

    1. HELM: the proposal to evaluate models across many scenarios and metrics at once, not on a single number
    2. A large context window does not mean the context is used: models lose what sits in the middle
    3. The Leaderboard Illusion: how the Arena favors labs that test private variants and get more data
    4. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks
    5. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks
    6. GPT-6 Astra: same price, new benchmark ceiling — and a system card that says reasoning is getting harder to monitor
  3. AI in Brazil: what is at stake

    The law that will not get voted, the authority testing before regulating, and the public money that finally became a tender. Three readings to understand the board.

    1. PL 2338: approved by the Senate in December 2024, Brazil's AI bill has been stalled in the Chamber for sixteen months
    2. ANPD's AI sandbox has three companies testing explainability under the LGPD — and a second cycle announced for late 2026
    3. The government published the R$ 959 million tender for the Rio Grande do Norte AI supercomputer — bids due 8 October