Skip to content
GenAI BR
Archive · Theme7 entries in this theme

Agents

Systems that plan, use tools and act — and how to tell when they work.

  1. Tencent open-sourced a 770B MoE under Apache 2.0 — and it edges Kimi K3 on code with a quarter of the parameters

    It is the largest Apache-2.0 open-weight release of the period, and it confirms the pattern: Chinese labs shipping frontier-adjacent agentic coders at a fraction of proprietary cost.

    Open modelstechnicaltechnical, decision-makerannouncement · Tencent · 28 Aug 2026
  2. On real scientific software, the best coding agent solves fewer than half the bugs

    It quantifies the gap between saturated leaderboards and real domain-heavy engineering — exactly where engineers will be asked to deploy agents next.

    Evaluationtechnicaltechnicalpaper · Universidade Fudan · 01 Sept 2026
  3. ReAct: interleaving reasoning and action is what made the first language agents work

    If you are going to build an agent, this is the conceptual starting point: the think-act-observe loop.

    Agentsintroductorytechnical, beginnerpaper · Princeton / Google Brain · ICLR 2023 · 06 Oct 2022
  4. Workflows are not agents: the distinction that prevents most failed agent projects

    It is the vocabulary product and engineering teams need to share before discussing “putting an agent” on anything.

    Agentsintroductorytechnical, decision-makerpost · Anthropic · 19 Dec 2024
  5. MCP became the tool standard for agents — and since December 2025 it no longer belongs to Anthropic

    Anyone building agent integrations today builds on MCP; knowing who governs the standard is knowing who your integration depends on.

    Agentstechnicaltechnicalcode · Agentic AI Foundation · 28 Jul 2026
  6. SWE-bench: the coding benchmark that went from 2% to 97% in three years — and what that says about benchmarks

    It is the number every launch cites; knowing how it is measured is the difference between reading an announcement and understanding it.

    Evaluationtechnicaltechnicalpaper · Princeton / UChicago · ICLR 2024 · 10 Oct 2023
  7. Terminal-Bench: 89 hand-verified terminal tasks, and the benchmark that replaced SWE-bench in launch decks

    It is the benchmark that currently discriminates between frontier models on agentic work — and what will appear in the next headlines.

    Evaluationtechnicaltechnicalpaper · Laude Institute / Stanford · 17 Jan 2026