Fields / 03

Agents

Agents are reliable inside well-instrumented environments such as codebases and browsers with clear goals. Open-ended autonomy over long horizons is still fragile.

EmergingAcceleratingRevised 5 Oct 2026Draft — awaiting editor review
What to watch
  1. Task length an agent can complete unattended
  2. Standard interfaces for tools and agent-to-agent communication
  3. Security, especially prompt injection through tool results

The state of play

An agent is a model in a loop: observe, reason, call a tool, observe the result, repeat. ReAct gave the loop its canonical shape in 2022. What changed since is that models got good enough at tool use and planning for the loop to finish real tasks. Computer-use agents in 2024 and terminal coding agents in 2025 brought this to everyday software.

Infrastructure is converging too. The Model Context Protocol turned tool integration into something you write once and reuse across clients.

What works today

  • Software engineering tasks with tests to check against.
  • Research and browsing tasks that end in a document a human reviews.
  • Narrow operational workflows with a small, well-defined toolset.

What doesn’t yet

  • Long-horizon tasks without checkpoints. Errors compound.
  • Acting safely on untrusted input. Any tool result can carry instructions.
  • Cost predictability for open-ended runs.

Key ideas

Idea In one line
Tool design Fewer, sharper tools beat many overlapping ones.
Environment feedback Agents improve dramatically when they can check their own work.
Human checkpoints Approval at irreversible steps is still the main safety mechanism.

Open questions

What is the right unit of delegation: a task, a role, or a whole project?

Evidence

Lab reports & log entries