Skip to content

Agents and Tool Use

This is the research on AI that does more than answer — it takes action. Papers here cover agents that plan multi-step work, call tools and APIs, browse, write and run code, and recover when a step fails.

Flashy agent demos are common; agents that hold up in a real workflow are rarer. The briefs in this theme focus on what held up under testing: where agents got real work done, and where they broke down mid-task. If you're deciding whether to put an agent into a real workflow, start here.

In this section

Page Last updated
There's No Silver Bullet for Coding-Agent Rewards
Why a single reward can't keep a coding agent honest, and what a layered verification stack does instead.
Updated 2026-06-29