Agents and Tool Use¶
This is the research on AI that does more than answer — it takes action. Papers here cover agents that plan multi-step work, call tools and APIs, browse, write and run code, and recover when a step fails.
Flashy agent demos are common; agents that hold up in a real workflow are rarer. The briefs in this theme focus on what held up under testing: where agents got real work done, and where they broke down mid-task. If you're deciding whether to put an agent into a real workflow, start here.
In this section¶
| Page | Last updated |
|---|---|
| There's No Silver Bullet for Coding-Agent Rewards Why a single reward can't keep a coding agent honest, and what a layered verification stack does instead. |
Updated 2026-06-29 |