Skip to content

Reasoning and Capabilities

How well do models actually reason? Papers here look at how a model arrives at a good answer — step-by-step reasoning, planning, math and logic, and the failure modes that show up when a problem gets a little harder than the examples a model saw.

These papers test whether a capability holds up beyond the demo. The briefs in this theme are careful to separate what a model can do from what it looks like it can do. When a paper claims a new reasoning skill, we ask whether it survives a harder version of the test.

In this section

Page Last updated
Why AI Drops Its Caution the Moment You Ask for Advice
AI that hedges a correlation in analysis asserts it as cause-and-effect once you ask for advice — and a re-prompt brings the caution back.
Updated 2026-06-29