self-repair · active

Fixing its own mistakes

Whether and how a model can diagnose and correct a wrong answer once it's already made.

Tags: agentic, coding-agent, math

This topic covers self-correction and self-critique behavior: does reviewing or revising an answer make it better or worse, and what turns out to matter is what's driving the revision — the model's own unaided judgment, or a concrete external signal like a failing test. The two are easy to conflate under one banner ("self-correction") and behave very differently, which is exactly the kind of thing a bare "capability score" would have hidden.

Claims

Techniques

Related: Using the tools it is given, Tracking state through a long task, Generating and editing working code
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.