repository-context-files-do-not-raise-task-success-on-benchmark-coding
observationsingle papercontestedpending review

Giving a coding agent a repository context file — AGENTS.md, CLAUDE.md — does not raise its success rate on benchmark coding tasks and costs about 20% more inference: across 4 agents, 2 benchmarks and 3 conditions, LLM-generated files hurt slightly in 5 of 8 settings while developer-written ones gained 2.4% (p=0.21), and agents obey the files — which is why they spend more — so the files do not carry success-relevant information rather than being ignored.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Following an unfamiliar procedure · Coding agent, Context and memory

Observed on

2026, Claude Code / Codex / Qwen Code on SWE-bench Lite and CTXbench.

Sources

Suspected axis of disagreement(a guess, not verified)

Probably the outcome variable, and possibly the file's origin. The study measures success on tasks the agent has never seen, in repositories where the file was written generically or generated by the agent; the practitioner claim is about a file that accumulates the exact failures one team's agent made in one codebase, measured by whether those failures recur. A file of non-obvious project deviations might do the second job while adding nothing to the first — the study's own developer-written condition trends that way (+2.4%, not significant). If that is right, the technique's value is in what goes in the file, not in having one. Not tested against each other.

Status: pending-reviewLast checked: 2026-09-11Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Filed the day the guide-file technique was, because the first measured study of it cuts against the practitioner consensus it came from. That is the catalogue working as intended: the technique record says what a guide file is, and this claim says the evidence on whether it helps is split, and along which axis.