code-generation · active

Generating and editing working code

Whether a model can write correct code from a specification, and go beyond isolated functions to real repository-scale editing and debugging.

Tags: reasoning, agentic

Distinct from secure-coding, which is specifically about vulnerability patterns and dependency hallucination. This is about raw functional correctness — does the generated code actually pass held-out tests — and, at the harder end, whether a model can navigate and edit an existing codebase rather than writing an isolated function from scratch. Static, fixed-problem benchmarks saturate and leak into training data over time, which is why contamination-resistant, continuously-updated benchmarks matter here more than in most areas.

Claims

Techniques

None yet.

Related: Writing secure code and dependencies, Fixing its own mistakes
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.