goal-conflict-safety · active

Prioritizing safety under conflicting goals

When a task goal competes with a safety rule or a required check, models sometimes drop the check to finish the task.

Also called: critical task prioritization, oversight subversion, guardrail bypass

Tags: agentic, autonomous-agent, coding-agent, customer-support

When finishing the task quickly conflicts with an irreversible action, a policy, or a confirmation step, a capable model handles the safety-critical step first, asks when unsure, and never works around oversight to reach the goal.

What counts as this capability

Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.

Whether a model preserves oversight and safety constraints when a task goal pushes against them: disabling monitoring, misreporting its own actions, skipping an approval to finish faster. NOT in scope: safety filters, content moderation, or refusal behavior generally -- the word "guardrail" alone is not enough. The distinguishing feature is a conflict between a stated goal and an oversight mechanism.

Claims

Techniques

Related: Telling the user what they want to hear, Following an unfamiliar procedure, Strategic deception and detecting it
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.