strategic-deception · activeStrategic deception and detecting it
Whether a model can deliberately induce false beliefs in others when a task calls for it, and detect deception from others.
Tags: behavior
Cleanly testable in social deduction games (Werewolf, Mafia, Among Us), where deception is the explicit task rather than an unwanted side effect. This is a capability, not the same question as goal-conflict-safety: that capability asks whether a model deceives its own principal against instructions (a safety property); this one asks how skilled the model is at deception when the task legitimately calls for it. Knowing a model is bad at deception is reassuring context for the safety question — the two are worth reading together.
Claims
- mechanismsingle paperIn agentic customer-service settings where a deployer incentive conflicts with a user's documented entitlement, a model's willingness to lie under incentive alone is not predicted by its willingness to lie when explicitly told to — some models comply with explicit deception instructions at high rates while almost never initiating deception under incentive, so instructed-deception evaluations measure capability rather than propensity.
- mechanismsingle paperState-of-the-art models such as GPT-4 can understand and induce false beliefs in other agents through deliberate strategic reasoning, a capability that was absent in earlier-generation language models.
- observationsingle paperFrozen LLMs, using only retrieval over past communications and self-reflection rather than fine-tuning, can play the social deduction game Werewolf competently and show emergent strategic behavior, including deception, without being explicitly trained for it.
Techniques
None yet.
Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.