From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
A causal framework separates deceptive-looking model output from the mechanism producing it, giving agent evaluators a stricter basis for claims about intent or agency.
The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.
When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.
The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings. When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled. Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.
This narrows deception evaluation from detecting suspicious outputs to testing whether a proposed deceptive mechanism actually caused them. It reinforces counterfactual intervention as an audit method and warns that stronger elicitation alone cannot establish intent: recipient knowledge, internal preference, sampled output, and objective origin must be separated, and even positive mechanism evidence does not establish agency.