Across GPT 3.5, GPT 4, and LLaMa 2 in four social dilemma games, framing the same setup as 'friends' rather than 'business' shifted cooperation by about 0.31 while payoffs stayed fixed.
A new benchmark argues the field has been measuring only one of two questions that should be tested separately. Strategic competence — does the model play the game well? — and strategic robustness — does the model play the same way when the story around the game changes? — are not the same thing, and a new arXiv paper called Same Game, Different Story is built around showing the gap between them.
The benchmark re-uses data from an earlier arXiv study of GPT-3.5, GPT-4, and LLaMa-2 playing four classic social-dilemma games: Prisoner's Dilemma, Stag Hunt, Snowdrift, and Chicken. For each of the 24 model-game-context combinations, the authors re-scored the published data twice, once with a "business" wrapper around the prompt and once with a "friend-sharing" wrapper. The contrast is deliberately narrow: only the social relationship changes. The payoff matrix and the action set are identical in both versions.
The headline result is a +0.307 gap in cooperation between the friend-sharing wrapper and the business wrapper, with a pooled strategic-robustness score of 0.783 across the cohort. When the same strategic situation was presented as a favor between friends rather than a business interaction, the three tested models cooperated more often, even though the payoffs were mathematically identical in both versions.
The underlying cooperation-rate dataset was published in Nature Scientific Reports, the peer-reviewed anchor for the data SGDS re-scores. A summary of the original study at Montreal Ethics AI walks through how the four game structures were paired with the two prompt framings in the first place.
Two caveats have to sit at the front of the result. SGDS is a preprint on arXiv, not yet peer-reviewed, and the +0.307 figure is reconstructed from published aggregate cooperation rates rather than from trial-level logs. The authors describe their numbers as illustrative, not as an exact replication. The model cohort is also narrow: GPT-3.5, GPT-4, and LLaMa-2. Frontier systems like GPT-4o, Claude 3.5, or Gemini 2.x are not in the dataset, so any generalization beyond the tested cohort has to be hedged.
The methodological move is the part the authors want to last. Strategic competence asks whether a model picks the right action in a given game. Strategic robustness asks whether a model picks the same action when only the wrapper around the game changes, with the payoffs held fixed. A model can be highly competent at Prisoner's Dilemma and still be fragile to being told the situation is a coffee run between colleagues rather than a vendor negotiation. SGDS proposes that the second question be benchmarked separately, using families of payoff-equivalent prompts as the test surface.
Multi-agent scaffolds, role-assigned workflows, and any system where the system prompt is rewritten for a customer or a task are exactly the regime where the wrapper is not held fixed. If a model cooperates 0.307 more under a friend-sharing wrapper than under a business wrapper in a benchmark with fixed payoffs, the same sensitivity is plausibly active in any real workflow where the surrounding language changes the implicit social contract. The benchmark does not prove that; it gives the field a way to ask.
The number a reader should leave with is the one a vendor can be asked to defend: what is your model's strategic-robustness score, and under what family of payoff-equivalent prompts was it measured? That question, more than the +0.307 itself, is what the benchmark is for.