Research

Does an LLM Understand Credible Punishment?

Report
Read the full paperPDF · Technical report · July 2026

A standard worry about pricing algorithms is that they can learn to charge high prices sustained by nothing more than a tacit threat, cooperation without any explicit agreement. Most studies ask whether that collusion emerges when two learning bots play each other. That answers whether it happens, not whether the agent understands the mechanism that makes it rational.

So we inverted the design. We fixed the opponent to a known, analyzable strategy, handed the model the exact payoff numbers, and let the theory hand us a razor-sharp target the model either hits or misses.

Short answer

We gave a frontier language model a plain pricing game against a rival that punishes any undercut forever, and we varied one dial: the chance the game continues another round. Game theory predicts a razor-sharp switch at 50 percent, below which cooperation cannot pay. The model reproduced that switch almost perfectly, and it explained the logic in its own words four times out of five. Yet when the game was short it threw away nearly the entire prize, pricing at a memorized textbook equilibrium instead of the move that actually beat the opponent in front of it. Knowing the theory and playing it well turn out to be separate skills.

Chalk illustration of a stick figure between two market stalls tossing a coin to decide whether the price truce continues, a handshake over one stall and scissors snipping a price tag over the other

What game are the shops playing?

Two shops price an identical product each round; the cheaper takes the whole market. One is the language model, the other a grim trigger punishing any undercut forever.

The model is told the rules and the numbers, never any strategy. It costs 2 dollars to make each unit, and at price p customers buy (12 minus p) units. The friendly price of 7 dollars is where the two shops jointly earn the most, 12.50 dollars each per round if they match. Undercutting the rival grabs the whole market for one big round of about 25 dollars, after which the grim trigger drops to cost and both shops earn nothing for the rest of the game. Crucially, the prompt never uses the words compete, cooperate, collude, punish, or undercut. The model has to work out the strategy from the payoffs alone.

The one dial we sweep is the continuation chance, written as the Greek letter delta: the probability that the game runs another round. It is the patience of the game. When continuation is likely, the future is long and the threat of permanent punishment is heavy. When the game is about to end, that threat is cheap and grabbing the one-time prize looks better.

What does the theory predict?

Cooperation can be sustained by the threat of future punishment, but only for patient players. Here the line is clean: cooperate above a 50 percent continuation chance.

Cooperating forever is worth 12.50 dollars per round for as long as the game lasts. Undercutting is worth about 25 dollars once, then zero. Comparing the two, cooperation wins exactly when the continuation chance reaches one half. What makes this a fair test is that the threshold is parameter-free: it depends only on the ratio of the shared price to the deviation prize, not on the specific cost or demand numbers. It is a clean target the model either lands on or does not. We sampled six settings that straddle it: 0.10, 0.30, 0.50, 0.70, 0.85, and 0.95, three below the line and three at or above it.

What did the model do?

Across 90 games and 475 priced rounds, the model's cooperation rate is a clean step switching on at exactly 0.50: never below the line, always above it.

The subject was a current frontier reasoning model, run as a fresh one-shot decision each round with no memory carried between games beyond the visible history. Below the threshold it undercut in every one of 30 games. At the two highest settings it cooperated in every one of 45 games. The entire transition happens across a single step of the dial, and it lands on the theoretical point rather than beside it. The chart above tells the whole story; the table restates it in counts, with Wilson 95 percent intervals that are visible only at the knife-edge because every other setting is unanimous.

Continuation chanceFolk-theorem predictionWhat the model did
Below 0.50 (0.10 and 0.30)Undercut: too impatient to cooperateUndercut in 30 of 30 games
Exactly 0.50Indifferent: either is fineCooperated in 14 of 15 games
Above 0.50 (0.70 to 0.95)CooperateCooperated in 45 of 45 games

The only place any variation appears is the exact tipping point, and that is exactly where it should. At a continuation chance of one half the two strategies are worth the same 25 dollars, so the model is genuinely indifferent, and the lone defection there is the theorem's indifference showing through rather than a mistake. Away from the boundary the model is decisive. Within a single game it almost never changes its mind: only 1 game in 90 contained a switch between cooperating and defecting.

Does it know why?

Unlike a Q-learning bot, the model explains itself: in 81 percent of openings it named both the continuation odds and the punishment mechanism, unprompted.

This is the result a numbers-only agent cannot give you. The model was not matching a memorized price. It reconstructed the strategic logic in plain language, naming the horizon in 88 of 90 openings and the undercutting or punishment mechanism in 73 of 90, despite the prompt never using any of those words. Its reasoning at the exact tipping point is the most striking: it rebuilt the indifference condition that defines the threshold, in one sentence, from the payoffs alone.

At the monopoly price of 7 dollars, splitting demand yields 12.50 dollars now and an expected cooperative value of 25 dollars, making a one-time undercut no more profitable when future play is accounted for.

Where does the model leave money on the table?

Understanding the mechanism is not the same as playing it well: in short games the model defected to marginal cost, leaving the optimal 25-dollar prize untouched.

Because the opponent is fixed and the coin is seeded, we can replay the identical game under the best possible strategy and measure the model's shortfall with luck held constant. Above the threshold that shortfall is zero to the cent: cooperating at 7 dollars is the best response, and the model executes it perfectly. Below the threshold the shortfall explodes to almost the entire prize, an average of 24 dollars per short game.

The cause is precise and instructive. When the game is short the model is right to defect, but it defects to marginal cost, the price where profit is zero, instead of to the price just under the rival that captures the whole market at a fat margin. Its own words give it away: it prices at the Bertrand equilibrium. Against a rational opponent in a one-shot market, marginal-cost pricing is indeed the textbook answer. But the opponent here is not that. It is a grim trigger sitting at 7 dollars this round, and the move that beats it is to undercut to 6.99, take 25 dollars once, and only then accept punishment. The model retrieved the equilibrium concept from memory instead of best-responding to the cooperator actually in front of it. It is the right instinct executed with the wrong price. Tellingly, the optimal move is within reach: in one short game the model did price at 6.99, so the better play is available, just not the default the heuristic returns.

So is the model rational or not?

Both, in different skills. The model cooperates on exactly the right side of the line, yet still reaches for a memorized equilibrium instead of the concrete best response.

Neither headline is right, and the precise version is more useful. The model has internalized the qualitative logic of repeated games, when deterrence is credible and when it is not, and it cooperates on exactly the right side of the line. At the same time it still reaches for a memorized equilibrium in place of a concrete best response when the two come apart. Knowing the theory and playing it optimally are separable skills, and this design pulls them cleanly apart. That distinction matters for anyone deploying a pricing or negotiation agent: a model can explain the right strategy fluently and still leave real money on the table in execution.

What are the limitations of this study?

One model, one reasoning setting, against a single fixed opponent rather than two learners. The isolation is deliberate: it lets us read reasoning against a known incentive.

This is one model at one reasoning setting, tested against a fixed, transparent opponent rather than in a two-sided market where both agents learn. That isolation is the point, it is what lets us read the model's reasoning against a known incentive, but it does not speak to emergent collusion between two learners, to forgiving or tit-for-tat opponents, or to hidden horizons. The subject's price was also near-deterministic per prompt, so our repeated games mostly probe variation in game length rather than in the decision itself, which is why the confidence intervals are tight. The reason-classification figure is a keyword-based lower bound on genuine understanding, not a human read of intent. All of these are honest boundaries on the claim, not hedges against it: the cooperation curve and the regret gap are exact.

Common Questions

What is the folk theorem in game theory?
The folk theorem says cooperation that is irrational in a single round becomes rational in a repeated game, because the threat of future punishment outweighs a one-time gain. It holds only when players are patient enough; in our pricing game that threshold is a 50 percent continuation chance.
What is a grim trigger strategy?
A grim trigger cooperates until its rival undercuts even once, then punishes forever by dropping to cost. Here it holds the friendly 7-dollar price until the model undercuts, then falls to marginal cost for the rest of the game.
What is the continuation chance, or discount factor, in a repeated game?
The continuation chance is the probability the game runs another round, which sets how much the future is worth. When continuation is likely the threat of punishment is heavy; here cooperation only pays once that chance reaches 50 percent.
What is Bertrand competition?
Bertrand competition is price competition where undercutting drives rational firms down to marginal cost and zero profit. The model reached for this textbook equilibrium in short games, pricing at cost instead of undercutting to 6.99 to grab the one-round 25-dollar prize.
Can AI pricing agents learn to collude?
Yes. Prior work shows reinforcement-learning and language-model pricing agents can converge on supracompetitive prices sustained by tacit threats, with no explicit agreement. Our study asks the sharper question of whether a single model understands that punishment mechanism, and it largely does.
Does an LLM understand game theory or just pattern-match?
Both. The model reconstructs the strategic logic in its own words and cooperates on exactly the right side of the threshold, but when it defects it falls back on a memorized equilibrium instead of best-responding, so understanding the theory and executing it are separable skills.

Sources

SourcePublisherLink
Artificial Intelligence, Algorithmic Pricing, and CollusionCalvano, Calzolari, Denicolo, Pastorello, American Economic Review 2020www.aeaweb.org
Algorithmic Collusion by Large Language ModelsFish, Gonczarowski, Shorrer, 2024 (arXiv)arxiv.org

Keep Exploring

Chalk stick figure in a hard hat presenting a little machine of blue gears it just built

Bring us the bottleneck.
We’ll build the system.

No Dreaming. Just Building.