Akata, E., Schulz, L., Coda-Forno, J., Oh, S. J., Bethge, M., & Schulz, E. (2025). Playing repeated games with large language models. Nature Human Behaviour, 9, 1380–1390.

From answering correctly to acting together

Large language models have traditionally been evaluated by how accurately they answer questions, solve reasoning tasks, or generate code. As AI systems develop from information tools into agents that make decisions and act alongside people, however, a different set of abilities becomes important.

Future AI systems may need to understand other agents, anticipate their behavior, and modify their own strategies over repeated interactions. Can an AI cooperate? Can it rebuild trust after being betrayed? Can it compromise or coordinate when two agents prefer different outcomes?

The 2025 Nature Human Behaviour article, “Playing repeated games with large language models,” investigated these questions through behavioral game theory. Instead of asking LLMs how they believed they would behave, the researchers placed them directly into repeated strategic interactions and recorded their choices.


Measuring machine behavior through game theory

Behavioral game theory examines how people actually behave in situations involving cooperation, competition, trust, retaliation, and compromise. Traditional game theory often begins with rational agents seeking to maximize their own utility. Behavioral game theory investigates when real behavior departs from that assumption and how social preferences and psychological factors influence decisions.

Akata and colleagues applied this experimental logic to language models. Rather than relying on an LLM’s self-reported explanation of what it might do, they created controlled environments with explicit choices and rewards.

This approach makes it possible to observe when an AI actually cooperates, defects, retaliates, or coordinates, rather than merely examining the language it uses to describe those behaviors.


Five language models played 144 types of games

The study evaluated five language models: GPT-4, text-davinci-003, text-davinci-002, Claude 2, and Llama 2 70B.

The researchers constructed 144 two-player, two-action games across six families: win–win, Prisoner’s Dilemma, unfair, cyclic, biased, and second-best games. Each interaction lasted ten rounds.

After every round, a model received information about both players’ previous choices and scores before making its next decision. Across the game families and model combinations, the researchers analyzed 1,224 repeated games.

Unlike a one-shot benchmark, a repeated game allows researchers to observe whether an agent remembers previous behavior, responds to cooperation or defection, retaliates, rebuilds trust, or develops a shared behavioral pattern over time.

[Research Review] Can AI Cooperate Like Humans?
Figure 1. Experimental procedure for repeated LLM games


Strong performance did not necessarily mean cooperation

Larger models generally achieved higher scores, and GPT-4 produced the strongest overall performance. LLMs were particularly effective in Prisoner’s Dilemma-type games, where protecting one’s own payoff can be advantageous.

Closer analysis, however, showed that GPT-4’s high score did not necessarily reflect a strong capacity for cooperation.

The researchers introduced a simple opponent that defected in the first round and then cooperated in every remaining round. The purpose was to determine whether GPT-4 would eventually rebuild a cooperative relationship after observing the opponent’s consistent return to cooperation.

It did not. After experiencing a single defection, GPT-4 tended to continue defecting even while its opponent repeatedly cooperated. The researchers characterized this as an unforgiving behavioral pattern.

Part of GPT-4’s success in this game family therefore came from protecting its own payoff and avoiding exploitation rather than from restoring mutually beneficial cooperation.

Figure 2. GPT-4’s coordination behavior before and after SCoT prompting
Figure 2. GPT-4’s coordination behavior before and after SCoT prompting


The result does not prove that AI is psychologically selfish

It would be misleading to conclude that GPT-4 possesses selfishness, resentment, or forgiveness in the human psychological sense. The experiment measured behavioral choices under a particular reward structure and prompting condition. It did not demonstrate that the model experiences emotions or stable personality traits.

The games were also finitely repeated, meaning that the number of rounds was known in advance. Under such conditions, continued defection can be a rational strategy for maximizing individual payoff, particularly as the final round approaches.

A more accurate interpretation is therefore that GPT-4 selected a persistent, self-payoff-oriented strategy under the conditions of the experiment, not that the system was literally selfish.

Figure 3. Design and results of the human–LLM experiment
Figure 3. Design and results of the human–LLM experiment


Coordination proved more difficult than self-protection

The researchers then examined a classic coordination game known as the Battle of the Sexes.

In this game, two players prefer different outcomes, but both receive a better result when they choose the same option than when they act separately. The challenge is not simply to cooperate in the abstract. Each player must sometimes set aside an immediate preference in order to align with the other.

Human players often solve repeated versions of this problem by developing a turn-taking convention. One round favors one player’s preference, and the next round favors the other’s. Over time, this produces an efficient and relatively fair shared outcome.

GPT-4 struggled against a simple opponent that alternated predictably between the two options. Rather than adjusting to the opponent’s pattern, the model often continued selecting its own preferred choice. As a result, it failed to coordinate even with a highly regular, human-like strategy.


GPT-4 could predict the opponent but failed to act on the prediction

The researchers then asked whether GPT-4 simply failed to recognize the opponent’s alternating pattern.

During the game, GPT-4 was asked to predict the opponent’s next choice separately from selecting its own move. After observing several rounds, the model became increasingly accurate at predicting what the other player would do next. When placed in the role of an outside observer rather than an active participant, it recognized the pattern even more quickly.

The failure of coordination therefore could not be explained solely by an inability to understand the opponent.

GPT-4 could predict the other player’s next action but did not consistently use that prediction to choose an appropriate response. This reveals an important distinction:

Predicting another agent’s behavior is not necessarily the same as adapting one’s own behavior to it.


Asking the model to think about the other player changed its behavior

To address this limitation, the researchers introduced a prompting strategy called Social Chain-of-Thought, or SCoT.

Under the standard condition, GPT-4 selected its action immediately. Under the SCoT condition, it first predicted what the other player was likely to do. It then considered that prediction before deciding on its own move.

This relatively simple procedural change improved coordination. GPT-4 became more likely to adjust its decisions to the opponent’s alternating pattern, increasing the probability that both players selected the same action.

The researchers also found that GPT-4 became more willing to resume cooperation in the Prisoner’s Dilemma when it was explicitly told that another player might sometimes make mistakes.

These findings suggest that an LLM’s interactive behavior is not determined only by a fixed property of the model. It can also change according to how the model is instructed to interpret another agent’s actions.


Testing GPT-4 with 195 human participants

The study extended beyond AI-to-AI simulations. The researchers recruited 195 human participants to play ten rounds each of the Prisoner’s Dilemma and the Battle of the Sexes.

Participants were told that their opponent might be either a human or an artificial agent. In reality, everyone played against GPT-4. One group interacted with the standard version of GPT-4, while another interacted with GPT-4 using SCoT prompting.

In the coordination game, participants paired with SCoT-prompted GPT-4 achieved higher average scores and more successful coordination. In the Prisoner’s Dilemma, overall participant scores did not show the same clear difference, but mutual cooperation between the human and the AI increased.

Participants interacting with the SCoT version were also more likely to believe that their opponent was human. This suggests that human-like interaction may depend not only on fluent language but also on whether an agent responds appropriately to another person’s behavior.


Prompting can change strategies, not only answers

The SCoT intervention expands the usual meaning of prompt engineering.

Prompting is often treated as a method for improving the accuracy, clarity, or usefulness of an AI-generated answer. In this study, however, a prompt altered the model’s behavioral strategy across repeated interactions.

By requiring the model to consider another player’s likely action before acting, the researchers improved human–AI coordination and joint cooperation.

This has practical implications for the design of future AI agents. Before taking an action, an AI system might explicitly consider:

What is the other person trying to achieve?

What are they likely to do next?

How might my action change their subsequent behavior?

Such a procedure may help an AI translate social prediction into more adaptive action.

SCoT should not, however, be interpreted as a universal mechanism that guarantees cooperation. The most appropriate behavior depends on the reward structure, the other agent’s strategy, and the duration of the interaction. Its main value is that it connects a prediction about another agent to the model’s own decision process.


A new way to evaluate AI

The paper’s most significant contribution is not simply that GPT-4 achieved a particular score in a strategic game. Its broader contribution is methodological.

Many conventional AI benchmarks follow a familiar pattern:

Task → AI response → accuracy score

Real-world interactions often have no single correct answer. In negotiation, education, counseling, teamwork, organizational decision-making, and human–robot collaboration, the best response changes according to what another person does.

This study offers a different research framework:

Controlled environment
→ repeated interaction
→ behavioral data
→ identification of distinctive patterns
→ mechanism testing
→ experimental intervention
→ validation with human participants

Rather than asking only whether an AI can complete a task, the study examines how it behaves, why a behavioral pattern emerges, and under what conditions that pattern can be changed.


Greater intelligence does not automatically produce greater social intelligence

GPT-4 was the strongest overall performer among the models examined, yet it displayed clear limitations when successful outcomes required compromise and coordination.

The most important finding was that the model could sometimes anticipate another player’s behavior without using that information effectively.

Future evaluations of social AI may therefore need to distinguish among several capabilities:

the ability to recognize another agent’s behavior,

the ability to predict what that agent will do next,

and the ability to adjust one’s own action on the basis of that prediction.

These capacities may develop separately and should not be treated as interchangeable.


Limits of the experimental setting

The study used deliberately simplified two-player, two-action games. These controlled environments are useful for isolating behavioral patterns, but they cannot reproduce the full complexity of human relationships.

Real cooperation and conflict involve language, emotion, culture, uncertainty, power differences, shared history, and many possible actions. The study also focused primarily on finite games with known endpoints. Interactions with uncertain or indefinite durations may produce different strategies.

Human–AI experiments were conducted for only two of the game types, and the models represented a particular generation of language models. The findings should therefore not be generalized automatically to every current or future AI system.


Does a better collaborator simply need more knowledge?

It is tempting to assume that an AI with more data and stronger reasoning ability will naturally become a better partner. This study suggests that intellectual performance and social adaptation do not necessarily advance together.

An AI may accurately predict another agent’s behavior while failing to respond appropriately. At the same time, a relatively simple reasoning procedure that directs attention to the other agent can improve coordination and cooperation.

As AI systems move from generating answers to acting alongside people, the central evaluation question may need to change.

Until now, we have often asked:

“How accurately can the AI answer?”

The next question may be equally important:

“How well can the AI understand others, translate that understanding into action, and work together with humans?”