OASISSocial Reward
Research update

Can LLM agents
learn from society?

Exploring social interaction as a source of reward.
Early experiments in a world where agents learn from one another.

A society as a learning signalConceptual social network, with 196 nodes representing the initial population. Interactions generate feedback, which updates the policy before agents return to the environment. Connections are schematic.OASIS / SOCIAL REWARDA SOCIETY AS A LEARNING SIGNALMany agents.A shared world.Posts, comments, likes and reposts.Every interactioncan become feedback.Feedback becomes a policy update.01Interact02Observe reward03Update policy04Return to society
A society as a learning signalConceptual network of interacting agents. Feedback becomes a policy update.A SOCIETY AS A LEARNING SIGNALInteract. Learn. Return.Social feedback becomes a policy update.
A conceptual social network. Nodes represent agents; connections are schematic.

From reward models
to reward worlds.

What if the environment an agent acts in could also generate the signal it learns from?

Many post-training pipelines ask a model or a human to score an answer. With OASIS Social Reward, we explore a different source of feedback: the responses of other agents in a simulated society.

In OASIS, agents create posts, comment, repost, like content, and follow one another.[1] These interactions produce feedback at scale. We turn that feedback into reward, use reinforcement learning to update the policy, and return the updated agents to the environment.

The social simulator becomes both a place to act and a source of reward.

This is an early exploration. The experiments show adaptation to socially generated feedback and mixed transfer to external benchmarks. They also reveal a central challenge: what a society rewards shapes what its agents learn.

Interact. Learn. Return.

Instructions enter a social environment. As agents interact, the simulator produces trajectories—prompts and model outputs—along with rewards derived from social feedback. Proximal Policy Optimization (PPO)[2] then updates the policy parameters.

The updated agents return to the simulator and generate the next round of experience. Both the agents and the feedback they create become part of an evolving learning environment.

01 / System architectureClick a figure to enlarge ↗
Figure 1. The Social Reward training loop. Each round collects trajectories and rewards, performs a PPO update, and returns updated agents to the social simulator. Likes serve as the initial reward signal.

Can a model adapt
to social feedback?

We started with a simulated society of 196 LLM agents. Agents repeatedly posted and reacted to one another. Likes were used as a simple proxy for social influence, allowing the simulator to generate rewards without a new human preference label for every action.

A qualitative example offers a first look at how the policy changes. Across two Social RL epochs, the responses become more directly addressed to an audience and more oriented toward interaction.

Base modelLlama-3-8B-Instruct[3]
“Feeling overwhelmed with homework? Take a deep breath and break it down into smaller tasks! You got this! #homeworkhelp #studytips”
Social RLEpoch 1
“Hey students! Don't let homework stress you out! Break it down into smaller tasks, take regular breaks, and ask for help when you need it. Remember, every assignment is an opportunity to learn and grow.”
Social RLEpoch 2
“Hey friends! I know homework can be overwhelming, but don't forget to take breaks and stay focused… Keep pushing through and don't hesitate to ask for help when you need it.”
Figure 2. Example generations from the initial experiments. This illustrates a change in style; it does not by itself establish an improvement in general capability.

More participation.
More engagement.

Inside the simulated society, we observed a smaller fraction of zero-like posts and a larger fraction receiving positive engagement. Average likes per post increased from 0.6767 to 0.7065, while total posts rose from 43,443 to 52,790.

+4.4%Average likes per post
Relative change
+21.5%Total posts
Relative change
196LLM agents
Initial population
03 / Engagement inside the simulatorClick a figure to enlarge ↗
Figure 3. Average likes per post and total posting volume before and after Social RL. Relative changes are calculated from the reported values. Both axes start at zero.

Does learning transfer
beyond the simulator?

We evaluated checkpoints on coding and reasoning benchmarks, including HumanEval[4], MBPP[5], and BIG-Bench Hard (BBH)[6], as well as instruction-following evaluation with AlpacaEval 2[7]. The picture is mixed: some scores improve, while others remain flat or decline.

In the general social environment, HumanEval rises from 62.2 to 65.24, while Math and AlpacaEval 2 decline slightly. In a code-oriented environment, the reported BBH average rises from 52.4 to 64.5. These observations motivate closer study of how the social environment shapes transfer.

04 / Transfer to external benchmarksClick a figure to enlarge ↗
Figure 4. Selected preliminary benchmark results. Numbers at the right of each benchmark show the after-minus-before difference in percentage points. The environments report different benchmark sets, so these results do not establish that one environment is universally better.
View the exact benchmark values
Preliminary external evaluation results
BenchmarkBaseAfter Social RLChange (pp)
General social environment
HumanEval62.265.24+3.04
MBPP57.858.2+0.40
Math27.727.46-0.24
AlpacaEval 222.9220.25-2.67
Code-oriented environment
HumanEval62.264.0+1.80
MBPP57.858.2+0.40
BBH avg.52.464.5+12.10
Math 0-shot27.728.2+0.50

These are exploratory checkpoint comparisons. Repeated-run variation and confidence intervals are not available for these preliminary results, so small score differences should be interpreted cautiously.

Social influence is not
the same as correctness.

A popular answer can still be wrong. An agent can learn to attract attention without becoming more useful.

Once agents learn from one another, these incentives can compound. More reward inside a society does not necessarily translate into stronger capabilities outside it.

Early signs of reward hacking

Our early experiments suggest that increasing interaction data in a single RL stage from roughly 30k to 60k samples can bring coding performance back toward the base model. Output formatting also becomes less reliable, sometimes causing action parsing to fail.

Capability driftHigher social reward does not guarantee higher external benchmark scores.
Protocol driftGenerated outputs may violate the structured action format the environment expects.

The next objective is to reward useful contribution: helping another agent solve a problem, producing a verifiable solution, improving a collective outcome, or making collaboration more effective.

From social influence
to social contribution.

Likes were a deliberately simple starting point. A richer society offers other signals: whether code is reused, whether a discussion converges on a correct answer, or whether a group solves a problem that its members could not solve alone.

01

Social influence

Study how likes, reposts, comments, and participation change behavior.

02

Task-grounded reward

Connect social feedback to verification, correctness, collaboration, and task success.

03

Continual societies

Interact, update, and return to an evolving multi-agent environment.

04

Social dynamics and safety

Study herd effects, misinformation, polarization, and emergent norms.

What kinds of societies generate the right reward for learning?

OASIS Social Reward does not yet establish social feedback as a universally reliable route to stronger models. It does show that agents adapt to socially generated rewards, that some learning can transfer beyond the simulator, and that the design of the environment matters.

The longer-term direction is to place agents in worlds where new experience is continuously created—and investigate whether those worlds can support useful, sustained learning.

Research status: early-stage exploration. Results and examples are based on the project’s internal presentation materials. All reported results are preliminary. Network connections are conceptual.

References

These sources describe the simulator, algorithm, base model, and evaluation tools. The Social Reward results reported here come from the project’s internal experiments, not from these references. The source materials label the math evaluations as “Math” and “Math 0-shot” without specifying the dataset or version.

  1. Ziyi Yang, Zaibin Zhang, et al. OASIS: Open Agent Social Interaction Simulations with One Million Agents. 2024. ↩
  2. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal Policy Optimization Algorithms. 2017. ↩
  3. Meta. Meta-Llama-3-8B-Instruct: Model Card. 2024. ↩
  4. Mark Chen et al. Evaluating Large Language Models Trained on Code. 2021. ↩
  5. Jacob Austin et al. Program Synthesis with Large Language Models. 2021. ↩
  6. Mirac Suzgun et al. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. 2022. ↩
  7. Tatsu Lab. AlpacaEval: Official Repository and AlpacaEval 2.0 Documentation. ↩

Cite this post

OASIS Project. “Can LLM agents learn from society?” OASIS Social Reward, September 2026. Research blog.

@misc{oasis2026socialreward,
  author = {{OASIS Project}},
  title = {Can {LLM} agents learn from society?},
  year = {2026},
  month = sep,
  howpublished = {OASIS Social Reward research blog},
  url = {https://zhangzaibin.github.io/evolving-machines/research/oasis-social-reward/}
}