From reward models
to reward worlds.
What if the environment an agent acts in could also generate the signal it learns from?
Many post-training pipelines ask a model or a human to score an answer. With OASIS Social Reward, we explore a different source of feedback: the responses of other agents in a simulated society.
In OASIS, agents create posts, comment, repost, like content, and follow one another.[1] These interactions produce feedback at scale. We turn that feedback into reward, use reinforcement learning to update the policy, and return the updated agents to the environment.
The social simulator becomes both a place to act and a source of reward.
This is an early exploration. The experiments show adaptation to socially generated feedback and mixed transfer to external benchmarks. They also reveal a central challenge: what a society rewards shapes what its agents learn.
Interact. Learn. Return.
Instructions enter a social environment. As agents interact, the simulator produces trajectories—prompts and model outputs—along with rewards derived from social feedback. Proximal Policy Optimization (PPO)[2] then updates the policy parameters.
The updated agents return to the simulator and generate the next round of experience. Both the agents and the feedback they create become part of an evolving learning environment.
Can a model adapt
to social feedback?
We started with a simulated society of 196 LLM agents. Agents repeatedly posted and reacted to one another. Likes were used as a simple proxy for social influence, allowing the simulator to generate rewards without a new human preference label for every action.
A qualitative example offers a first look at how the policy changes. Across two Social RL epochs, the responses become more directly addressed to an audience and more oriented toward interaction.
“Feeling overwhelmed with homework? Take a deep breath and break it down into smaller tasks! You got this! #homeworkhelp #studytips”
“Hey students! Don't let homework stress you out! Break it down into smaller tasks, take regular breaks, and ask for help when you need it. Remember, every assignment is an opportunity to learn and grow.”
“Hey friends! I know homework can be overwhelming, but don't forget to take breaks and stay focused… Keep pushing through and don't hesitate to ask for help when you need it.”
More participation.
More engagement.
Inside the simulated society, we observed a smaller fraction of zero-like posts and a larger fraction receiving positive engagement. Average likes per post increased from 0.6767 to 0.7065, while total posts rose from 43,443 to 52,790.
Relative change
Relative change
Initial population
Does learning transfer
beyond the simulator?
We evaluated checkpoints on coding and reasoning benchmarks, including HumanEval[4], MBPP[5], and BIG-Bench Hard (BBH)[6], as well as instruction-following evaluation with AlpacaEval 2[7]. The picture is mixed: some scores improve, while others remain flat or decline.
In the general social environment, HumanEval rises from 62.2 to 65.24, while Math and AlpacaEval 2 decline slightly. In a code-oriented environment, the reported BBH average rises from 52.4 to 64.5. These observations motivate closer study of how the social environment shapes transfer.
View the exact benchmark values
| Benchmark | Base | After Social RL | Change (pp) |
|---|---|---|---|
| General social environment | |||
| HumanEval | 62.2 | 65.24 | +3.04 |
| MBPP | 57.8 | 58.2 | +0.40 |
| Math | 27.7 | 27.46 | -0.24 |
| AlpacaEval 2 | 22.92 | 20.25 | -2.67 |
| Code-oriented environment | |||
| HumanEval | 62.2 | 64.0 | +1.80 |
| MBPP | 57.8 | 58.2 | +0.40 |
| BBH avg. | 52.4 | 64.5 | +12.10 |
| Math 0-shot | 27.7 | 28.2 | +0.50 |
These are exploratory checkpoint comparisons. Repeated-run variation and confidence intervals are not available for these preliminary results, so small score differences should be interpreted cautiously.
Social influence is not
the same as correctness.
A popular answer can still be wrong. An agent can learn to attract attention without becoming more useful.
Once agents learn from one another, these incentives can compound. More reward inside a society does not necessarily translate into stronger capabilities outside it.
Early signs of reward hacking
Our early experiments suggest that increasing interaction data in a single RL stage from roughly 30k to 60k samples can bring coding performance back toward the base model. Output formatting also becomes less reliable, sometimes causing action parsing to fail.
The next objective is to reward useful contribution: helping another agent solve a problem, producing a verifiable solution, improving a collective outcome, or making collaboration more effective.
From social influence
to social contribution.
Likes were a deliberately simple starting point. A richer society offers other signals: whether code is reused, whether a discussion converges on a correct answer, or whether a group solves a problem that its members could not solve alone.
Social influence
Study how likes, reposts, comments, and participation change behavior.
Task-grounded reward
Connect social feedback to verification, correctness, collaboration, and task success.
Continual societies
Interact, update, and return to an evolving multi-agent environment.
Social dynamics and safety
Study herd effects, misinformation, polarization, and emergent norms.
OASIS Social Reward does not yet establish social feedback as a universally reliable route to stronger models. It does show that agents adapt to socially generated rewards, that some learning can transfer beyond the simulator, and that the design of the environment matters.
The longer-term direction is to place agents in worlds where new experience is continuously created—and investigate whether those worlds can support useful, sustained learning.
Research status: early-stage exploration. Results and examples are based on the project’s internal presentation materials. All reported results are preliminary. Network connections are conceptual.
References
These sources describe the simulator, algorithm, base model, and evaluation tools. The Social Reward results reported here come from the project’s internal experiments, not from these references. The source materials label the math evaluations as “Math” and “Math 0-shot” without specifying the dataset or version.
- Ziyi Yang, Zaibin Zhang, et al. OASIS: Open Agent Social Interaction Simulations with One Million Agents. 2024. ↩
- John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal Policy Optimization Algorithms. 2017. ↩
- Meta. Meta-Llama-3-8B-Instruct: Model Card. 2024. ↩
- Mark Chen et al. Evaluating Large Language Models Trained on Code. 2021. ↩
- Jacob Austin et al. Program Synthesis with Large Language Models. 2021. ↩
- Mirac Suzgun et al. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. 2022. ↩
- Tatsu Lab. AlpacaEval: Official Repository and AlpacaEval 2.0 Documentation. ↩
Cite this post
OASIS Project. “Can LLM agents learn from society?” OASIS Social Reward, September 2026. Research blog.
@misc{oasis2026socialreward,
author = {{OASIS Project}},
title = {Can {LLM} agents learn from society?},
year = {2026},
month = sep,
howpublished = {OASIS Social Reward research blog},
url = {https://zhangzaibin.github.io/evolving-machines/research/oasis-social-reward/}
}