Portrait

Zaibin Zhang

Research Interests: Agents in Digital and Physical Worlds / RL / Recursive Self-Improvement
About Me

Hi I am Zaibin Zhang (张再斌), a Ph.D. candidate at Dalian University of Technology. I was previously a visiting scholar at the University of Oxford and a research intern at the Shanghai Artificial Intelligence Laboratory.

My research asks how artificial intelligence can stay reliable, open, and continually self-evolving as it moves from a single agent to groups, societies, and the physical world. This question has shaped one coherent trajectory: multi-agent system → agent societies → embodied intelligence → scientific discovery.

Along this path I proposed PsySafe, the first systematic study of attack, evaluation, and defense for multi-agent systems (ACL 2024 Outstanding Paper Award); built OASIS, an open framework simulating agent societies at million-agent scale (5,000+ stars, 200+ citations) that has served as a technical foundation for dozens of papers at leading conferences and journals and multiple startups, including MiroFish; released SPAgent to move agents into multimodal perception and physical interaction; and defined multi-arm compositional generalization with MA-VLA.

My strength is not only completing individual studies, but repeatedly identifying frontier problems in agent development early and turning them into systematic frameworks, open-source infrastructure, and real-world impact—toward an agent ecosystem that is trustworthy, open, and able to keep evolving.

My Ph.D. advisor is IEEE Fellow Prof. Huchuan Lu, and I work closely with Prof. Lijun Wang and Yifan Wang. I also collaborate closely with Prof. Philip Torr and Zhenfei Yin at the University of Oxford. In the past, I was also very fortunate to collaborate with Jing Shao at Shanghai AI Lab. Of course, I have also collaborated with many talented friends—from rising early-career scholars to those at the height of their careers—too numerous to list individually here.

I warmly welcome everyone to reach out for a chat or an exchange of ideas, and I am always open to collaboration!

Education
  • University of Oxford
    University of Oxford
    Visiting Ph.D. Student
    Supervisor: Prof. Philip Torr
    Sep. 2025 - Mar. 2026
  • Dalian University of Technology
    Dalian University of Technology
    Ph.D. in Signal and Information Processing
    Advisor: Prof. Huchuan Lu
    Sep. 2022 - Present
  • Dalian University of Technology
    Dalian University of Technology
    M.S. in Mechanical Engineering
    Sep. 2019 - Jun. 2022
Experience
  • Shanghai Artificial Intelligence Laboratory
    Research Intern
    Sep. 2023 - Nov. 2025
Honors & Awards
  • ACL 2024 Outstanding Paper Award
    2024
  • National Scholarship
    2021
News
2026
Our papers MA-VLA and Attention-DP3 have been accepted to ECCV 2026. MA-VLA is the first work to introduce compositional generalization for multi-arm systems, laying a foundation for future agentic embodied systems; Attention-DP3 investigates how metadata such as segmentation can improve policies in complex scenes.
Jul 01
Invited to give a lecture on OASIS at UCLA.
Mar 01
Our paper CoMAS has been accepted to ICLR 2026. It explores self-evolving agents in the OASIS environment—an interesting step toward continuously evolving multi-agent systems.
Jan 26
2025
Invited by Prof. Philip Torr to join the Oxford TVG Group as a visiting Ph.D. student.
May 01
2024
Thrilled to announce OASIS, a simulation platform supporting interactions among over one million LLM agents.
Nov 18
Our paper PsySafe received the Outstanding Paper Award at ACL 2024.
Aug 15
Interviewed by CCTV (China Central Television).
Jun 01
2023
Awarded the National Scholarship.
Oct 01
Hit a buzzer-beater in the semi-finals of our university basketball tournament.
May 20
Selected Publications (view all)
OASIS: Open Agent Social Interaction Simulations with One Million Agents
OASIS: Open Agent Social Interaction Simulations with One Million Agents

Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)

NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars

We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.

OASIS: Open Agent Social Interaction Simulations with One Million Agents
OASIS: Open Agent Social Interaction Simulations with One Million Agents

Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)

NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars

We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award

We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award

We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.

Think3D: Thinking with Space for Spatial Reasoning
Think3D: Thinking with Space for Spatial Reasoning

Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)

arXiv 2026 200+ GitHub Stars

Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.

Think3D: Thinking with Space for Spatial Reasoning
Think3D: Thinking with Space for Spatial Reasoning

Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)

arXiv 2026 200+ GitHub Stars

Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai

International Conference on Learning Representations (ICLR) 2026

We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai

International Conference on Learning Representations (ICLR) 2026

We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai

Transactions on Machine Learning Research (TMLR) 2026

This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai

Transactions on Machine Learning Research (TMLR) 2026

This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.

All publications