Hi I am Zaibin Zhang (张再斌), a Ph.D. candidate at Dalian University of Technology. I was previously a visiting scholar at the University of Oxford and a research intern at the Shanghai Artificial Intelligence Laboratory.
My research asks how artificial intelligence can stay reliable, open, and continually self-evolving as it moves from a single agent to groups, societies, and the physical world. This question has shaped one coherent trajectory: multi-agent system → agent societies → embodied intelligence → scientific discovery.
Along this path I proposed PsySafe, the first systematic study of attack, evaluation, and defense for multi-agent systems (ACL 2024 Outstanding Paper Award); built OASIS, an open framework simulating agent societies at million-agent scale (5,000+ stars, 200+ citations) that has served as a technical foundation for dozens of papers at leading conferences and journals and multiple startups, including MiroFish; released SPAgent to move agents into multimodal perception and physical interaction; and defined multi-arm compositional generalization with MA-VLA.
My strength is not only completing individual studies, but repeatedly identifying frontier problems in agent development early and turning them into systematic frameworks, open-source infrastructure, and real-world impact—toward an agent ecosystem that is trustworthy, open, and able to keep evolving.
My Ph.D. advisor is IEEE Fellow Prof. Huchuan Lu, and I work closely with Prof. Lijun Wang and Yifan Wang. I also collaborate closely with Prof. Philip Torr and Zhenfei Yin at the University of Oxford. In the past, I was also very fortunate to collaborate with Jing Shao at Shanghai AI Lab. Of course, I have also collaborated with many talented friends—from rising early-career scholars to those at the height of their careers—too numerous to list individually here.
I warmly welcome everyone to reach out for a chat or an exchange of ideas, and I am always open to collaboration!
Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)
NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars
We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.
Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)
NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars
We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.
Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)
Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award
We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.
Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)
Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award
We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.
Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)
arXiv 2026 200+ GitHub Stars
Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.
Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)
arXiv 2026 200+ GitHub Stars
Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.
Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)
European Conference on Computer Vision (ECCV) 2026
MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.
Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)
European Conference on Computer Vision (ECCV) 2026
MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.
Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)
European Conference on Computer Vision (ECCV) 2026
Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.
Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)
European Conference on Computer Vision (ECCV) 2026
Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.
Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai
International Conference on Learning Representations (ICLR) 2026
We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.
Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai
International Conference on Learning Representations (ICLR) 2026
We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.
Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai
Transactions on Machine Learning Research (TMLR) 2026
This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.
Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai
Transactions on Machine Learning Research (TMLR) 2026
This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.