2026

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Zaibin Zhang*, Junlan Xiao*, Zhongbo Zhang*, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

MA-VLA studies multi-arm embodied collaboration through structured atomic action assignment. Instead of treating language as a single global instruction, it decomposes cooperative behavior into mid-level atomic prompts and assigns them to individual arms, enabling explicit division of labor, role-agnostic execution via Arm Shuffle, and compositional reuse across unseen collaboration patterns in simulation and real-world evaluations.

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Changbo Yan*, Zhongbo Zhang*, Zaibin Zhang*, Yifan Wang, Lijun Wang, Huchuan Lu (* equal contribution)

European Conference on Computer Vision (ECCV) 2026

Attention-DP3 introduces a spatially object-aware 3D diffusion policy for manipulation in cluttered scenes. It lifts open-vocabulary 2D segmentation masks into 3D and injects geometry-aligned object cues through Tri-field Attentional Conditioning, helping the policy attend to target geometry, suppress distractors, and remain robust as scene clutter increases.

Think3D: Thinking with Space for Spatial Reasoning
Think3D: Thinking with Space for Spatial Reasoning

Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)

arXiv 2026 200+ GitHub Stars

Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.

Think3D: Thinking with Space for Spatial Reasoning
Think3D: Thinking with Space for Spatial Reasoning

Zaibin Zhang*, Yuhan Wu*, Lianjie Jia*, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu (* equal contribution)

arXiv 2026 200+ GitHub Stars

Think3D equips VLM agents with interactive 3D chain-of-thought reasoning by combining 3D reconstruction, ego/global-view switching, and camera-based spatial manipulation. It improves spatial reasoning in a training-free setting and reveals emergent 3D exploration strategies in RL-trained open-weight models.

2025

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai

International Conference on Learning Representations (ICLR) 2026

We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai

International Conference on Learning Representations (ICLR) 2026

We propose CoMAS, a framework for self-evolution of LLM-based agents through interaction rewards. Agents interact in a forum-like setting, evaluate each other, and optimize policies from self-produced reward signals, enabling continuous co-evolution without external supervision.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai

Transactions on Machine Learning Research (TMLR) 2026

This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhong-Zhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai

Transactions on Machine Learning Research (TMLR) 2026

This survey formalizes agentic reinforcement learning as the shift from single-step LLM reinforcement learning to temporally extended decision-making in dynamic environments. Synthesizing more than 500 works, it organizes the field by agent capabilities and applications while cataloging open-source environments, benchmarks, and frameworks.

2024

OASIS: Open Agent Social Interaction Simulations with One Million Agents
OASIS: Open Agent Social Interaction Simulations with One Million Agents

Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)

NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars

We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.

OASIS: Open Agent Social Interaction Simulations with One Million Agents
OASIS: Open Agent Social Interaction Simulations with One Million Agents

Ziyi Yang*, Zaibin Zhang*, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, Jing Shao (* equal contribution)

NeurIPS 2024 Workshop on Open-World Agents 2024 5000+ GitHub Stars

We propose OASIS, a generalizable and scalable social media simulator based on real-world platforms. OASIS supports large-scale simulations with up to one million users, featuring dynamic environments, diverse action spaces, and recommendation systems. We replicate various social phenomena including information spreading, group polarization, and herd effects.

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award

We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Zaibin Zhang*, Yongting Zhang*, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL) 2024 Outstanding Paper Award

We explore multi-agent system safety through the lens of agent psychology, revealing that dark psychological states constitute significant threats. We propose PsySafe, a framework focusing on identifying risky behaviors from dark personality traits, evaluating safety from psychological perspectives, and devising mitigation strategies.