I am Sixu Yan (鄢思旭), a third-year Ph.D. student at the HUST Vision Lab (HUSTVL), Artificial Intelligence Institute, School of Electronic Information and Communications, Huazhong University of Science and Technology (HUST), advised by Prof. Xinggang Wang.
Prior to that, I received my M.S. degree from the School of Mechanical Engineering at Shanghai Jiao Tong University (SJTU), where I conducted research in the Robot Control and Machine Vision Lab (RCMVL) at the Institute of Robotics, under the supervision of Prof. Han Ding and Prof. Zhenhua Xiong.
My research goal is to develop general-purpose cognitive robots. Currently, I focus on scaling up robotic dexterous manipulation through large-scale synthetic data generation and sim-to-real transfer. My prior work includes motion planning and imitation learning. During my Ph.D., I have been fortunate to collaborate closely with Dr. Hangxin Liu, Dr. Zeyu Zhang, and Prof. Song-Chun Zhu from the Beijing Institute for General Artificial Intelligence (BIGAI).
“Stay curious. Stay humble. Keep building.”
🔥 News
- [2026/01] 🎉 ReCogDrive was accepted to ICLR 2026!
- [2025/05] 🎉 M3Bench was accepted to RA-L 2025!
- [2025/04] 🎉 DiffusionDrive was awarded as a CVPR 2025 Highlight!
- [2025/03] 🎉 M2Diffuser was accepted to T-PAMI 2025!
- [2025/02] 🎉 DiffusionDrive was accepted to CVPR 2025!
📝 Selected Publications
My research is broadly in Robotics and Computer Vision, with a particular focus on generalizable perception and dexterous manipulation in complex environments. For a complete publication list, please visit my Google Scholar profile.
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
Sixu Yan*, Shikang Wang*, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
* denotes equal contribution
Paper · arXiv · Project · Code · Simulation · YouTube · Bilibili · RedNote
We propose AdaRoboVLG, a task-adaptive vision-language-grasping framework that composes specialized foundation-model priors with a generalizable base policy, enabling physically feasible and context-aware grasp synthesis across different robotic hands.
BibTeX
@article{yan2026adarobovlg,
title = {Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis},
author = {Yan, Sixu and Wang, Shikang and Huang, Binhua and Tang, Xuanlai and Fan, Guohua and Huang, Fan and Li, Haoxuan and Li, Yongkang and Li, Yuhan and Liao, Bencheng and Zhang, Zeyu and Liu, Wenyu and Liu, Hangxin and Wang, Xinggang},
journal = {arXiv preprint arXiv:2609.04096},
year = {2026}
}
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
Yongkang Li, Lijun Zhou, Sixu Yan, Bencheng Liao, Tianyi Yan, Kaixin Xiong, Long Chen, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Haiyang Sun, Xinggang Wang
Paper · arXiv · Project · Code · Hugging Face · Dataset
We propose UniDriveVLA, a unified driving vision-language-action model that decouples understanding, spatial perception, and action planning through specialized experts, achieving strong performance across open-loop and closed-loop autonomous-driving benchmarks.
BibTeX
@article{li2026unidrivevla,
title = {{UniDriveVLA}: Unifying Understanding, Perception, and Action Planning for Autonomous Driving},
author = {Li, Yongkang and Zhou, Lijun and Yan, Sixu and Liao, Bencheng and Yan, Tianyi and Xiong, Kaixin and Chen, Long and Xie, Hongwei and Wang, Bing and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Sun, Haiyang and Wang, Xinggang},
journal = {arXiv preprint arXiv:2604.02190},
year = {2026}
}
OmniTrack: General Motion Tracking via Physics-Consistent Reference
Yuhan Li, Peiyuan Zhi, Yunshen Wang, Tengyu Liu, Sixu Yan, Wenyu Liu, Xinggang Wang, Baoxiong Jia, Siyuan Huang
Paper · arXiv · Project · Code · YouTube
We propose OmniTrack, a two-stage humanoid motion tracking framework that first generates physically feasible references in simulation and then learns a general policy to track them, improving accuracy and generalization to unseen motions.
BibTeX
@article{li2026omnitrack,
title = {OmniTrack: General motion tracking via physics-consistent reference},
author = {Li, Yuhan and Zhi, Peiyuan and Wang, Yunshen and Liu, Tengyu and Yan, Sixu and Liu, Wenyu and Wang, Xinggang and Jia, Baoxiong and Huang, Siyuan},
journal = {arXiv preprint arXiv:2602.23832},
year = {2026}
}
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
Yongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang
Paper · arXiv · Project · Code · Models · Dataset
We propose ReCogDrive, a reinforced cognitive framework that combines vision-language driving understanding with a diffusion planner and reinforcement learning to generate safer, more stable trajectories for end-to-end autonomous driving.
BibTeX
@inproceedings{li2026recogdrive,
title = {{ReCogDrive}: A Reinforced Cognitive Framework for End-to-End Autonomous Driving},
author = {Li, Yongkang and Xiong, Kaixin and Guo, Xiangyu and Li, Fang and Yan, Sixu and Xu, Gangwei and Zhou, Lijun and Chen, Long and Sun, Haiyang and Wang, Bing and Ma, Kun and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Wang, Xinggang},
booktitle = {International Conference on Learning Representations},
volume = {2026},
pages = {157518--157556},
year = {2026}
}
M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes
Zeyu Zhang*, Sixu Yan*, Muzhi Han, Zaijin Wang, Xinggang Wang, Song-Chun Zhu, Hangxin Liu
* denotes equal contribution
Paper · arXiv · Project · YouTube · Bilibili
We propose M3Bench, a large-scale benchmark and data generation toolkit for evaluating whole-body motion generation in mobile manipulation. It includes over 30,000 pick-and-place tasks across 119 realistic 3D scenes, with expert trajectories generated by a VKC planner.
BibTeX
@article{zhang2025m3bench,
title = {M${}^{3}$Bench: Benchmarking Whole-Body Motion Generation for Mobile Manipulation in 3D Scenes},
author = {Zhang, Zeyu and Yan, Sixu and Han, Muzhi and Wang, Zaijin and Wang, Xinggang and Zhu, Song-Chun and Liu, Hangxin},
journal = {IEEE Robotics and Automation Letters},
year = {2025},
volume = {10},
number = {7},
pages = {7286--7293},
publisher = {IEEE}
}
M2Diffuser: Diffusion-based Trajectory Optimization for Mobile Manipulation in 3D Scenes
Sixu Yan, Zeyu Zhang, Muzhi Han, Zaijin Wang, Qi Xie, Zhitian Li, Zhehan Li, Hangxin Liu, Xinggang Wang, Song-Chun Zhu
Paper · arXiv · Project · Code · YouTube · Bilibili
We propose M2Diffuser (Mobile Manipulation Diffuser), a conditional diffusion-based neural motion planner capable of generating full-body coordinated trajectories that satisfy both physical and task constraints.
BibTeX
@article{yan2025m2diffuser,
title = {M2Diffuser: Diffusion-based Trajectory Optimization for Mobile Manipulation in 3D Scenes},
author = {Yan, Sixu and Zhang, Zeyu and Han, Muzhi and Wang, Zaijin and Xie, Qi and Li, Zhitian and Li, Zhehan and Liu, Hangxin and Wang, Xinggang and Zhu, Song-Chun},
journal = {IEEE Transactions on Pattern Analysis and Machine Intelligence},
year = {2025},
publisher = {IEEE}
}
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, Xinggang Wang
Paper · arXiv · Project · Code · Hugging Face
We propose DiffusionDrive, a truncated diffusion-based planner for real-time end-to-end autonomous driving. By injecting multi-mode anchors and reducing denoising to two steps, it achieves fast inference while maintaining high-quality trajectory prediction.
BibTeX
@inproceedings{liao2025diffusiondrive,
author = {Liao, Bencheng and Chen, Shaoyu and Yin, Haoran and Jiang, Bo and Wang, Cheng and Yan, Sixu and Zhang, Xinbang and Li, Xiangyu and Zhang, Ying and Zhang, Qian and Wang, Xinggang},
title = {DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving},
booktitle = {Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
year = {2025},
pages = {12037--12047}
}
🎓 Education
-
Huazhong University of Science and Technology (HUST), 2024.09–Present
Ph.D. Student, HUST Vision Lab (HUSTVL)
Research advisor: Prof. Xinggang Wang -
Shanghai Jiao Tong University (SJTU), 2021.09–2024.06
M.S. Student, Robot Control and Machine Vision Lab (RCMVL)
Research advisor: Prof. Zhenhua Xiong; academic advisor: Prof. Han Ding -
Ocean University of China (OUC), 2017.09–2021.06
Undergraduate Student, GPA rank 1/62 during the application season
Research advisor: Prof. Xiaojie Tian
💼 Internships
- Beijing Institute for General Artificial Intelligence (BIGAI), 2023.07–2024.08
Research Intern, Robotics Lab
Research advisors: Dr. Hangxin Liu and Dr. Zeyu Zhang; academic advisor: Prof. Song-Chun Zhu
🤝 Academic Service
- Reviewer, IEEE Transactions on Automation Science and Engineering (T-ASE 2026)
- Reviewer, European Conference on Computer Vision (ECCV 2026)
- Reviewer, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026)
- Reviewer, Conference on Neural Information Processing Systems (NeurIPS 2025)
- Reviewer, IEEE Robotics and Automation Letters (RA-L 2024)
- Reviewer, IEEE International Conference on Robotics and Automation (ICRA 2024)
🏆 Selected Awards & Honors
- 2026 — Grand Prize, 14th Hubei Challenge Cup College Student Entrepreneurship Plan Competition, Huazhong University of Science and Technology
- 2025 — Best Poster Award, 4th Workshop on Mobile Manipulation and Embodied Intelligence (MOMA.v4), IROS 2025
- 2025 — Champion, Solo Dance, 2025 World Humanoid Robot Games
- 2025 — Best Paper Award (First Prize), 1st International Conference on General Artificial Intelligence
- 2025 — National Scholarship for Ph.D. Students, Huazhong University of Science and Technology
- 2025 — Outstanding Graduate Student, Huazhong University of Science and Technology
- 2023 — First-class Comprehensive Academic Scholarship, Shanghai Jiao Tong University
- 2022 — First-class Comprehensive Academic Scholarship, Shanghai Jiao Tong University
- 2021 — National Scholarship for Undergraduates, Ocean University of China
- 2021 — First-class Scholarship for Academic Excellence, Ocean University of China
- 2021 — Outstanding Individual of the 10th Role Model Program, College of Engineering, OUC
- 2021 — Outstanding Bachelor’s Thesis Award, Ocean University of China
- 2019 — First-class Scholarship for Academic Excellence, Ocean University of China
- 2019 — Scholarship for Social Practice, Ocean University of China
- 2019 — First Prize, 13th National College Student Energy Saving and Emission Reduction Competition
- 2019 — Honorable Mention (Second Prize), Mathematical Contest in Modeling (MCM)
- 2019 — Second Prize, 10th National Undergraduate Mathematics Competition
- 2019 — First Prize, 9th Shandong Undergraduate Mathematics Competition
- 2019 — First Prize, 2nd Shandong Undergraduate Physics Competition
- 2018 — First-class Scholarship for Academic Excellence, Ocean University of China
- 2018 — Scholarship for Technological Innovation, Ocean University of China
This website is based on the Academic Pages template.