Prior to this, I obtained my Bachelor's degree in Computer Science and Engineering, from Sichuan University in 2020. Then I received my Master degree in Computer Science and Engineering, at Sun Yat-Sen University in 2023, where I was advised by Prof. Wei-Shi Zheng.
My research focuses on generalizable robotic manipulation, particularly on investigating pre-training data and scalable model architectures for robot foundation models.
I am currently seeking a full-time position in industry. Please feel free to reach out if you know of a suitable opportunity.
A native video-action foundation model for generalizable robot control, featuring semantic visual-action tokenization, causal pre-training, a sparse MoE backbone, and asynchronous closed-loop inference.
A multi-chunk prediction framework for causal world modeling that improves training convergence and rollout accuracy while enabling 2x faster parallel inference.
A zero-shot long-horizon manipulation framework that mimics human long-range activities via demonstrations and achieves robust execution via generating visual future.
The first weakly supervised framework for developing efficient end-to-end action recognition models on long videos, which gives birth to a new weakly supervised pipeline for downstream long-video tasks.
A large-scale action video description dataset named ActionHub is proposed, which is the first, and the largest dataset that provides millions of video descriptions to describe thousands of human actions.
The first benchmark, named XOV-Action, for the cross-domain open-vocabulary action recognition task, and a simple yet effective method to address the scene bias for the task.