Hi, I'm Ding Zou. I currently work as a Large Language Model Algorithm Engineer (Team Leader) at
ZTE Corporation for the Blue Sword Project, where I lead a
research team focusing on Agentic Training research (Algorithm & Infra). I hold a Master's degree in
Computer Science and Technology (2021–2024) and a Bachelor's degree in Electronic Engineering
(2017–2021), both from Huazhong University of Science and Technology (HUST).
My research spans multimodal large language models, reinforcement learning post-training, embodied
intelligence, and knowledge-aware recommendation.
Revisits multimodal RL post-training data sampling from a difficulty-distinguish perspective, with two complementary difficulty metrics (PISM & CMAB) and a hierarchical post-training framework.
A large-scale embodied planning model trained with multimodal post-training and Step-GRPO reinforcement learning, significantly outperforming RoboBrain2.0 on multimodal reasoning, spatial perception, and task planning benchmarks.
Introduces a curriculum reinforcement learning strategy from the reward-acquisition-difficulty perspective, enabling a Qwen2.5VL-3B model to outperform InternVL2.5-26B on multiple general benchmarks.
A knowledge-enhanced multi-intent transformer that integrates global heterogeneous information to model user intents while selectively filtering intent-irrelevant knowledge triples.
Disentangles latent user intents from multiple views with graph neural networks for bundle recommendation.
Experience
ZTE Corporation, BlueSword Program 2024.07 - Present
Algorithm Expert — Large Language Model Algorithm Engineer (Team Leader), Agentic Training research (Algorithm & Infra)