I am a Ph.D. researcher at HKUST(GZ), advised by
Prof. Lei Zhu. My research connects
agentic post-training, recursive self-improvement, and
long-horizon planning and harness design. Before that, I received my M.S. in Artificial Intelligence from Tianjin University.
01
Research Interests
My research asks how agents can sustain open-ended work, learn from execution, and improve the systems around them:
Long-horizon agent systems & harnesses
Multi-agent coordination is the key to reliable long-horizon execution: specialist agents decompose and coordinate evolving workflows, while shared context and harnesses preserve continuity across extended tasks in JarvisHub and Seedance Studio.
Agentic RL & environment feedback
Trajectory rollouts, structured rewards, rubrics, verifiers, evaluator design, and preference optimization in JarvisVid and GenEvolve.
Self-improving agents / RSI
Auto research and recursive self-improvement (RSI) through the joint evolution of skills, trajectory data, evaluators, and harnesses. I apply this framework in SeedanceGame to iteratively improve its agent system.
02
Experience
Algorithm Intern
ByteDance
Building multi-agent systems and harnesses for long-horizon workflows, with persistent project context, traceable execution, and coordinated branching. Exploring recursive self-improvement (RSI) through trajectory data, evaluators, and iterative harness evolution.
Qingyun Program
Tencent
Built video–action data pipelines and action-conditioned world models, including progressive denoising and distillation for minute-scale autoregressive generation.
Algorithm Intern
Hedra AI
Developed AgentShot and MSBench, large-scale V2V lip-sync generation, and RL-based human-centric video captioning with FactorizedGRPO.
03
Selected Research
JarvisVid: Agentic Video Storytelling via Closed-Loop Trajectory Planning
Models multi-shot creation as a sequence of agent decisions across plot, storyboard, visual anchors, and motion. Grid rewards, CoT-SFT, online planning experience distillation, and DPO improve policy quality and cross-stage consistency.
Agentic RLTrajectory learningVideo generation
Technical Report · 2026Agent harness
JarvisHub: A Canvas-Native Agent Harness
Uses a visual canvas as the user workspace, shared project state, and external memory. Typed nodes, traceable actions, and a protocol bridge support inspectable and recoverable long-running work.
Unifies search, reference selection, skills, and prompt construction into tool-use trajectories, then distills structured visual experience from high-reward rollouts.
AgentShot: Towards Agentic Multi-Shot Video Generation and Benchmarking
Introduces MSBench and a five-stage, nine-agent framework for prompt understanding, storyboard planning, and consistent multi-shot generation using open-source video models.