Agentic intelligence · Generative systems

I am a Ph.D. researcher at HKUST(GZ), advised by Prof. Lei Zhu. My research connects agentic post-training, recursive self-improvement, and long-horizon planning and harness design. Before that, I received my M.S. in Artificial Intelligence from Tianjin University.

01

Research Interests

My research asks how agents can sustain open-ended work, learn from execution, and improve the systems around them:

  1. Long-horizon agent systems & harnesses

    Multi-agent coordination is the key to reliable long-horizon execution: specialist agents decompose and coordinate evolving workflows, while shared context and harnesses preserve continuity across extended tasks in JarvisHub and Seedance Studio.

  2. Agentic RL & environment feedback

    Trajectory rollouts, structured rewards, rubrics, verifiers, evaluator design, and preference optimization in JarvisVid and GenEvolve.

  3. Self-improving agents / RSI

    Auto research and recursive self-improvement (RSI) through the joint evolution of skills, trajectory data, evaluators, and harnesses. I apply this framework in SeedanceGame to iteratively improve its agent system.

02

Experience

Algorithm Intern

ByteDance

Building multi-agent systems and harnesses for long-horizon workflows, with persistent project context, traceable execution, and coordinated branching. Exploring recursive self-improvement (RSI) through trajectory data, evaluators, and iterative harness evolution.

Qingyun Program

Tencent

Built video–action data pipelines and action-conditioned world models, including progressive denoising and distillation for minute-scale autoregressive generation.

Algorithm Intern

Hedra AI

Developed AgentShot and MSBench, large-scale V2V lip-sync generation, and RL-based human-centric video captioning with FactorizedGRPO.

03

Selected Research

JarvisVid: Agentic Video Storytelling via Closed-Loop Trajectory Planning

Models multi-shot creation as a sequence of agent decisions across plot, storyboard, visual anchors, and motion. Grid rewards, CoT-SFT, online planning experience distillation, and DPO improve policy quality and cross-stage consistency.

Agentic RLTrajectory learningVideo generation
Technical Report · 2026Agent harness

JarvisHub: A Canvas-Native Agent Harness

Uses a visual canvas as the user workspace, shared project state, and external memory. Typed nodes, traceable actions, and a protocol bridge support inspectable and recoverable long-running work.

arXiv · 2026Agentic generation

GenEvolve: Self-Evolving Image Generation Agents

Unifies search, reference selection, skills, and prompt construction into tool-use trajectories, then distills structured visual experience from high-reward rollouts.

AgentShot: Towards Agentic Multi-Shot Video Generation and Benchmarking

Introduces MSBench and a five-stage, nine-agent framework for prompt understanding, storyboard planning, and consistent multi-shot generation using open-source video models.

Multi-agentBenchmarkVideo generation