Zhuofan Zong 宗卓凡 · Ph.D. Candidate at CUHK

I build generative models that understand, reason, and create across modalities. I am a final-year Ph.D. candidate at MMLab, CUHK, advised by Prof. Hongsheng Li and Prof. Xiaogang Wang, and currently a Research Intern at ByteDance Seed. My current research focuses on agentic visual generation and autoregressive video generation.

Portrait of Zhuofan Zong
Seeking 2027 Research Scientist opportunities. I expect to graduate in July 2027 and am looking for roles focused on visual generation and world models. Get in touch →

Research

01

Multimodal generation

Generative models for high-quality, controllable image and video synthesis, including diffusion and unified generative models.

02

Multimodal understanding

Multimodal large language models that perceive diverse visual inputs and reason over complex multimodal content.

03

Large vision foundation models

Scalable visual representation learning, model architectures, and efficient training for general-purpose vision systems.

Experience

Beijing, China

ByteDance Seed

Research Intern · Seed Vision · Mentor: Liyang Liu

Working on enhancing Seedream's agentic generation capabilities.

Beijing, China

SenseTime Research

Research Intern · Interaction Model R&D · Mentor: Mingjie Zhan

Designed visual output capabilities for interactive foundation models, with an emphasis on real-time, synchronized, and high-fidelity audio-video generation.

Beijing, China

SenseTime Research

Research Intern · Base Model R&D · Mentors: Guanglu Song & Yu Liu

Researched diffusion models, multimodal LLMs, and large vision foundation models. A core founding member of frontier R&D projects including a large vision foundation model, a multimodal interactive model, and SenseMirage.

Multimodal Models

2023 — 2026
Overview of the VoCa unified autoregressive audio-video generation framework

VoCa: Unified Autoregressive Modeling for Talking Audio-Video Generation

Zhuofan Zong, Jiale Yuan, Yufei Liu, Dongzhi Jiang, Hao Shao, Zimu Lu, Ke Wang, Yunqiao Yang, Mingjie Zhan, and Hongsheng Li

A unified autoregressive framework for synchronized talking audio-video generation.

ECCV
2026

Vision Foundation Models

2020 — 2023

* Equal contribution. Author lists are abbreviated only where shown.

Education & Service

Beihang University

MPhil in Computer Technology

Aug 2020 — Jan 2023 · Advisor: Prof. Biao Leng

Beihang University

BEng in Computer Science and Engineering

Aug 2016 — Jun 2020 · Advisors: Prof. Biao Leng & Prof. Xianglong Liu