PhD Defence • Artificial Intelligence | Machine Learning • Controllable Image and Video Generation with Foundation Models

Wednesday, October 21, 2026 2:30 pm - 5:30 pm EDT (GMT -04:00)

Please note: This PhD defence will take place online.

Cong Wei, PhD candidate
David R. Cheriton School of Computer Science

Supervisor: Professor Wenhu Chen

Diffusion and flow-based generative models have made substantial progress in synthesizing images and videos from text descriptions. However, traditional text-to-image and text-to-video models lack the controllability required for practical applications. In visual creation, users may wish to edit existing images or videos, provide visual references, and specify which details should remain unchanged. For video creation, users may additionally want characters’ performances to synchronize with speech or long videos to remain coherent over extended durations. These demands call for generation methods that offer control beyond traditional text conditioning. Yet accepting additional conditions does not guarantee that a model follows them, motivating feedback that can guide improvement. This thesis advances controllable image and video generation with foundation models through four connected challenges: flexible instruction-guided generation and editing, temporal control through speech and history, reward feedback for instruction adherence, and benchmarks that reveal limits beyond visual appearance.

Across image and video generation and editing, users need flexible interfaces to express their intent through multimodal instructions comprising text, images, and videos. Supporting these requests requires models to identify the intended task and interpret the roles of the supplied inputs. We develop generalist systems that support diverse tasks through these flexible interfaces. OmniEdit uses specialist supervision to bring multiple image-editing operations into a generalist model. AnyV2V composes an image editor with an image-to-video generator to extend image-editing capabilities to video without task-specific tuning. To connect instruction understanding with visual synthesis, we introduce UniVideo, which brings multimodal understanding, generation, and editing into a shared model. The shared model can interpret multimodal instructions and distinguish the intended task, supporting nine image and video tasks within a unified framework and demonstrating strong generalization to unseen tasks.

Compared with image generation, video generation introduces an additional temporal dimension of control. Beyond specifying content and edits, we study how speech and historical context can guide the unfolding video. We introduce MoCha for speech-conditioned video generation that aligns performance with audio and coordinates multi-character dialogue, enabling movie-grade storytelling. For long-video generation, maintaining consistency with earlier context becomes increasingly difficult as the video grows longer. We address this with Context Forcing, which combines a long-context teacher that guides the use of historical context with Slow-Fast Memory that balances recent detail with longer-term context. This supports real-time long-video generation while mitigating forgetting and drift.

Enabling models to accept additional control inputs does not ensure that their outputs follow the intended requirements. We study this gap in image generation and editing, using reward feedback to evaluate the requested transformation together with source preservation and visual quality. RationalRewards learns structured critiques from preference data and uses them to guide both reinforcement learning and inference-time refinement. With RewardHarness, we further show that adapting a reward system need not require updating its model weights: evaluation guidelines and tools can instead be learned from a small set of preference demonstrations. Across these approaches, downstream image-generation and editing experiments show that feedback can guide improvement rather than merely judge completed outputs.

When generation requests depend on world knowledge or visual reasoning, visual quality alone cannot establish correctness. We therefore design benchmarks that make these demands explicit and expose model limitations. SearchGen-Bench tests long-tailed and evolving knowledge requirements, revealing failures to depict requested entities and facts accurately. These failures motivate selective retrieval and co-training of a search agent and generator, so that external information can help satisfy requests beyond the generator’s existing knowledge. VGI-Bench evaluates visual reasoning in video generation, examining whether models can solve visually grounded problems given an initial image, a goal, and task constraints. The results reveal emerging visual reasoning abilities, but reliable performance under these constraints remains challenging. These benchmarks make knowledge accuracy and visual reasoning explicit criteria for evaluating whether generated content satisfies the request.

Together, these studies advance visual generation from synthesizing visually plausible outputs toward understanding and fulfilling users’ creative intent by broadening how that intent can be expressed, improving how models respond to it, and evaluating success beyond visual appearance.


Attend this PhD defence virtually on Zoom.