PhD Defence • Artificial Intelligence | Machine Learning • Understanding and Generating Visual Worlds with Foundation Models

Monday, October 19, 2026 1:00 pm - 4:00 pm EDT (GMT -04:00)

Please note: This PhD defence will take place online.

Weiming Ren, PhD candidate
David R. Cheriton School of Computer Science

Supervisors: Professors Wenhu Chen, Jimmy Lin

Foundation models have transformed how visual information is understood and generated, but they still typically operate on isolated fragments of visual experience, such as a short caption, a single image, or a few seconds of video. In contrast, people perceive and create within coherent visual worlds that comprise multiple modalities, evolve over time, and extend across long horizons. This thesis argues that closing this gap for foundation models requires not only larger parameter sizes, but also broader interfaces: models must be able to accept richer inputs and produce a wider range of outputs. It studies how these interfaces can be expanded along three axes: control, the conditioning signals that a generative model can faithfully follow; horizon, the spatial and temporal extent of the visual content that a model can process; and modality, the range of inputs and outputs that a single model can support.

The first part expands the control axis by allowing visual generation to be guided by multiple modalities rather than by text alone. I present ConsistI2V, an image-to-video diffusion model that adapts a pretrained text-to-image model into a video generator conditioned on a single keyframe. It introduces fine-grained spatiotemporal conditioning on the first frame, preserving its appearance, identity, and layout while allowing the generated content to move over time. It also introduces FrameInit, an inference-time noise initialization method that uses the low-frequency information of the first frame to guide the layout of the generated video. Building on this model, I present AnyV2V, a tuning-free framework that formulates a wide range of video editing tasks as first-frame-conditioned generation. This formulation enables diverse forms of video editing without task-specific training.

The second part expands the horizon axis by addressing architecture, training data, and evaluation as closely connected challenges. I present Vamba, a hybrid Mamba-Transformer model that avoids the quadratic cost of causal self-attention. In Vamba, text tokens access visual information through cross-attention, while video tokens are updated using linear-time Mamba-2 blocks. Without discarding or compressing video tokens, Vamba processes more than 1,024 frames on a single GPU and outperforms previous efficient video models on hour-long video benchmarks. To address the limited availability of long/high-resolution video training data, I present VISTA, a spatiotemporal augmentation framework that synthesizes long-duration and high-resolution video instruction data from existing video-caption pairs. Training with this synthetic dataset consistently improves multiple video models. To evaluate these advances more reliably, I present VideoEval-Pro, which demonstrates that widely used multiple-choice benchmarks for long video understanding are often inflated by answer-choice shortcuts and prior knowledge. It instead uses carefully filtered open-ended questions, on which the strongest proprietary models achieves only 40.8% accuracy.

The third part expands the modality axis by combining visual understanding and generation within a single model. I present Tuna, a native unified multimodal model based on a shared continuous visual representation. This representation is produced by cascading a VAE encoder with a representation encoder, creating a feature space that preserves both the low-level visual structure needed for generation and the semantic information needed for understanding. A single LLM decoder performs autoregressive text generation for understanding and flow matching denoising for generation. The experiments demonstrate that joint training on multimodal understanding and generation data improves both capabilities for Tuna, compared with training on either type of data alone. To further simplify the model architecture, I present Tuna-2, which removes pretrained vision encoders entirely and operates directly on pixel embeddings, produced by simple patch embedding layers. With a minimalist model design, Tuna-2 matches its encoder-based counterpart in generation while achieving better performance on fine-grained visual understanding, demonstrating the effectiveness and scalability of pixel-space unified models.

Together, these contributions broaden the interface between foundation models and the visual world along all three axes. They trace a progression from generating and editing visual content under increasingly rich forms of control, to understanding visual information over extended periods, and finally to supporting understanding and generation within a single model. These advances represent concrete steps toward foundation models that process visual experience as a coherent, temporally extended, and multimodal world rather than as a collection of isolated fragments.


Attend this PhD defence virtually on Zoom.