A comprehensive tutorial on the architecture design, representation learning, training dynamics, and evaluation of unified multimodal models that integrate understanding and generation within a single framework.
A systematic taxonomy of UMM architectures — External Expert Integration, Modular Joint Modeling, and End-to-End Unified Modeling — with trade-off analysis between autoregressive, diffusion, and hybrid approaches.
The "Unified Tokenizer" debate: continuous representations (e.g., CLIP) vs. discrete tokens (e.g., VQ-VAE), and hybrid encoding strategies balancing semantic understanding with generative fidelity.
The full training lifecycle — from constructing interleaved image-text data to unified pre-training objectives and advanced post-training alignment methods such as DPO and GRPO.
From isolated multimodal understanding or generation systems to unified multimodal foundation models capable of handling both tasks simultaneously.
A taxonomy of architectures including External Expert Integration, Modular Joint Modeling, and End-to-End Unified Modeling, with comparisons between autoregressive, diffusion, and hybrid approaches.
Continuous versus discrete representations, their advantages and limitations, and emerging hybrid encoding strategies that balance semantic understanding and generative fidelity.
Construction of modality-interleaved datasets, unified pre-training objectives, and post-training alignment methods such as DPO and GRPO.
Evaluation protocols, real-world applications in robotics and autonomous driving, and future directions such as scalable unified tokenizers and unified world models.
Download the tutorial slides for the CVPR 2026 unified multimodal models tutorial.
Download Slides →An annotated compilation of all references discussed in the tutorial as a comprehensive reading list.
Coming SoonOpen-source unified multimodal codebase with annotated pointers to models (e.g., Emu, Janus) and datasets.
Coming Soon@misc{wang2026roadconvergence,
title = {The Road to Convergence: Evolution of Unified Multimodal Models},
author = {Wang, Jindong and Chen, Hao and Hu, Jiakui and Su, Zhaolong and Li, Sharon},
year = {2026},
howpublished = {CVPR 2026 Tutorial},
url = {https://umm-tutorial.github.io}
}