First Workshop on Visual Storytelling with Unified Multimodal Understanding and Generation
In conjunction with the Winter Conference on Applications of Computer Vision (WACV) 2027
We are pleased to announce the First Workshop on Visual Storytelling with Unified Multimodal Understanding and Generation (UniMUG 2027), to be held in conjunction with WACV 2027.
The field of artificial intelligence is rapidly advancing toward foundation models capable of both understanding and generating video content. While significant progress has been made through scaling compute, data, and algorithmic innovations, the domains of video understanding and video generation have largely evolved in parallel despite operating on the same underlying data modality. Internally, generative models learn latent concepts to compose spatial and temporal elements, while understanding models learn similar latent concepts to reason and produce textual responses.
This workshop aims to bridge these two complementary domains by providing a platform for academic and industry researchers to explore how shared latent representations and world concepts can be cross-leveraged to advance both understanding and generation capabilities.
We seek to address several fundamental questions at the intersection of understanding and generation, for example:
We invite submissions on, but not limited to, the following topics.
A full-day program of keynotes, presentations, and discussions.
Distinguished researchers from academia and industry.
Instructions for preparing and submitting your paper.