WACV 2027 Workshop

UniMUG 2027

First Workshop on Visual Storytelling with Unified Multimodal Understanding and Generation

In conjunction with the Winter Conference on Applications of Computer Vision (WACV) 2027

Submission
TBD
Notification
TBD
Camera-Ready
TBD
Workshop
Full Day

Overview

We are pleased to announce the First Workshop on Visual Storytelling with Unified Multimodal Understanding and Generation (UniMUG 2027), to be held in conjunction with WACV 2027.

The field of artificial intelligence is rapidly advancing toward foundation models capable of both understanding and generating video content. While significant progress has been made through scaling compute, data, and algorithmic innovations, the domains of video understanding and video generation have largely evolved in parallel despite operating on the same underlying data modality. Internally, generative models learn latent concepts to compose spatial and temporal elements, while understanding models learn similar latent concepts to reason and produce textual responses.

This workshop aims to bridge these two complementary domains by providing a platform for academic and industry researchers to explore how shared latent representations and world concepts can be cross-leveraged to advance both understanding and generation capabilities.

We seek to address several fundamental questions at the intersection of understanding and generation, for example:

How can internal latent concepts be enriched in richness and reasoning abilities?
How can they be cross-leveraged to improve both understanding and generative outcomes?
How can we leverage both to conduct rigorous evaluation of both reciprocally?

Important Dates

Paper Submission
TBD
Notification
TBD
Camera-Ready
TBD
Workshop
Full Day

Topics of Interest

We invite submissions on, but not limited to, the following topics.

Joint Understanding & Generation

  • Novel architectures for unified multimodal understanding and generation
  • Shared latent representations across understanding and generation tasks
  • Multi-task learning frameworks bridging perception and synthesis
  • Methods to train unified models for joint understanding and generation

Generation-Enhanced Understanding

  • Leveraging generative priors for improved video comprehension
  • World models for video reasoning and prediction
  • Generative pre-training for downstream understanding tasks

Understanding-Enhanced Generation

  • Incorporating semantic understanding signals into video generation
  • Controllable generation through scene understanding
  • Consistency enforcement via understanding modules

Efficiency & Scalability

  • Teacher-student knowledge distillation from expert models
  • Efficient unified architectures for resource-constrained deployment
  • Model compression techniques for joint understanding-generation models

Evaluation & Benchmarks

  • Novel metrics for unified model assessment
  • Benchmark datasets for joint understanding and generation
  • Human evaluation methodologies

Applications

  • Long-form video storytelling and narrative generation
  • Interactive video content creation
  • Video summarization and recap generation
  • Multi-modal generation and understanding
  • Cross-modal alignment

Workshop Schedule

A full-day program of keynotes, presentations, and discussions.

Time
Session
08:50 - 09:00
Opening Remarks
09:00 - 09:40
Keynote Talk 1
09:40 - 10:20
Keynote Talk 2
10:20 - 10:40
Coffee Break
10:40 - 11:20
Keynote Talk 3
11:20 - 12:20
Panel Discussion
12:20 - 13:30
Lunch Break
13:30 - 14:10
Keynote Talk 4
14:10 - 14:50
Keynote Talk 5
14:50 - 15:00
Coffee Break
15:00 - 16:00
Lightning Talks
16:00 - 16:30
Keynote Talk 6
16:30 - 17:30
Poster Session
17:30 - 17:45
Awards & Closing Remarks

Invited Speakers

Distinguished researchers from academia and industry.

HZ
HKU
Tentative
SX
NYU & Google
Tentative
KG
University of Texas at Austin
Confirmed
AK
Wayve
Tentative
AL
TwelveLabs
Tentative
SH
Amazon
Confirmed

Organizers

General Chairs
GM
Amazon
Confirmed
MA
Amazon
Confirmed
Program Chairs
RZ
Amazon
AG
Amazon
BS
Amazon
MP
Adobe
FX
Felix Juefei Xu
Google DeepMind
YY
Arizona State University, Amazon

Submission Guidelines

Instructions for preparing and submitting your paper.

Format WACV conference style, maximum 14 pages (excluding references)
Review Process Double-blind review
Supplementary Materials Optional additional materials permitted
Dual Submission Policy To be specified based on WACV 2027 guidelines
Submit via OpenReview

Sponsors