Skip to content
PAPER N.3

FrescoDiffusion

2026
READ PAPER

Animating Frescoes

FrescoDiffusion turns a single ultra-high-resolution image into video at the same scale, preserving fine details that standard image-to-video models lose when complex scenes are resized.

Instead of generating one low-resolution clip and simply upscaling it, the method performs tiled diffusion directly on a large latent canvas.

4.0

A Global Prior for Local Detail

A low-resolution animation first captures the global structure and motion of the image. Its latent trajectory is then upscaled and used as a prior to guide every step of the 4K tiled denoising process.

This prior regularization keeps separate tiles coherent while still allowing the model to create new details and motion at full resolution.

4.1
image-3a460c36840a713263869efd0e9251ceddfeaa25-1000x604-webp
image-f734a03f51508ad4895e6c83c79bd4a4787d7a51-1000x1000-webp

Controlling What Moves

FrescoDiffusion can treat different regions independently, relaxing the prior earlier on active areas while keeping backgrounds constrained for longer. This creates motion without breaking the overall composition.

Tests on FrescoArchive and VBench-I2V show improved coherence, fidelity and controllability over tiled baselines, with only a small runtime overhead.

4.2
0
100
1