Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model

stp2yJanuary 16, 20250 Comments

AmazUtah_NLP at SemEval-2024 Task 9: A MultiChoice Question Answering System for Commonsense Defying Reasoning

[Submitted on 2 Jul 2024 (v1), last revised 15 Jan 2025 (this version, v2)]

View a PDF of the paper titled Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model, by Cong Cao and 3 other authors

View PDF
HTML (experimental)

Abstract:Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the first framework for zero-shot video restoration and enhancement based on the pre-trained image diffusion model. By replacing the spatial self-attention layer with the proposed short-long-range (SLR) temporal attention layer, the pre-trained image diffusion model can take advantage of the temporal correlation between frames. We further propose temporal consistency guidance, spatial-temporal noise sharing, and an early stopping sampling strategy to improve temporally consistent sampling. Our method is a plug-and-play module that can be inserted into any diffusion-based image restoration or enhancement methods to further improve their performance. Experimental results demonstrate the superiority of our proposed method. Our code is available at this https URL.

Submission history

From: Cong Cao [view email]
[v1]
Tue, 2 Jul 2024 05:31:59 UTC (11,366 KB)
[v2]
Wed, 15 Jan 2025 06:06:31 UTC (21,175 KB)

Source link
lol

By stp2y