OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution

Abstract

Denoising Diffusion Probabilistic Models (DDPM) and Flow Matching (FM)generative models show promising potential for one-step Real-World ImageSuper-Resolution (Real-ISR). Recent one-step Real-ISR models typically inject aLow-Quality (LQ) image latent distribution at the initial timestep. However, afundamental gap exists between the LQ image latent distribution and theGaussian noisy latent distribution, limiting the effective utilization ofgenerative priors. We observe that the noisy latent distribution at DDPM/FMmid-timesteps aligns more closely with the LQ image latent distribution. Basedon this insight, we present One Mid-timestep Guidance Real-ISR (OMGSR), auniversal framework applicable to DDPM/FM-based generative models. OMGSRinjects the LQ image latent distribution at a pre-computed mid-timestep,incorporating the proposed Latent Distribution Refinement loss to alleviate thelatent distribution gap. We also design the Overlap-Chunked LPIPS/GAN loss toeliminate checkerboard artifacts in image generation. Within this framework, weinstantiate OMGSR for DDPM/FM-based generative models with two variants:OMGSR-S (SD-Turbo) and OMGSR-F (FLUX.1-dev). Experimental results demonstratethat OMGSR-S/F achieves balanced/excellent performance across quantitative andqualitative metrics at 512-resolution. Notably, OMGSR-F establishesoverwhelming dominance in all reference metrics. We further train a1k-resolution OMGSR-F to match the default resolution of FLUX.1-dev, whichyields excellent results, especially in the details of the image generation. Wealso generate 2k-resolution images by the 1k-resolution OMGSR-F using ourtwo-stage Tiled VAE & Diffusion.

Quick Read (beta)

loading the full paper ...