UltraGen: High-Resolution Video Generation with Hierarchical Attention

  • 2025-10-21 16:23:21
  • Teng Hu, Jiangning Zhang, Zihan Su, Ran Yi
  • 0

Abstract

Recent advances in video generation have made it possible to produce visuallycompelling videos, with wide-ranging applications in content creation,entertainment, and virtual reality. However, most existing diffusiontransformer based video generation models are limited to low-resolution outputs(<=720P) due to the quadratic computational complexity of the attentionmechanism with respect to the output width and height. This computationalbottleneck makes native high-resolution video generation (1080P/2K/4K)impractical for both training and inference. To address this challenge, wepresent UltraGen, a novel video generation framework that enables i) efficientand ii) end-to-end native high-resolution video synthesis. Specifically,UltraGen features a hierarchical dual-branch attention architecture based onglobal-local attention decomposition, which decouples full attention into alocal attention branch for high-fidelity regional content and a globalattention branch for overall semantic consistency. We further propose aspatially compressed global modeling strategy to efficiently learn globaldependencies, and a hierarchical cross-window local attention mechanism toreduce computational costs while enhancing information flow across differentlocal windows. Extensive experiments demonstrate that UltraGen can effectivelyscale pre-trained low-resolution video models to 1080P and even 4K resolutionfor the first time, outperforming existing state-of-the-art methods andsuper-resolution based two-stage pipelines in both qualitative and quantitativeevaluations.

 

Quick Read (beta)

loading the full paper ...