T3-Video: Transform Trained Transformer for Accelerating Native 4K Video Generation

1Zhejiang University   2Youtu Lab, Tencent   3National University of Singapore   4Peking University  

T3-Video achieves native UHD-4K video generation via efficient fine-tuning on just the 42K UltraVideo dataset.

4K Vision World demo generated by our T3-Video-Wan2.1-T2V-1.3B, where the prompt for each video is generated by GPT-4o and sorted by GDP.

High-lights for the T3 Module

1) Only modifies the inference logic of Full-Attention to achieve a global receptive field within a single layer.
2) Multi-scale window-attention realizes a linear transformation of the attention computation complexity from O(HW) to O(hwN), where hw denotes the window size and N denotes the number of windows.
3) Plug-and-play with elegant one-line code replacement, while being compatible with existing attention acceleration libraries such as FlashAttention, SageAttention, etc.

data-composition

T3-Video is much faster and better

data-composition

Pseudocode with one-line code replacement

data-composition

Compatible with multiple models

BibTeX


      @inproceedings{t3video,
        title={Transform Trained Transformer for Accelerating Native 4K Video Generation},
        author={Jiangning Zhang and Junwei Zhu and Teng Hu and Yabiao Wang and Donghao Luo and Weijian Cao and Zhenye Gan and Xiaobin Hu and Zhucun Xue and Xiangtai Li and Chengjie Wang and Yong Liu},
        booktitle={Forty-third International Conference on Machine Learning},
        year={2026},
        url={https://openreview.net/forum?id=5dabjiOBpS}
      }