Wan2.2 TI2V 5B Diffusers
Overview
Highlights
- Native Diffusers integration for streamlined Python deployment
- Strong spatial consistency from source image to video
- Efficient 5B parameter scale for faster inference
- Permissive Apache-2.0 license for commercial flexibility
- Optimized for high-fidelity temporal motion and animation
Usage
# Install Hugging Face transformers
pip install transformers torch
# Load model with transformers
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("Wan-AI/Wan2.2-TI2V-5B-Diffusers")
tokenizer = AutoTokenizer.from_pretrained("Wan-AI/Wan2.2-TI2V-5B-Diffusers")
Hugging Face Download
We recommend downloading the model via the Hugging Face CLI or Hub SDK.
Guidance:Before downloading, install huggingface_hub with:
pip install -U huggingface_hub
CLI Download
Download the full repository
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B-Diffusers
Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B-Diffusers config.json --local-dir ./dir
See the official docs for more CLI options
SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('Wan-AI/Wan2.2-TI2V-5B-Diffusers')
Git Download
Make sure git-lfs is installed first
git lfs install
git clone https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers
To skip LFS large-file downloads, use:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers
Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.
PyTorch / Transformers Usage
Install Transformers
pip install -U transformers torch
Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('Wan-AI/Wan2.2-TI2V-5B-Diffusers')
tokenizer = AutoTokenizer.from_pretrained('Wan-AI/Wan2.2-TI2V-5B-Diffusers')
Model Download
We recommend downloading the model via the ModelScope CLI or SDK.
Guidance:Before downloading, install ModelScope with:
pip install modelscope
CLI Download
Download the full repository
modelscope download --model Wan-AI/Wan2.2-TI2V-5B-Diffusers
Download a single file to a local folder (e.g. README.md into ./dir)
modelscope download --model Wan-AI/Wan2.2-TI2V-5B-Diffusers README.md --local_dir ./dir
See the docs for more CLI options
SDK Download
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('Wan-AI/Wan2.2-TI2V-5B-Diffusers')
Git Download
Make sure git-lfs is installed first
git lfs install
git clone https://www.modelscope.cn/Wan-AI/Wan2.2-TI2V-5B-Diffusers.git
To skip LFS large-file downloads, use:
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/Wan-AI/Wan2.2-TI2V-5B-Diffusers.git
ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。
Notebook Quickstart
Install the ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html
Load the model and run inference
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks
p = pipeline('text-generation', 'Wan-AI/Wan2.2-TI2V-5B-Diffusers')
Full Documentation
---
license: apache-2.0
language:
- en
- zh
pipeline_tag: text-to-video
---
Wan2.2
<p align="center">
<img src="assets/logo.png" width="400"/>
<p>
<p align="center">
💜 <a href="https://wan.video"><b>Wan</b></a>    |    🖥️ <a href="https://github.com/Wan-Video/Wan2.2">GitHub</a>    |   🤗 <a href="https://huggingface.co/Wan-AI/">Hugging Face</a>   |   🤖 <a href="https://modelscope.cn/organization/Wan-AI">ModelScope</a>   |    📑 <a href="https://arxiv.org/abs/2503.20314">Technical Report</a>    |    📑 <a href="https://wan.video/welcome?spm=a2ty_o02.30011076.0.0.6c9ee41eCcluqg">Blog</a>    |   💬 <a href="https://gw.alicdn.com/imgextra/i2/O1CN01tqjWFi1ByuyehkTSB_!!6000000000015-0-tps-611-1279.jpg">WeChat Group</a>   |    📖 <a href="https://discord.gg/AKNgpMK4Yj">Discord</a>  
<br>
-----
Wan: Open and Advanced Large-Scale Video Generative Models <be>
We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:
- 👍 Effective MoE Architecture: Wan2.2 introduces a Mixture-of-Experts (MoE) architecture into video diffusion models. By separating the denoising process cross timesteps with specialized powerful expert models, this enlarges the overall model capacity while maintaining the same computational cost.
- 👍 Cinematic-level Aesthetics: Wan2.2 incorporates meticulously curated aesthetic data, complete with detailed labels for lighting, composition, contrast, color tone, and more. This allows for more precise and controllable cinematic style generation, facilitating the creation of videos with customizable aesthetic preferences.
- 👍 Complex Motion Generation: Compared to Wan2.1, Wan2.2 is trained on a significantly larger data, with +65.6% more images and +83.2% more videos. This expansion notably enhances the model's generalization across multiple dimensions such as motions, semantics, and aesthetics, achieving TOP performance among all open-sourced and closed-sourced models.
- 👍 Efficient High-Definition Hybrid TI2V: Wan2.2 open-sources a 5B model built with our advanced Wan2.2-VAE that achieves a compression ratio of 16×16×4. This model supports both text-to-video and image-to-video generation at 720P resolution with 24fps and can also run on consumer-grade graphics cards like 4090. It is one of the fastest 720P@24fps models currently available, capable of serving both the industrial and academic sectors simultaneously.
This repository contains our TI2V-5B model, built with the advanced Wan2.2-VAE that achieves a compression ratio of 16×16×4. This model supports both text-to-video and image-to-video generation at 720P resolution with 24fps and can runs on single consumer-grade GPU such as the 4090. It is one of the fastest 720P@24fps models available, meeting the needs of both industrial applications and academic research.
Video Demos
<div align="center">
<video width="80%" controls>
<source src="https://cloud.video.taobao.com/vod/4szTT1B0LqXvJzmuEURfGRA-nllnqN_G2AT0ZWkQXoQ.mp4" type="video/mp4">
Your browser does not support the video tag.
</video>
</div>
🔥 Latest News!!
- Jul 28, 2025: 👋 We've released the inference code and model weights of Wan2.2.
Community Works
If your research or project builds upon Wan2.1 or Wan2.2, we welcome you to share it with us so we can highlight it for the broader community.📑 Todo List
- Wan2.2 Text-to-Video
- Wan2.2 Image-to-Video
- Wan2.2 Text-Image-to-Video
Run Wan2.2
#### Installation
Clone the repo:
git clone https://github.com/Wan-Video/Wan2.2.git
cd Wan2.2Install dependencies:
# Ensure torch >= 2.4.0
pip install -r requirements.txt#### Model Download
| Models | Download Links | Description |
|--------------------|---------------------------------------------------------------------------------------------------------------------------------------------|-------------|
| T2V-A14B | 🤗 Huggingface 🤖 ModelScope | Text-to-Video MoE model, supports 480P & 720P |
| I2V-A14B | 🤗 Huggingface 🤖 ModelScope | Image-to-Video MoE model, supports 480P & 720P |
| TI2V-5B | 🤗 Huggingface 🤖 ModelScope | High-compression VAE, T2V+I2V, supports 720P |
> 💡Note:
> The TI2V-5B model supports 720P video generation at 24 FPS.
Download models using huggingface-cli: This repository supports the > This command can run on a GPU with at least 24GB VRAM (e.g, RTX 4090 GPU). > 💡If you are running on a GPU with at least 80GB VRAM, you can remove the > 💡Similar to Image-to-Video, the
`` sh
pip install "huggingface_hub[cli]"
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B --local-dir ./Wan2.2-TI2V-5BDownload models using modelscope-cli:
pip install modelscope
modelscope download Wan-AI/Wan2.2-TI2V-5B --local_dir ./Wan2.2-TI2V-5B#### Run Text-Image-to-Video Generation
Wan2.2-TI2V-5B Text-Image-to-Video model and can support video generation at 720P resolutions.
> 💡Unlike other tasks, the 720P resolution of the Text-Image-to-Video task is 1280*704 or 704*1280.
--offload_model True, --convert_model_dtype and --t5_cpu options to speed up execution.
> 💡If the image parameter is configured, it is an Image-to-Video generation; otherwise, it defaults to a Text-to-Video generation.
size` parameter represents the area of the generated video, with the aspect ratio following that of the original input image.