> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-zh/dui-bi/video-gen-comparison.md).

# 视频生成对比

比较可在 Clore.ai GPU 服务器上部署的领先开源视频生成模型。

{% hint style="info" %}
**AI 视频生成** 在 2024-2025 年爆发式增长。本指南比较了顶级开源模型——Hunyuan Video、Wan2.1、CogVideoX、Mochi 1 和 LTX-Video——涵盖质量、速度、VRAM 需求和使用场景。
{% endhint %}

***

## 快速决策矩阵

|               | Hunyuan Video | Wan2.1     | CogVideoX  | Mochi 1    | LTX-Video  |
| ------------- | ------------- | ---------- | ---------- | ---------- | ---------- |
| **开发者**       | 腾讯            | 阿里巴巴       | 智谱 AI      | Genmo      | LightRicks |
| **质量**        | ⭐⭐⭐⭐⭐         | ⭐⭐⭐⭐⭐      | ⭐⭐⭐⭐       | ⭐⭐⭐⭐       | ⭐⭐⭐        |
| **速度**        | 慢             | 中等         | 中等         | 中等         | **快**      |
| **最低显存**      | 24GB          | 16GB       | 16GB       | 24GB       | **8GB**    |
| **最大分辨率**     | 1280×720      | 1280×720   | 1440×960   | 848×480    | 1216×704   |
| **最长时长**      | 5s            | 5s         | 6秒         | 5.4秒       | 2分钟        |
| **许可证**       | CLA           | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| **GitHub 星标** | 1万以上          | 7000以上     | 6000以上     | 4000以上     | 5000以上     |

***

## 概览

### Hunyuan Video

截至 2025 年初，腾讯的 Hunyuan Video 被广泛认为是最好的开源视频生成模型。它采用基于 Transformer 的架构，运动质量极佳。

**关键规格**：130 亿参数，720p 下 5 秒，需要 24GB+ VRAM

### Wan2.1

阿里巴巴的 Wan（文影）2.1 是 Hunyuan 的强劲竞争对手，在更低的最低 VRAM 需求下提供相近的质量。提供 13 亿和 140 亿参数版本。

**关键规格**：13 亿（轻量版）或 140 亿，720p 下 5 秒，13 亿版本需要 16GB+ VRAM

### CogVideoX

智谱 AI 的 CogVideoX 侧重于精确跟随文本和连贯的长视频。它尤其擅长电影感内容和叙事型生成。

**关键规格**：50 亿/100 亿参数，1440×960 下 6 秒，需要 16GB+ VRAM

### Mochi 1

Genmo 的 Mochi 1 以平滑、流畅的运动和逼真的物理效果而闻名。它使用新颖的 AsymmDiT 架构。完全开源（权重 + 训练代码）。

**关键规格**：100 亿参数，848×480 下 5.4 秒，需要 24GB VRAM

### LTX-Video

LightRicks 的 LTX-Video 将推理速度放在首位。它可以在现代 GPU 上实现实时或近实时的视频生成——非常适合交互式应用。

**关键规格**：20 亿参数，最长可达 2 分钟的视频，需要 8GB VRAM

***

## 质量对比

### EvalCrafter 基准（2025）

{% hint style="info" %}
质量具有主观性。这些分数反映了来自 VBench 和 EvalCrafter 基准的社区共识。
{% endhint %}

| 模型            | VBench 分数 | 运动质量   | 文本对齐    | 美学  |
| ------------- | --------- | ------ | ------- | --- |
| Hunyuan Video | **83.2**  | **优秀** | 优秀      | 优秀  |
| Wan2.1（14B）   | **82.8**  | 优秀     | 优秀      | 优秀  |
| CogVideoX-5B  | 79.6      | 好      | **非常好** | 好   |
| Mochi 1       | 77.4      | 非常好    | 好       | 好   |
| LTX-Video     | 71.2      | 好      | 好       | 可接受 |

### 定性优势

| 模型            | 最擅长                | 弱点           |
| ------------- | ------------------ | ------------ |
| Hunyuan Video | 整体质量、电影感           | 非常慢，显存消耗大    |
| Wan2.1        | 质量/效率平衡，图像转视频（I2V） | 偶尔过饱和        |
| CogVideoX     | 长篇叙事、文本准确性         | 运动不够动态       |
| Mochi 1       | 流畅运动、物理效果          | 分辨率上限较低      |
| LTX-Video     | 速度、长视频             | 与其他模型相比质量有差距 |

***

## 速度基准

### 生成时间（A100 80GB，单 GPU）

| 模型            | 480p 5秒 | 720p 5秒  | 1080p 5秒 |
| ------------- | ------- | -------- | -------- |
| Hunyuan Video | 45分钟    | 约 3 小时   | ❌ 显存溢出   |
| Wan2.1（14B）   | 15分钟    | 45分钟     | ❌ 显存溢出   |
| Wan2.1（1.3B）  | 3 分钟    | 8分钟      | ❌ 显存溢出   |
| CogVideoX-5B  | 10分钟    | 25分钟     | ❌ 显存溢出   |
| Mochi 1       | 8分钟     | ❌ 显存溢出   | ❌ 显存溢出   |
| LTX-Video     | **45秒** | **3 分钟** | 8分钟      |

{% hint style="warning" %}
**时间为近似值** 并会随采样步数（20-50）、引导尺度和硬件而变化。预览时使用更少步数。
{% endhint %}

### 采用优化（TeaCache / FORA / 步数蒸馏）

优化推理可以显著缩短生成时间：

| 模型            | 使用缓存          | 加速倍数 |
| ------------- | ------------- | ---- |
| Hunyuan Video | 约 15 分钟（720p） | 4×   |
| Wan2.1        | 约 12 分钟（720p） | 约 4× |
| CogVideoX     | 约 8 分钟（720p）  | 约 3× |
| LTX-Video     | 约 45 秒（720p）  | 4×   |

***

## 显存要求

### 按模型和分辨率划分的最低 VRAM

| 模型            | 480p    | 720p  | 1080p |
| ------------- | ------- | ----- | ----- |
| Hunyuan Video | 24GB    | 40GB+ | ❌     |
| Wan2.1（14B）   | 24GB    | 40GB+ | ❌     |
| Wan2.1（1.3B）  | **8GB** | 16GB  | 24GB  |
| CogVideoX-5B  | 16GB    | 24GB  | ❌     |
| CogVideoX-2B  | **8GB** | 16GB  | ❌     |
| Mochi 1       | 24GB    | ❌     | ❌     |
| LTX-Video     | **8GB** | 12GB  | 24GB  |

### 显存优化技巧

#### 量化

```python
# 采用 8 位量化的 CogVideoX（显存减半）
from diffusers import CogVideoXPipeline
import torch

pipe = CogVideoXPipeline.from_pretrained(
    "THUDM/CogVideoX-5b"，
    torch_dtype=torch.float16
)
pipe.enable_model_cpu_offload()  # 进一步减少显存
pipe.vae.enable_slicing()
pipe.vae.enable_tiling()
```

#### CPU 卸载

```python
# 使用 CPU 卸载的 Wan2.1 以降低显存需求
from diffusers import WanPipeline

pipe = WanPipeline.from_pretrained(
    "Wan-AI/Wan2.1-T2V-1.3B-Diffusers"，
    torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
```

***

## Hunyuan Video：深度解析

### 架构

* **130 亿 DiT** （扩散 Transformer）参数
* 对所有空间和时间 token 进行全注意力
* 在 10 亿以上视频片段上训练

### 在 Clore.ai 上部署

```bash
# 克隆并安装
git clone https://github.com/Tencent/HunyuanVideo
cd HunyuanVideo
pip install -r requirements.txt

# 下载权重（约 87GB）
huggingface-cli download tencent/HunyuanVideo --local-dir ./weights

# 生成
python sample_video.py \
  --video-size 720 1280 \
  --video-length 129 \
  --infer-steps 50 \
  --prompt "一只雄伟的雄鹰在白雪覆盖的群山上空翱翔" \
  --flow-shift 7.0 \
  --embedded-cfg-scale 6.0 \
  --save-path ./outputs
```

### 通过 ComfyUI

```bash
# 为 ComfyUI 安装 HunyuanVideo 节点
cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-HunyuanVideoWrapper
pip install -r ComfyUI-HunyuanVideoWrapper/requirements.txt
```

**最适合**：最高质量的电影感视频生成，不受 VRAM 约束

***

## Wan2.1：深度解析

### 架构

* **两个版本**：Wan2.1-T2V-1.3B 和 Wan2.1-T2V-14B
* **图像转视频** （I2V）模型也可用
* 强大的多语言（中文 + 英文）提示词支持

### 在 Clore.ai 上部署

```python
from diffusers import WanPipeline
from diffusers.utils import export_to_video
import torch

# 1.3B 模型——可放入 8-16GB VRAM
pipe = WanPipeline.from_pretrained(
    "Wan-AI/Wan2.1-T2V-1.3B-Diffusers"，
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

output = pipe(
    prompt="一个宁静的日本庭院，樱花飘落",
    negative_prompt="低质量，模糊",
    height=480,
    width=832,
    num_frames=81,
    num_inference_steps=50,
    guidance_scale=5.0,
).frames[0]

export_to_video(output, "wan_video.mp4", fps=16)
```

### 使用 Wan2.1 进行图像转视频

```python
from diffusers import WanImageToVideoPipeline
from PIL import Image

pipe = WanImageToVideoPipeline.from_pretrained(
    "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers"，
    torch_dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()

image = Image.open("input.jpg")
output = pipe(
    image=image,
    prompt="这个人自信地向前走",
    num_frames=81,
).frames[0]
```

**最适合**：质量与效率平衡、I2V、多语言

***

## CogVideoX：深度解析

### 架构

* **专家级 Transformer** 配备 3D 全注意力
* **50 亿和 100 亿** 参数版本
* 用于提升视觉质量的 CogView3 图像编码器

### 在 Clore.ai 上部署

```python
from diffusers import CogVideoXPipeline
from diffusers.utils import export_to_video
import torch

pipe = CogVideoXPipeline.from_pretrained(
    "THUDM/CogVideoX-5b"，
    torch_dtype=torch.bfloat16
)
pipe.to("cuda")
pipe.vae.enable_slicing()
pipe.vae.enable_tiling()

video = pipe(
    prompt="城市夜景的延时摄影，车灯轨迹清晰可见",
    num_videos_per_prompt=1,
    num_inference_steps=50,
    num_frames=49,
    guidance_scale=6,
    generator=torch.Generator(device="cuda").manual_seed(42),
).frames[0]

export_to_video(video, "cogvideo.mp4", fps=8)
```

**最适合**：精确的文本转视频、叙事内容、长篇生成

***

## Mochi 1：深度解析

### 架构

* **AsymmDiT** ——非对称扩散 Transformer
* 专注于时间一致性和流畅运动
* 完全开源，包括训练代码

### 在 Clore.ai 上部署

```bash
pip install mochi-preview

python -c "
from mochi_preview.pipelines import DecoderModelFactory, DitModelFactory, MochiSingleGPUPipeline, T5ModelFactory
import tempfile
from pathlib import Path

pipeline = MochiSingleGPUPipeline(
    text_encoder_factory=T5ModelFactory(),
    dit_factory=DitModelFactory(model_path='./weights/mochi-dit.safetensors'),
    decoder_factory=DecoderModelFactory(model_path='./weights/mochi-vae.safetensors'),
    cpu_offload=True,
    decode_type='tiled_full',
)

video = pipeline(
    height=480, width=848,
    num_frames=163,
    num_inference_steps=64,
    sigma_schedule_type='linear_quadratic',
    cfg_schedule_type='linear',
    conditioning_args={'prompt': '一只海豚在日落时跃过海浪'},
)
"
```

**最适合**：流畅运动、逼真物理效果、研究用途

***

## LTX-Video：深度解析

### 架构

* **20 亿参数** DiT——更小、更快
* 原生 **长视频** 支持（最长 2 分钟）
* 专为实时或近实时生成而设计

### 在 Clore.ai 上部署

```python
from diffusers import LTXPipeline
from diffusers.utils import export_to_video
import torch

pipe = LTXPipeline.from_pretrained(
    "Lightricks/LTX-Video"，
    torch_dtype=torch.bfloat16
)
pipe.to("cuda")

video = pipe(
    prompt="夏日花园里，一只蝴蝶停在花朵上",
    negative_prompt="最差质量、运动不连贯、模糊",
    width=704,
    height=480,
    num_frames=161,
    decode_timestep=0.03,
    decode_noise_scale=0.025,
    num_inference_steps=50,
).frames[0]

export_to_video(video, "ltx_video.mp4", fps=24)
```

**最适合**：快速生成、交互式应用、长视频、有限 VRAM（8GB）

***

## 功能对比

### 能力概览

| 功能         | Hunyuan | Wan2.1 | CogVideoX | Mochi | LTX |
| ---------- | ------- | ------ | --------- | ----- | --- |
| 文本转视频      | ✅       | ✅      | ✅         | ✅     | ✅   |
| 图像转视频      | ✅       | ✅      | ✅         | ❌     | ✅   |
| 视频转视频      | ❌       | ❌      | ✅         | ❌     | ✅   |
| ControlNet | 部分      | ❌      | ✅         | ❌     | ❌   |
| LoRA 支持    | ✅       | ✅      | ✅         | ❌     | ✅   |
| ComfyUI 节点 | ✅       | ✅      | ✅         | ✅     | ✅   |
| 长视频（>10秒）  | ❌       | ❌      | 部分        | ❌     | ✅   |
| 中文提示词      | ✅       | ✅      | ✅         | ❌     | ❌   |

***

## Clore.ai GPU 推荐

### 针对每个模型

| 模型            | 最低 GPU         | 推荐          | 理想         |
| ------------- | -------------- | ----------- | ---------- |
| Hunyuan Video | RTX 3090（24GB） | A6000（48GB） | A100（80GB） |
| Wan2.1 14B    | RTX 3090（24GB） | A6000（48GB） | A100（80GB） |
| Wan2.1 1.3B   | RTX 3080（10GB） | RTX 3090    | RTX 4090   |
| CogVideoX-5B  | RTX 3090（24GB） | A6000（48GB） | A100       |
| CogVideoX-2B  | RTX 3080（10GB） | RTX 3090    | RTX 4090   |
| Mochi 1       | RTX 3090（24GB） | A6000（48GB） | A100       |
| LTX-Video     | RTX 3080（10GB） | RTX 4080    | RTX 4090   |

### 每个视频的成本估算

```
在 A100 80GB 上运行的 Hunyuan Video（720p，5秒）（[裸金属](https://clore.ai/bare-metal)）：
  时间：约 45 分钟 → 成本：每个视频约 $1.12

在 RTX 3090 上运行的 Wan2.1-1.3B（480p，5秒）（$0.07–0.21/小时）：
  时间：约 3 分钟 → 成本：每个视频约 $0.025

在 RTX 4090 上运行的 LTX-Video（720p，5秒）（$0.14–0.42/小时）：
  时间：约 3 分钟 → 成本：每个视频约 $0.030
```

***

## 何时使用哪一个

### 决策指南

```
追求最高质量（不考虑成本）？
  → A100 上的 Hunyuan Video

追求最佳质量/成本平衡？
  → A6000 上的 Wan2.1 14B

VRAM 有限（8-12GB）？
  → LTX-Video 或 Wan2.1 1.3B

需要快速生成？
  → LTX-Video

需要图像转视频？
  → Wan2.1 I2V 或 CogVideoX

需要长视频（>10秒）？
  → LTX-Video

研究/微调？
  → Mochi 1（开放训练代码）或 CogVideoX

需要 ComfyUI 工作流？
  → 都支持，Hunyuan/Wan 的节点最好
```

***

## 有用链接

* [Hunyuan Video GitHub](https://github.com/Tencent/HunyuanVideo)
* [HuggingFace 上的 Wan2.1](https://huggingface.co/Wan-AI)
* [CogVideoX GitHub](https://github.com/THUDM/CogVideo)
* [Mochi 1 GitHub](https://github.com/genmoai/mochi)
* [LTX-Video GitHub](https://github.com/Lightricks/LTX-Video)
* [视频生成排行榜](https://huggingface.co/spaces/ArtificialAnalysis/video-generation-arena-leaderboard)

***

## 总结

| 模型                | 适用场景                |
| ----------------- | ------------------- |
| **Hunyuan Video** | 最看重最佳质量，且有 A100+ 可用 |
| **Wan2.1**        | 质量与效率的最佳平衡          |
| **CogVideoX**     | 精确的文本转视频、长篇叙事       |
| **Mochi 1**       | 流畅运动、物理效果、开放研究      |
| **LTX-Video**     | 速度、低 VRAM、长视频       |

开源视频生成生态发展非常快。对于大多数 Clore.ai 部署， **Wan2.1** （预算选 1.3B，追求质量选 14B）在质量、速度和资源效率之间提供了最佳组合。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-zh/dui-bi/video-gen-comparison.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
