> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-zh/shi-jue-mo-xing/qwen-vl.md).

# Qwen2.5-VL 视觉语言模型

在 Clore.ai GPU 上运行 Qwen2.5-VL——领先的开源视觉语言模型，用于图像/视频/文档理解。

阿里巴巴的 Qwen2.5-VL（2024 年 12 月）是表现最佳的开源权重视觉语言模型（VLM）。它提供 3B、7B 和 72B 参数规模，能够理解图像、视频帧、PDF、图表以及复杂的视觉布局。7B 版本恰到好处——它在基准测试中超过许多更大的模型，同时在单张 24 GB GPU 上也能轻松运行。

在 [Clore.ai](https://clore.ai/) 你可以租用你所需的确切 GPU——从用于 7B 模型的 RTX 3090，到用于 72B 版本的多 GPU 配置——并在几分钟内开始分析视觉内容。

## 主要特性

* **多模态输入** ——在单一模型中处理图像、视频、PDF、截图、图表和示意图。
* **三种规模** ——3B（边缘/移动端）、7B（生产最佳平衡点）、72B（SOTA 级质量）。
* **动态分辨率** ——以图像的原生分辨率进行处理；不会强制缩放到 224×224。
* **视频理解** ——接受多帧视频输入，并具备时间推理能力。
* **文档 OCR** ——从扫描文档、收据和手写便签中提取文本。
* **多语言** ——在英语、中文及 20 多种其他语言上表现出色。
* **支持 Ollama** ——可在本地运行，使用 `ollama run qwen2.5vl:7b` 即可零代码部署。
* **Transformers 集成** — `Qwen2_5_VLForConditionalGeneration` 在 HuggingFace 中 `transformers`.

## 需求

{% hint style="warning" %}
**Clore.ai 市场上未列出多 GPU 的 80GB 级机型。** 目前列出的最大配置是 4× RTX PRO 6000 Blackwell（每张 96GB，共 380GB）以及 8–11× RTX 5090（每张 32GB）。A100 / H200 / B200 容量可按 [裸机](https://clore.ai/bare-metal) 需求提供。部署前请查看 [GPU 价格与可用性](/guides/guides_v2-zh/ru-men-zhi-nan/pricing.md) 。
{% endhint %}

| 组件     | 3B    | 7B       | 72B           |
| ------ | ----- | -------- | ------------- |
| GPU 显存 | 8 GB  | 16–24 GB | 80+ GB（多 GPU） |
| 系统内存   | 16 GB | 32 GB    | 128 GB        |
| 磁盘     | 10 GB | 20 GB    | 150 GB        |
| Python | 3.10+ | 3.10+    | 3.10+         |
| CUDA   | 12.8+ | 12.8+    | 12.8+         |

**Clore.ai GPU 推荐：** 对于 **7B 模型**， **RTX 4090** （24 GB，$0.14–0.42/小时）或 **RTX 3090** （24 GB，$0.07–0.21/小时）是理想选择。对于 **72B**，在市场中筛选 **A100 80 GB** 或多 GPU 配置。

## 快速开始

### 选项 A：Ollama（最简单）

```bash
# 安装 ollama
curl -fsSL https://ollama.ai/install.sh | sh

# 拉取并运行 7B 视觉模型
ollama run qwen2.5vl:7b
```

然后在 ollama 提示符中：

```
>>> 描述这张图片：/path/to/photo.jpg
```

### 选项 B：Python / Transformers

```bash
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install transformers accelerate qwen-vl-utils pillow
```

## 使用示例

### 使用 Transformers 进行图像理解

```python
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info

model_name = "Qwen/Qwen2.5-VL-7B-Instruct"

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_name)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://upload.wikimedia.org/wikipedia/commons/a/a7/Camponotus_flavomarginatus_ant.jpg"},
            {"type": "text", "text": "这是什么昆虫物种？描述其关键识别特征。"},
        ],
    }
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)

inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(model.device)

output_ids = model.generate(**inputs, max_new_tokens=512)
response = processor.batch_decode(
    output_ids[:, inputs.input_ids.shape[1]:],
    skip_special_tokens=True,
)[0]

打印(response)
```

### 视频分析

```python
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "Qwen/Qwen2.5-VL-7B-Instruct",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-7B-Instruct")

messages = [
    {
        "role": "user",
        "content": [
            {"type": "video", "video": "file:///workspace/clip.mp4", "max_pixels": 360 * 420, "fps": 1.0},
            {"type": "text", "text": "总结这个视频中发生了什么。按顺序列出关键事件。"},
        ],
    }
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)

inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(model.device)

output_ids = model.generate(**inputs, max_new_tokens=1024)
print(processor.batch_decode(output_ids[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])
```

### 文档 OCR 与提取

```python
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "file:///workspace/receipt.jpg"},
            {"type": "text", "text": "从这张收据中提取所有项目、数量和价格。以 JSON 形式返回。"},
        ],
    }
]

# 使用上面的相同模型/处理器设置进行处理
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, videos=video_inputs, padding=True, return_tensors="pt").to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=2048)
print(processor.batch_decode(output_ids[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])
```

### 用于批量处理的 Ollama API

```python
import ollama
import base64
from pathlib import Path

def analyze_image(image_path: str, question: str) -> str:
    """通过 Ollama API 将图像发送到 Qwen2.5-VL。"""
    image_data = base64.b64encode(Path(image_path).read_bytes()).decode()
    response = ollama.chat(
        model="qwen2.5vl:7b",
        messages=[{
            "role": "user",
            "content": question,
            "images": [image_data],
        }],
    )
    return response["message"]["content"]

# 批量处理一个图像文件夹
from pathlib import Path
for img in sorted(Path("./photos").glob("*.jpg"))】【：】【“】【t_a4cbbc64":"result = analyze_image(str(img), \"用一句话描述这张图片。\")","t_1e5b26a7":"print(f\"{img.name}: {result}\")","t_d9216bed":"Ollama 可用于快速部署","t_5c8a8105":"是能让你最快搭建起可用 VLM 的路径。交互式使用无需 Python 代码。","t_0211a171":"7B 是最佳平衡点","t_93c42ebc":"——7B Instruct 版本在 4-bit 量化下可容纳于 16 GB VRAM，并且能提供与大得多的模型相当的质量。","t_6e71e1b5":"动态分辨率很重要","t_e9100024":"——Qwen2.5-VL 以原生分辨率处理图像。对于大图像（>4K），请将最大宽度缩放到 1920px，以避免过高的 VRAM 占用。","t_86e82429":"视频 fps 设置","t_97e159b9":"——对于视频输入，将","t_b548b427":"fps=1.0","t_ee0b8e20":"设置为每秒采样 1 帧。更高的值会很快消耗 VRAM；对于大多数分析任务，1 fps 就足够了。","t_ccef17c3":"持久化存储","t_d3ca0922":"——将","t_edfe28ea":"HF_HOME=/workspace/hf_cache","t_a0523658":"；7B 模型约为 15 GB。对于 ollama，模型会保存到","t_88b7163e":"~/.ollama/models/","t_026c65bb":"结构化输出","t_fad108c1":"——Qwen2.5-VL 对 JSON 格式指令的遵循度很好。让它“以 JSON 返回”，大多数时候你都会得到可解析的输出。","t_35d4189f":"多图像比较","t_a45242c9":"——你可以在一条消息中传入多张图像用于比较任务（例如，“这两个产品中哪一个看起来更高端？”）。","t_45b77fb7":"tmux","t_0c4d43df":"——始终在","t_19eca478":"Clore.ai 租用实例上运行。","t_f12b0134":"修复","t_1a7527f3":"使用 7B","t_58920087":"在","t_54192663":"from_pretrained()","t_15d31f48":"使用","t_0845b274":"；或者使用 3B 版本","t_470ff15b":"ollama pull qwen2.5vl:7b","t_9d608b38":"——确保你使用了正确的标签","t_6d7a2e33":"视频处理缓慢","t_4a5cc645":"fps","t_ecb03ac6":"降到 0.5，并将","t_21dfeebc":"max_pixels","t_bd857aec":"设为","t_f565583c":"；帧数越少 = 推理越快","t_80b71283":"输出乱码或为空","t_2c65d168":"增大","t_3d63beee":"max_new_tokens","t_65c73406":"；默认值对于详细描述来说可能太低","t_baae7577":"ImportError：qwen_vl_utils","t_7ffb73d9":"pip install qwen-vl-utils","t_a6053858":"——这是","t_f7eba64c":"process_vision_info()","t_977b9586":"72B 模型无法容纳","t_e09ac483":"使用 2× A100 80 GB，配合","t_dc2f6d3b":"，或应用 AWQ 量化","t_532ed82e":"找不到图像路径","t_bcb1ea69":"对于消息中的本地文件，请使用","t_d89b065a":"file:///绝对路径","t_2a992b2f":"格式","t_ebfc0232":"用英语提示时输出中文","t_ad145933":"在你的提示中添加“仅用英语回答。”"}】<|endoftext|>}】>ാ} ]}]}`}】]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}>ৱ}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]】]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}}}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}]}
    result = analyze_image(str(img), "用一句话描述这张图片。")
    print(f"{img.name}: {result}")
```

## 给 Clore.ai 用户的建议

1. **Ollama 可用于快速部署** — `ollama run qwen2.5vl:7b` 是能让你最快搭建起可用 VLM 的路径。交互式使用无需 Python 代码。
2. **7B 是最佳平衡点** ——7B Instruct 版本在 4-bit 量化下可容纳于 16 GB VRAM，并且能提供与大得多的模型相当的质量。
3. **动态分辨率很重要** ——Qwen2.5-VL 以原生分辨率处理图像。对于大图像（>4K），请将最大宽度缩放到 1920px，以避免过高的 VRAM 占用。
4. **视频 fps 设置** ——对于视频输入，将 `fps=1.0` 设置为每秒采样 1 帧。更高的值会很快消耗 VRAM；对于大多数分析任务，1 fps 就足够了。
5. **持久化存储** ——将 `HF_HOME=/workspace/hf_cache`；7B 模型约为 15 GB。对于 ollama，模型会保存到 `~/.ollama/models/`.
6. **结构化输出** ——Qwen2.5-VL 对 JSON 格式指令的遵循度很好。让它“以 JSON 返回”，大多数时候你都会得到可解析的输出。
7. **多图像比较** ——你可以在一条消息中传入多张图像用于比较任务（例如，“这两个产品中哪一个看起来更高端？”）。
8. **tmux** ——始终在 `tmux` Clore.ai 租用实例上运行。

## 故障排查

| 问题                          | 修复                                                                        |
| --------------------------- | ------------------------------------------------------------------------- |
| `OutOfMemoryError` 使用 7B    | 使用 `load_in_4bit=True` 在 `from_pretrained()` 使用 `bitsandbytes`；或者使用 3B 版本 |
| 未找到 Ollama 模型               | `ollama pull qwen2.5vl:7b` ——确保你使用了正确的标签                                  |
| 视频处理缓慢                      | 减少 `fps` 降到 0.5，并将 `max_pixels` 设为 `256 * 256`；帧数越少 = 推理越快                |
| 输出乱码或为空                     | 增大 `max_new_tokens`；默认值对于详细描述来说可能太低                                       |
| `ImportError：qwen_vl_utils` | `pip install qwen-vl-utils` ——这是 `process_vision_info()`                  |
| 72B 模型无法容纳                  | 使用 2× A100 80 GB，配合 `device_map="auto"` ，或应用 AWQ 量化                       |
| 找不到图像路径                     | 对于消息中的本地文件，请使用 `file:///绝对路径` 格式                                          |
| 用英语提示时输出中文                  | 在你的提示中添加“仅用英语回答。”                                                         |


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-zh/shi-jue-mo-xing/qwen-vl.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
