> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-hi/image-generation/hunyuan-image3.md).

# HunyuanImage 3.0

Clore.ai GPUs पर Tencent का 80B MoE मल्टिमोडल इमेज जनरेशन और एडिटिंग मॉडल HunyuanImage 3.0 चलाएँ

Tencent का HunyuanImage 3.0 है **दुनिया का सबसे बड़ा ओपन-सोर्स इमेज जनरेशन मॉडल** 80B कुल पैरामीटर (इन्फ़रेंस के दौरान 13B सक्रिय) के साथ। 26 जनवरी, 2026 को जारी किया गया, यह इमेज जनरेशन, संपादन और समझ को एक ही ऑटोरिग्रेसिव मॉडल में एकीकृत करके पारंपरिक ढांचे को तोड़ता है — text-to-image और image-to-image के लिए अब अलग-अलग पाइपलाइनों की जरूरत नहीं। यह फ़ोटो-यथार्थवादी छवियाँ बनाता है, तत्वों को सुरक्षित रखते हुए सटीक संपादन करता है, शैली परिवर्तन संभालता है, और यहाँ तक कि मल्टी-इमेज फ्यूज़न भी करता है, सब कुछ एक ही मॉडल से।

**HuggingFace:** [tencent/HunyuanImage-3.0-Instruct](https://huggingface.co/tencent/HunyuanImage-3.0-Instruct) **GitHub:** [Tencent-Hunyuan/HunyuanImage-3.0](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0) **लाइसेंस:** Tencent Hunyuan सामुदायिक लाइसेंस (100M MAU तक अनुसंधान और व्यावसायिक उपयोग के लिए मुफ़्त)

## मुख्य विशेषताएँ

* **80B कुल / 13B सक्रिय पैरामीटर** — सबसे बड़ा ओपन-सोर्स इमेज MoE मॉडल; प्रति इन्फ़रेंस केवल 13B पैरामीटर सक्रिय करता है
* **एकीकृत मल्टीमोडल आर्किटेक्चर** — एक ही मॉडल में टेक्स्ट-टू-इमेज, इमेज संपादन, शैली रूपांतरण, और मल्टी-इमेज संयोजन
* **निर्देश-आधारित संपादन** — प्राकृतिक भाषा में बताएं कि आप क्या बदलना चाहते हैं, और जो तत्व अपरिवर्तित हैं उन्हें सुरक्षित रखें
* **डिस्टिल्ड चेकपॉइंट उपलब्ध है** — `HunyuanImage-3.0-Instruct-Distil` तेज़ जनरेशन के लिए सिर्फ 8 सैंपलिंग स्टेप में चलता है
* **vLLM एक्सेलेरेशन** — प्रोडक्शन में बहुत तेज़ इन्फ़रेंस के लिए नेटिव vLLM सपोर्ट
* **ऑटोरिग्रेसिव फ्रेमवर्क** — DiT-आधारित मॉडलों (FLUX, SD3.5) के विपरीत, समझ और जनरेशन दोनों के लिए एकीकृत AR दृष्टिकोण का उपयोग करता है

## मॉडल वेरिएंट

| मॉडल                                 | उपयोग-प्रकरण                          | स्टेप्स | HuggingFace                                |
| ------------------------------------ | ------------------------------------- | ------- | ------------------------------------------ |
| **HunyuanImage-3.0**                 | केवल टेक्स्ट-टू-इमेज                  | 30–50   | `tencent/HunyuanImage-3.0`                 |
| **HunyuanImage-3.0-Instruct**        | टेक्स्ट-टू-इमेज + संपादन + मल्टी-इमेज | 30–50   | `tencent/HunyuanImage-3.0-Instruct`        |
| **HunyuanImage-3.0-Instruct-Distil** | तेज़ इन्फ़रेंस (8 स्टेप)              | 8       | `tencent/HunyuanImage-3.0-Instruct-Distil` |

## आवश्यकताएँ

{% hint style="warning" %}
**Clore.ai marketplace पर multi-GPU 80GB-class rigs सूचीबद्ध नहीं हैं।** आज सूचीबद्ध सबसे बड़े boxes 4× RTX PRO 6000 Blackwell (प्रत्येक 96GB, कुल 380GB) और 8–11× RTX 5090 (प्रत्येक 32GB) हैं। A100 / H200 / B200 क्षमता [bare metal](https://clore.ai/bare-metal) के रूप में अनुरोध पर बेची जाती है। देखें [GPU मूल्य और उपलब्धता](/guides/guides_v2-hi/getting-started/pricing.md) किसी deployment का आकार तय करने से पहले।
{% endhint %}

| कॉन्फ़िगरेशन | एकल GPU (ऑफ़लोडिंग)       | अनुशंसित     | बहु-GPU उत्पादन |
| ------------ | ------------------------- | ------------ | --------------- |
| GPU          | 1× RTX 4090 24GB          | 1× A100 80GB | 2–3× A100 80GB  |
| VRAM         | 24GB (लेयर ऑफ़लोड के साथ) | 80GB         | 160–240GB       |
| RAM          | 128GB                     | 128GB        | 256GB           |
| डिस्क        | 200GB                     | 200GB        | 200GB           |
| CUDA         | 12.8+                     | 12.8+        | 12.8+           |

**अनुशंसित Clore.ai सेटअप:**

* **मार्केटप्लेस पर सबसे अच्छा मूल्य:** 1× RTX PRO 6000 Blackwell 96GB ($0.92–1.38/घंटा) — ऑफ़लोडिंग के बिना पूरे मॉडल को आराम से चलाता है। A100 80GB क्लासिक विकल्प है, लेकिन यह केवल [bare metal](https://clore.ai/bare-metal)
* **बजट विकल्प:** 1× RTX 4090 ($0.14–0.42/घंटा) — CPU ऑफ़लोडिंग के साथ काम करता है (धीमा, लेकिन उपयोगी)
* **तेज़ उत्पादन:** 2× A100 80GB ([bare metal](https://clore.ai/bare-metal)) — बैच जनरेशन और Instruct मॉडल के लिए

## त्वरित शुरुआत

### इंस्टॉलेशन

```bash
# रिपॉज़िटरी क्लोन करें
git clone https://github.com/Tencent-Hunyuan/HunyuanImage-3.0.git
cd HunyuanImage-3.0

# वातावरण बनाएँ
pip install -r requirements.txt
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128

# मॉडल वेट्स डाउनलोड करें
huggingface-cli download tencent/HunyuanImage-3.0-Instruct --local-dir ./ckpts/HunyuanImage-3-Instruct
```

### Transformers के साथ टेक्स्ट-टू-इमेज

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

# मॉडल लोड करें (पूर्ण प्रिसिशन के लिए लगभग 80GB VRAM की आवश्यकता)
model_path = "./ckpts/HunyuanImage-3-Instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

# टेक्स्ट से इमेज जनरेट करें
prompt = "शरद ऋतु में एक शांत जापानी उद्यान, एक क्रिस्टल-सा साफ तालाब में तैरती कोई मछलियाँ, गिरते सुनहरे मेपल पत्ते, वॉटरकलर पेंटिंग शैली"
output = model.generate_image(prompt, num_inference_steps=30)
output.save("japanese_garden.png")
```

### Gradio वेब इंटरफ़ेस का उपयोग

सभी फीचर्स के साथ प्रयोग करने का सबसे आसान तरीका:

```bash
cd HunyuanImage-3.0

# Gradio इंस्टॉल करें
pip install gradio

# वेब इंटरफ़ेस लॉन्च करें
python gradio_demo.py \
    --model-path ./ckpts/HunyuanImage-3-Instruct \
    --server-name 0.0.0.0 \
    --server-port 7860
```

फिर SSH टनल के माध्यम से पहुँचें: `ssh -L 7860:localhost:7860 root@<clore-ip>`

## उपयोग के उदाहरण

### 1. टेक्स्ट-टू-इमेज जनरेशन (CLI)

```bash
cd HunyuanImage-3.0

python inference.py \
    --model-path ./ckpts/HunyuanImage-3-Instruct \
    --prompt "रात में एक साइबरपंक शहर का दृश्य, बारिश से भीगी सड़कों पर प्रतिबिंबित नीयन-रोशनी वाली गगनचुंबी इमारतें, उड़ती कारें, वॉल्यूमेट्रिक धुंध, 8K" \
    --output-path output.png \
    --num-inference-steps 30 \
    --guidance-scale 5.0
```

### 2. प्राकृतिक भाषा के साथ इमेज संपादन

HunyuanImage 3.0 की खास विशेषताओं में से एक — परिवर्तनों का वर्णन करके मौजूदा छवियों को संपादित करें:

```bash
python inference.py \
    --model-path ./ckpts/HunyuanImage-3-Instruct \
    --prompt "मौसम को सर्दियों में बदलें, जहाँ बर्फ़ पेड़ों को ढक रही हो" \
    --image-path input_photo.jpg \
    --output-path edited_winter.png \
    --num-inference-steps 30
```

### 3. डिस्टिल्ड मॉडल के साथ तेज़ जनरेशन (8 स्टेप)

```bash
# डिस्टिल्ड चेकपॉइंट डाउनलोड करें
huggingface-cli download tencent/HunyuanImage-3.0-Instruct-Distil \
    --local-dir ./ckpts/HunyuanImage-3-Instruct-Distil

# केवल 8 स्टेप के साथ जनरेट करें (5-6× तेज़)
python inference.py \
    --model-path ./ckpts/HunyuanImage-3-Instruct-Distil \
    --prompt "मंगल ग्रह पर घोड़े की सवारी करता हुआ एक अंतरिक्ष यात्री, फ़ोटो-यथार्थवादी" \
    --output-path astronaut.png \
    --num-inference-steps 8
```

## अन्य इमेज मॉडलों के साथ तुलना

| विशेषता            | HunyuanImage 3.0         | FLUX.2 Klein                | SD 3.5 Large             |
| ------------------ | ------------------------ | --------------------------- | ------------------------ |
| पैरामीटर           | 80B MoE (13B सक्रिय)     | 32B DiT                     | 8B DiT                   |
| आर्किटेक्चर        | ऑटोरिग्रेसिव MoE         | डिफ्यूज़न ट्रांसफ़ॉर्मर     | डिफ्यूज़न ट्रांसफ़ॉर्मर  |
| इमेज संपादन        | ✅ मूल रूप से समर्थित     | ❌ ControlNet की आवश्यकता है | ❌ img2img की आवश्यकता है |
| मल्टी-इमेज फ्यूज़न | ✅ मूल रूप से समर्थित     | ❌                           | ❌                        |
| शैली रूपांतरण      | ✅ मूल रूप से समर्थित     | ❌ LoRA की आवश्यकता है       | ❌ LoRA की आवश्यकता है    |
| न्यूनतम VRAM       | \~24GB (ऑफ़लोड किया गया) | 16GB                        | 8GB                      |
| गति (A100)         | \~15–30 सेकंड            | \~0.3 सेकंड                 | \~5 सेकंड                |
| लाइसेंस            | Tencent समुदाय           | Apache 2.0                  | Stability AI CL          |

## Clore.ai उपयोगकर्ताओं के लिए सुझाव

1. **गति के लिए डिस्टिल्ड मॉडल का उपयोग करें** — `HunyuanImage-3.0-Instruct-Distil` 30–50 के बजाय 8 स्टेप में जनरेट करता है, जिससे इन्फ़रेंस समय 4–6× तक घट जाता है। गुणवत्ता आश्चर्यजनक रूप से पूर्ण मॉडल के काफ़ी करीब रहती है।
2. **96GB कार्ड सबसे उपयुक्त है** — एक एकल RTX PRO 6000 Blackwell (या एक A100 80GB के माध्यम से [bare metal](https://clore.ai/bare-metal)) Instruct मॉडल को बिना किसी ऑफ़लोडिंग ट्रिक के चलाता है। यह CPU ऑफ़लोडिंग वाले RTX 4090 से कहीं तेज़ है।
3. **मॉडल पहले से डाउनलोड करें** — पूरा Instruct चेकपॉइंट लगभग 160GB का है। इसे एक स्थायी Clore.ai वॉल्यूम पर एक बार डाउनलोड करें ताकि हर बार नया instance शुरू करने पर फिर से डाउनलोड न करना पड़े।
4. **Gradio के लिए SSH टनलिंग का उपयोग करें** — पोर्ट 7860 को सार्वजनिक रूप से एक्सपोज़ न करें। उपयोग करें `ssh -L 7860:localhost:7860` अपने ब्राउज़र से वेब इंटरफ़ेस को सुरक्षित रूप से एक्सेस करने के लिए।
5. **बैच कार्य के लिए vLLM बैकएंड आज़माएँ** — यदि आप बहुत सारी छवियाँ जनरेट कर रहे हैं, तो vLLM इन्फ़रेंस पथ (में `vllm_infer/` फ़ोल्डर) काफी बेहतर थ्रूपुट प्रदान करता है।

## समस्या निवारण

| समस्या                                        | समाधान                                                                                                                                   |
| --------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `CUDA out of memory` RTX 4090 पर              | उपयोग करें `device_map="auto"` CPU ऑफ़लोडिंग सक्षम करने के लिए, या Distil मॉडल पर स्विच करें                                             |
| डाउनलोड विफल / बहुत धीमा                      | सेट करें `HF_TOKEN` env वेरिएबल; उपयोग करें `huggingface-cli download` के साथ `--resume-download`                                        |
| HF मॉडल ID के माध्यम से मॉडल लोड नहीं कर सकता | नाम में डॉट होने के कारण, पहले लोकल रूप से क्लोन करें: `huggingface-cli download tencent/HunyuanImage-3.0-Instruct --local-dir ./ckpts/` |
| धुंधला या कम-गुणवत्ता वाला आउटपुट             | बढ़ाएँ `--num-inference-steps` को 40–50 तक; बढ़ाएँ `--guidance-scale` को 7.0 तक                                                          |
| इमेज संपादन निर्देशों को नज़रअंदाज़ करता है   | क्या बदलना है और क्या सुरक्षित रखना है, इस बारे में विशिष्ट रहें; छोटे, स्पष्ट प्रॉम्प्ट का उपयोग करें                                   |
| Gradio इंटरफ़ेस शुरू नहीं हो रहा              | सुनिश्चित करें `gradio>=4.0` इंस्टॉल है; जाँचें कि मॉडल पाथ सही डायरेक्टरी की ओर इशारा कर रहा है                                         |

## अधिक पढ़ें

* [GitHub रिपॉज़िटरी](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0) — आधिकारिक कोड, इन्फ़रेंस स्क्रिप्ट्स, Gradio डेमो
* [HunyuanImage 3.0-Instruct (HuggingFace)](https://huggingface.co/tencent/HunyuanImage-3.0-Instruct) — पूर्ण मॉडल वेट्स
* [डिस्टिल्ड चेकपॉइंट](https://huggingface.co/tencent/HunyuanImage-3.0-Instruct-Distil) — 8-स्टेप तेज़ इन्फ़रेंस
* [तकनीकी रिपोर्ट (arXiv)](https://arxiv.org/pdf/2509.23951) — आर्किटेक्चर विवरण और बेंचमार्क
* [ComfyUI एकीकरण](https://github.com/bgreene2/ComfyUI-Hunyuan-Image-3) — Community ComfyUI कस्टम नोड


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-hi/image-generation/hunyuan-image3.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
