> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-de/erste-schritte/gpu-comparison.md).

# GPU-Vergleich

Vollständiger GPU-Vergleichsleitfaden für KI-Workloads auf Clore.ai

Vollständiger Vergleich der auf CLORE.AI verfügbaren GPUs für KI-Workloads.

{% hint style="success" %}
Finden Sie die passende GPU für Ihre Aufgabe bei [CLORE.AI-Marktplatz](https://clore.ai/marketplace).
{% endhint %}

## Schnelle Empfehlung

| Ihre Aufgabe             | Preisgünstige Wahl | Bestes Preis-Leistungs-Verhältnis | Maximale Leistung |
| ------------------------ | ------------------ | --------------------------------- | ----------------- |
| Mit KI chatten (7B)      | RTX 3060 12 GB     | RTX 3090 24 GB                    | RTX 5090 32 GB    |
| Mit KI chatten (70B)     | RTX 3090 24 GB     | RTX 5090 32 GB                    | A100 80GB         |
| Bildgenerierung (SD 1.5) | RTX 3060 12 GB     | RTX 3090 24 GB                    | RTX 5090 32 GB    |
| Bildgenerierung (SDXL)   | RTX 3090 24 GB     | RTX 4090 24GB                     | RTX 5090 32 GB    |
| Bildgenerierung (FLUX)   | RTX 3090 24 GB     | RTX 5090 32 GB                    | A100 80GB         |
| Videogenerierung         | RTX 4090 24GB      | RTX 5090 32 GB                    | A100 80GB         |
| Modelltraining           | A100 40GB          | A100 80GB                         | H100 80 GB        |

## Consumer-GPUs

### NVIDIA RTX 3060 12GB

**Am besten für:** Budget-KI, SD 1.5, kleine LLMs

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 12GB GDDR6    |
| Speicherbandbreite | 360 GB/s      |
| FP16-Leistung      | 12,7 TFLOPS   |
| Tensor-Kerne       | 112 (3. Gen.) |
| TDP                | 170W          |
| \~Preis/Stunde     | $0.02-0.04    |

**Fähigkeiten:**

* ✅ Ollama mit 7B-Modellen (Q4)
* ✅ Stable Diffusion 1.5 (512x512)
* ✅ SDXL (768x768, langsam)
* ⚠️ FLUX schnell (mit CPU-Offload)
* ❌ Große Modelle (>13B)
* ❌ Videogenerierung

***

### NVIDIA RTX 3070/3070 Ti 8GB

**Am besten für:** SD 1.5, leichte Aufgaben

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 8GB GDDR6X    |
| Speicherbandbreite | 448-608 GB/s  |
| FP16-Leistung      | 20,3 TFLOPS   |
| Tensor-Kerne       | 184 (3. Gen.) |
| TDP                | 220-290W      |
| \~Preis/Stunde     | $0.02-0.04    |

**Fähigkeiten:**

* ✅ Ollama mit 7B-Modellen (Q4)
* ✅ Stable Diffusion 1.5 (512x512)
* ⚠️ SDXL (nur niedrige Auflösung)
* ❌ FLUX (zu wenig VRAM)
* ❌ Modelle >7B
* ❌ Videogenerierung

***

### NVIDIA RTX 3080/3080 Ti 10-12GB

**Am besten für:** Allgemeine KI-Aufgaben, gutes Gleichgewicht

| Spezifikation      | Wert              |
| ------------------ | ----------------- |
| VRAM               | 10-12GB GDDR6X    |
| Speicherbandbreite | 760-912 GB/s      |
| FP16-Leistung      | 29,8-34,1 TFLOPS  |
| Tensor-Kerne       | 272-320 (3. Gen.) |
| TDP                | 320-350W          |
| \~Preis/Stunde     | $0.04-0.06        |

**Fähigkeiten:**

* ✅ Ollama mit 13B-Modellen
* ✅ Stable Diffusion 1.5/2.1
* ✅ SDXL (1024x1024)
* ⚠️ FLUX schnell (mit Offload)
* ❌ Große Modelle (>13B)
* ❌ Videogenerierung

***

### NVIDIA RTX 3090/3090 Ti 24GB

**Am besten für:** SDXL, 13B-30B-LLMs, ControlNet

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 24GB GDDR6X   |
| Speicherbandbreite | 936 GB/s      |
| FP16-Leistung      | 35,6 TFLOPS   |
| Tensor-Kerne       | 328 (3. Gen.) |
| TDP                | 350-450W      |
| \~Preis/Stunde     | $0.05-0.08    |

**Fähigkeiten:**

* ✅ Ollama mit 30B-Modellen
* ✅ vLLM mit 13B-Modellen
* ✅ Alle Stable-Diffusion-Modelle
* ✅ SDXL + ControlNet
* ✅ FLUX schnell (1024x1024)
* ⚠️ FLUX dev (mit Offload)
* ⚠️ Video (kurze Clips)

***

### NVIDIA RTX 4070 Ti 12GB

**Am besten für:** Schnelles SD 1.5, effiziente Inferenz

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 12GB GDDR6X   |
| Speicherbandbreite | 504 GB/s      |
| FP16-Leistung      | 40,1 TFLOPS   |
| Tensor-Kerne       | 184 (4. Gen.) |
| TDP                | 285W          |
| \~Preis/Stunde     | $0.04-0.06    |

**Fähigkeiten:**

* ✅ Ollama mit 7B-Modellen (schnell)
* ✅ Stable Diffusion 1.5 (sehr schnell)
* ✅ SDXL (768x768)
* ⚠️ FLUX schnell (begrenzte Auflösung)
* ❌ Große Modelle (>13B)
* ❌ Videogenerierung

***

### NVIDIA RTX 4080 16GB

**Am besten für:** SDXL-Produktion, 13B-LLMs

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 16GB GDDR6X   |
| Speicherbandbreite | 717 GB/s      |
| FP16-Leistung      | 48,7 TFLOPS   |
| Tensor-Kerne       | 304 (4. Gen.) |
| TDP                | 320W          |
| \~Preis/Stunde     | $0.06-0.09    |

**Fähigkeiten:**

* ✅ Ollama mit 13B-Modellen (schnell)
* ✅ vLLM mit 7B-Modellen
* ✅ Alle Stable-Diffusion-Modelle
* ✅ SDXL + ControlNet
* ✅ FLUX schnell (1024x1024)
* ⚠️ FLUX dev (eingeschränkt)
* ⚠️ Kurze Videoclips

***

### NVIDIA RTX 4090 24GB

**Am besten für:** High-End-Consumer-Leistung, FLUX, Video

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 24GB GDDR6X   |
| Speicherbandbreite | 1008 GB/s     |
| FP16-Leistung      | 82,6 TFLOPS   |
| Tensor-Kerne       | 512 (4. Gen.) |
| TDP                | 450W          |
| \~Preis/Stunde     | $0.08-0.12    |

**Fähigkeiten:**

* ✅ Ollama mit 30B-Modellen (schnell)
* ✅ vLLM mit 13B-Modellen
* ✅ Alle Bildgenerierungsmodelle
* ✅ FLUX dev (1024x1024)
* ✅ Videogenerierung (kurz)
* ✅ AnimateDiff
* ⚠️ 70B-Modelle (nur Q4)

***

### NVIDIA RTX 5080 16GB *(Neu — Feb. 2025)*

**Am besten für:** Schnelles SDXL/FLUX, 13B-30B-LLMs, leistungsstarke Mittelklasse

| Spezifikation           | Wert          |
| ----------------------- | ------------- |
| VRAM                    | 16GB GDDR7    |
| Speicherbandbreite      | 960 GB/s      |
| FP16-Leistung           | \~80 TFLOPS   |
| Tensor-Kerne            | 336 (5. Gen.) |
| TDP                     | 360W          |
| \~Clore.ai Preis/Stunde | $1.50-2.00    |

**Fähigkeiten:**

* ✅ Ollama mit 13B-Modellen (schnell)
* ✅ vLLM mit 13B-Modellen
* ✅ Alle Stable-Diffusion-Modelle
* ✅ SDXL + ControlNet (sehr schnell)
* ✅ FLUX schnell/dev (1024x1024)
* ✅ Kurze Videoclips
* ⚠️ 30B-Modelle (nur Q4)
* ❌ 70B-Modelle

***

### NVIDIA RTX 5090 32GB *(Flaggschiff — Feb. 2025)*

**Am besten für:** Maximale Consumer-Leistung, 70B-Modelle, hochauflösende Videogenerierung

| Spezifikation           | Wert          |
| ----------------------- | ------------- |
| VRAM                    | 32GB GDDR7    |
| Speicherbandbreite      | 1792 GB/s     |
| FP16-Leistung           | \~120 TFLOPS  |
| Tensor-Kerne            | 680 (5. Gen.) |
| TDP                     | 575W          |
| \~Clore.ai Preis/Stunde | $3.00-4.00    |

**Fähigkeiten:**

* ✅ Ollama mit 70B-Modellen (Q4, schnell)
* ✅ vLLM mit 30B-Modellen
* ✅ Alle Bildgenerierungsmodelle
* ✅ FLUX dev (1536x1536)
* ✅ Videogenerierung (längere Clips)
* ✅ AnimateDiff + ControlNet
* ✅ Modell-Finetuning (LoRA, kleine Fine-Tunes)
* ✅ DeepSeek-R1 32B destilliert (FP16)

## Profi-/Rechenzentrums-GPUs

{% hint style="warning" %}
**Multi-GPU-Rigs der 80GB-Klasse sind auf dem Clore.ai-Marktplatz nicht gelistet.** Die größten heute gelisteten Systeme sind 4× RTX PRO 6000 Blackwell (je 96 GB, 380 GB gesamt) und 8–11× RTX 5090 (je 32 GB). Kapazitäten für A100 / H200 / B200 werden als [Bare Metal](https://clore.ai/bare-metal) auf Anfrage verkauft. Prüfe [GPU-Preise & Verfügbarkeit](/guides/guides_v2-de/erste-schritte/pricing.md) bevor du eine Bereitstellung dimensionierst.
{% endhint %}

### NVIDIA A100 40GB

**Am besten für:** Produktions-LLMs, Training, große Modelle

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 40GB HBM2e    |
| Speicherbandbreite | 1555 GB/s     |
| FP16-Leistung      | 77,97 TFLOPS  |
| Tensor-Kerne       | 432 (3. Gen.) |
| TDP                | 400W          |
| \~Preis/Stunde     | $0.15-0.20    |

**Fähigkeiten:**

* ✅ Ollama mit 70B-Modellen (Q4)
* ✅ vLLM-Serving in der Produktion
* ✅ Alle Bildgenerierungen
* ✅ FLUX dev (hohe Qualität)
* ✅ Videogenerierung
* ✅ Modell-Finetuning
* ⚠️ 70B FP16 (eng)

***

### NVIDIA A100 80GB

**Am besten für:** 70B+-Modelle, Video, Produktions-Workloads

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 80GB HBM2e    |
| Speicherbandbreite | 2039 GB/s     |
| FP16-Leistung      | 77,97 TFLOPS  |
| Tensor-Kerne       | 432 (3. Gen.) |
| TDP                | 400W          |
| \~Preis/Stunde     | $0.20-0.30    |

**Fähigkeiten:**

* ✅ Alle LLMs bis zu 70B (FP16)
* ✅ vLLM-Serving mit hohem Durchsatz
* ✅ Alle Bildgenerierungen
* ✅ Lange Videogenerierung
* ✅ Modell-Training
* ✅ DeepSeek-V3 (teilweise)
* ⚠️ 100B+-Modelle

***

### NVIDIA H100 80GB

**Am besten für:** Maximale Leistung, größte Modelle

| Spezifikation      | Wert          |
| ------------------ | ------------- |
| VRAM               | 80GB HBM3     |
| Speicherbandbreite | 3350 GB/s     |
| FP16-Leistung      | 267 TFLOPS    |
| Tensor-Kerne       | 528 (4. Gen.) |
| TDP                | 700W          |
| \~Preis/Stunde     | $0.40-0.60    |

**Fähigkeiten:**

* ✅ Alle Modelle mit maximaler Geschwindigkeit
* ✅ Modelle mit 100B+ Parametern
* ✅ Serving mehrerer Modelle
* ✅ Training im großen Maßstab
* ✅ Echtzeit-Videogenerierung
* ✅ DeepSeek-V3 (671B)

## Leistungsvergleiche

### LLM-Inferenz (Token/Sekunde)

| GPU            | Llama 3 8B | Llama 3 70B | Mixtral 8x7B | Clore.ai $/Std. |
| -------------- | ---------- | ----------- | ------------ | --------------- |
| RTX 3060 12 GB | 25         | -           | -            | $0.02-0.04      |
| RTX 3090 24 GB | 45         | 8\*         | 20\*         | $0.15-0.25      |
| RTX 4090 24GB  | 80         | 15\*        | 35\*         | $0.35-0.55      |
| RTX 5080 16GB  | 95         | -           | 40\*         | $1.50-2.00      |
| RTX 5090 32 GB | 150        | 30\*        | 65\*         | $3.00-4.00      |
| A100 40GB      | 100        | 25          | 45           | $0.80-1.20      |
| A100 80GB      | 110        | 40          | 55           | $1.20-1.80      |
| H100 80 GB     | 180        | 70          | 90           | $2.50-3.50      |

\*Mit Quantisierung (Q4/Q8)

### Geschwindigkeit der Bildgenerierung

| GPU            | SD 1.5 (512) | SDXL (1024) | FLUX schnell | Clore.ai $/Std. |
| -------------- | ------------ | ----------- | ------------ | --------------- |
| RTX 3060 12 GB | 4 Sek.       | 15 Sek.     | 25 Sek.\*    | $0.02-0.04      |
| RTX 3090 24 GB | 2 Sek.       | 7 Sek.      | 12 Sek.      | $0.15-0.25      |
| RTX 4090 24GB  | 1 Sek.       | 3 Sek.      | 5 Sek.       | $0.35-0.55      |
| RTX 5080 16GB  | 0,8 Sek.     | 2,5 Sek.    | 4 Sek.       | $1.50-2.00      |
| RTX 5090 32 GB | 0,6 Sek.     | 1,8 Sek.    | 3 Sek.       | $3.00-4.00      |
| A100 40GB      | 1,5 Sek.     | 4 Sek.      | 6 Sek.       | $0.80-1.20      |
| A100 80GB      | 1,5 Sek.     | 4 Sek.      | 5 Sek.       | $1.20-1.80      |

\*Mit CPU-Offload, geringere Auflösung

### Videogenerierung (5-Sekunden-Clip)

| GPU            | SVD      | Wan2.1   | Hunyuan  |
| -------------- | -------- | -------- | -------- |
| RTX 3090 24 GB | 3 Min.   | 5 Min.\* | -        |
| RTX 4090 24GB  | 1,5 Min. | 3 Min.   | 8 Min.\* |
| RTX 5090 32 GB | 1 Min.   | 2 Min.   | 5 Min.   |
| A100 40GB      | 1 Min.   | 2 Min.   | 5 Min.   |
| A100 80GB      | 45 Sek.  | 1,5 Min. | 3 Min.   |

\*Begrenzte Auflösung

## Preis-/Leistungsverhältnis

### Bestes Preis-Leistungs-Verhältnis je Aufgabe

**Chat/LLM (7B-30B-Modelle):**

1. 🥇 RTX 3090 24GB - Bestes Preis-Leistungs-Verhältnis
2. 🥈 RTX 3060 12GB - Niedrigste Kosten
3. 🥉 RTX 4090 24GB - Schnellste

**Bildgenerierung (SDXL/FLUX):**

1. 🥇 RTX 3090 24GB - Tolle Balance
2. 🥈 RTX 4090 24GB - 2x schneller
3. 🥉 A100 40GB - Produktionsstabilität

**Große Modelle (70B+):**

1. 🥇 A100 40GB - Bestes Preis-Leistungs-Verhältnis für 70B
2. 🥈 A100 80GB - Volle Präzision
3. 🥉 RTX 4090 24GB - Budget-Option (nur Q4)

**Videogenerierung:**

1. 🥇 A100 40GB - Gutes Gleichgewicht
2. 🥈 RTX 4090 24GB - Consumer-Option
3. 🥉 A100 80GB - Längste Clips

**Modelltraining:**

1. 🥇 A100 40GB - Standardwahl
2. 🥈 A100 80GB - Große Modelle
3. 🥉 RTX 4090 24GB - Kleine Modelle/LoRA

## Multi-GPU-Konfigurationen

Einige Aufgaben profitieren von mehreren GPUs:

| Konfiguration | Anwendungsfall                 | Gesamter VRAM |
| ------------- | ------------------------------ | ------------- |
| 2x RTX 3090   | 70B-Inferenz                   | 48GB          |
| 2x RTX 4090   | Schnelle 70B-Modelle, Training | 48GB          |
| 2x RTX 5090   | 70B FP16, schnelles Training   | 64 GB         |
| 4x RTX 5090   | 100B+-Modelle                  | 128GB         |
| 4x A100 40GB  | 100B+-Modelle                  | 160GB         |
| 8x A100 80 GB | DeepSeek-V3, Llama 405B        | 640GB         |

## Ihre GPU auswählen

### Entscheidungsflussdiagramm

```
Was ist Ihre Hauptaufgabe?
│
├─ Chat/LLM
│  ├─ Modellgröße?
│  │  ├─ ≤7B → RTX 3060 ($0,03–0,07/Std.)
│  │  ├─ 7B-30B → RTX 3090 ($0,07–0,21/Std.)
│  │  ├─ 30B-70B → RTX 5090 32GB oder 2× RTX 4090
│  │  └─ 70B+ → RTX PRO 6000 96GB oder 2× RTX 5090
│
├─ Bildgenerierung
│  ├─ Modell?
│  │  ├─ SD 1.5 → RTX 3060 ($0,03–0,07/Std.)
│  │  ├─ SDXL → RTX 3090 ($0,07–0,21/Std.)
│  │  └─ FLUX → RTX 4090 ($0,14–0,42/Std.)
│
├─ Videogenerierung
│  ├─ Länge?
│  │  ├─ Kurz (2-5 Sek.) → RTX 4090 ($0,14–0,42/Std.)
│  │  └─ Länger → 2× RTX 4090 oder RTX PRO 6000
│
└─ Training
   ├─ LoRA/klein → RTX 4090 ($0,14–0,42/Std.)
   └─ Vollständiges Fine-Tuning → 2× RTX 4090 oder RTX PRO 6000
```

## Tipps zum Geldsparen

1. **Spot-Orders nutzen** - etwa ein Drittel der Serverpreise im Spot-Modus liegt unter On-Demand, Median \~13 % Rabatt
2. **Klein anfangen** - Erst auf günstigeren GPUs testen
3. **Modelle quantisieren** - Q4/Q8 bringt größere Modelle in weniger VRAM unter
4. **Batch-Verarbeitung** - Mehrere Anfragen gleichzeitig verarbeiten
5. **Nebenzeiten** - Bessere Verfügbarkeit und manchmal niedrigere Preise

> 📚 Siehe auch: [Die 10 günstigsten GPUs für KI-Training 2025](https://blog.clore.ai/top-10-cheapest-gpus-for-ai-training/) | [Beste GPU für KI-Training — Detaillierter Leitfaden](https://blog.clore.ai/best-gpu-for-ai-training/)

## Nächste Schritte

* [Modellkompatibilitätsmatrix](/guides/guides_v2-de/erste-schritte/model-compatibility.md) - Welche Modelle auf welchen GPUs laufen
* [Docker-Image-Katalog](/guides/guides_v2-de/erste-schritte/docker-images.md) - Sofort einsatzbereite Images
* [Schnellstart-Anleitung](/guides/guides_v2-de/quickstart.md) - In 5 Minuten loslegen


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-de/erste-schritte/gpu-comparison.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
