> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/language-models.md).

# Überblick

Führen Sie große Sprachmodelle (LLMs) auf CLORE.AI-GPUs für Inferenz- und Chat-Anwendungen aus.

## Beliebte Tools

| Tool                                                                                 | Anwendungsfall                                | Schwierigkeit |
| ------------------------------------------------------------------------------------ | --------------------------------------------- | ------------- |
| [Ollama](/guides/guides_v2-de/sprachmodelle/ollama.md)                               | Einfachstes LLM-Setup                         | Anfänger      |
| [Open WebUI](/guides/guides_v2-de/sprachmodelle/open-webui.md)                       | ChatGPT-ähnliche Oberfläche                   | Anfänger      |
| [vLLM](/guides/guides_v2-de/sprachmodelle/vllm.md)                                   | Produktionsbereitstellung mit hohem Durchsatz | Mittel        |
| [Llama.cpp Server](/guides/guides_v2-de/sprachmodelle/llamacpp-server.md)            | Effiziente GGUF-Inferenz                      | Einfach       |
| [Text Generation WebUI](/guides/guides_v2-de/sprachmodelle/text-generation-webui.md) | Funktionsreiche Chat-Oberfläche               | Einfach       |
| [ExLlamaV2](/guides/guides_v2-de/sprachmodelle/exllamav2-fast.md)                    | Schnellste EXL2-Inferenz                      | Mittel        |
| [LocalAI](/guides/guides_v2-de/sprachmodelle/localai-openai-compatible.md)           | Mit OpenAI kompatible API                     | Mittel        |
| [SGLang](/guides/guides_v2-de/sprachmodelle/sglang.md)                               | Schnelle strukturierte Generierung            | Mittel        |
| [Text Generation Inference (TGI)](/guides/guides_v2-de/sprachmodelle/tgi.md)         | HuggingFace-Bereitstellungslösung             | Mittel        |
| [LMDeploy](/guides/guides_v2-de/sprachmodelle/lmdeploy.md)                           | MMlab-Bereitstellungstoolkit                  | Mittel        |
| [Aphrodite Engine](/guides/guides_v2-de/sprachmodelle/aphrodite-engine.md)           | vLLM-Fork mit zusätzlichen Funktionen         | Mittel        |
| [MLC-LLM](/guides/guides_v2-de/sprachmodelle/mlc-llm.md)                             | Kompilierung für maschinelles Lernen          | Schwer        |
| [LiteLLM](/guides/guides_v2-de/sprachmodelle/litellm.md)                             | Einheitlicher API-Proxy                       | Mittel        |
| [PowerInfer](/guides/guides_v2-de/sprachmodelle/powerinfer.md)                       | Inferenz für spärliche Modelle                | Schwer        |
| [Mistral.rs](/guides/guides_v2-de/sprachmodelle/mistral-rs.md)                       | Rust-basierte Inferenz-Engine                 | Mittel        |

## Modellanleitungen

### Neueste & beste Modelle

| Modell                                                                     | Parameter                 | Lizenz            | Passt auf                 |
| -------------------------------------------------------------------------- | ------------------------- | ----------------- | ------------------------- |
| [Qwen3.8-27B](/guides/guides_v2-de/sprachmodelle/qwen38-27b.md)            | 27B dicht                 | Apache 2.0        | 1× RTX 4090 (Q4)          |
| [Mistral Small 4](/guides/guides_v2-de/sprachmodelle/mistral-small4.md)    | 119B / 6,5B aktiv         | Apache 2.0        | 2× RTX 4090 (Q2)          |
| [GLM-5.2](/guides/guides_v2-de/sprachmodelle/glm-5-2.md)                   | \~750B / 40B aktiv        | MIT               | 10× RTX 5090 (Q2)         |
| [DeepSeek V4](/guides/guides_v2-de/sprachmodelle/deepseek-v4.md)           | Flash 167 GB · Pro 893 GB | MIT               | 6–8× RTX 5090 (Flash)     |
| [MiniMax M3](/guides/guides_v2-de/sprachmodelle/minimax-m3.md)             | 428B / 23B aktiv          | minimax-community | 6–8× RTX 5090 (Q2–Q3)     |
| [Nemotron 3 Ultra](/guides/guides_v2-de/sprachmodelle/nemotron-3-ultra.md) | 550B / 55B aktiv          | OpenMDW-1.1       | 4× RTX PRO 6000 (NVFP4)   |
| [Kimi K3](/guides/guides_v2-de/sprachmodelle/kimi-k3.md)                   | 2,8T                      | Kimi-K3-Lizenz    | Nichts auf dem Marktplatz |
| [GLM-5.1](/guides/guides_v2-de/sprachmodelle/glm-5-1.md)                   | 744B / 40B aktiv          | MIT               | Große Rigs                |
| [Llama 4](/guides/guides_v2-de/sprachmodelle/llama4.md)                    | Scout & Maverick          | Llama-Community   | 1× RTX 4090 (Scout Q4)    |
| [DeepSeek-R1](/guides/guides_v2-de/sprachmodelle/deepseek-r1.md)           | 671B MoE                  | MIT               | Große Rigs                |
| [Qwen2.5](/guides/guides_v2-de/sprachmodelle/qwen25.md)                    | 0,5B–72B                  | Apache 2.0        | Jede Karte                |

### Spezialisierte Modelle

| Modell                                                                 | Parameter                 | Am besten geeignet für        |
| ---------------------------------------------------------------------- | ------------------------- | ----------------------------- |
| [DeepSeek Coder](/guides/guides_v2-de/sprachmodelle/deepseek-coder.md) | 6,7B–33B                  | Code-Generierung              |
| [CodeLlama](/guides/guides_v2-de/sprachmodelle/codellama.md)           | 7B–34B                    | Code-Vervollständigung        |
| [GLM-4.7-Flash](/guides/guides_v2-de/sprachmodelle/glm-47-flash.md)    | 4,7B                      | Schnelles Chinesisch/Englisch |
| [GLM-5](/guides/guides_v2-de/sprachmodelle/glm5.md)                    | Wird noch bekannt gegeben | Neueste Zhipu AI              |
| [Kimi K2.5](/guides/guides_v2-de/sprachmodelle/kimi-k2.md)             | Wird noch bekannt gegeben | Moonshot-AI-Modell            |
| [Ling-2.5-1T](/guides/guides_v2-de/sprachmodelle/ling25.md)            | 1T                        | Riesiges Open-Source-LLM      |
| [LFM2-24B](/guides/guides_v2-de/sprachmodelle/lfm2-24b.md)             | 24B                       | Liquid AI-Modell              |
| [MiMo-V2-Flash](/guides/guides_v2-de/sprachmodelle/mimo-v2-flash.md)   | Wird noch bekannt gegeben | Schnelles Inferenzmodell      |

### Effiziente Modelle

| Modell                                                                   | Parameter                 | Am besten geeignet für            |
| ------------------------------------------------------------------------ | ------------------------- | --------------------------------- |
| [Gemma 2](/guides/guides_v2-de/sprachmodelle/gemma2.md)                  | 2B–27B                    | Effiziente Inferenz               |
| [Gemma 3](/guides/guides_v2-de/sprachmodelle/gemma3.md)                  | Wird noch bekannt gegeben | Googles neuestes kompaktes Modell |
| [Phi-4](/guides/guides_v2-de/sprachmodelle/phi4.md)                      | 14B                       | Klein, aber leistungsfähig        |
| [Mistral/Mixtral](/guides/guides_v2-de/sprachmodelle/mistral-mixtral.md) | 7B / 8x7B                 | Allgemeiner Zweck                 |
| [Mistral Large 3](/guides/guides_v2-de/sprachmodelle/mistral-large3.md)  | 675B MoE                  | Für Unternehmen geeignet          |
| [Mistral Small 3.1](/guides/guides_v2-de/sprachmodelle/mistral-small.md) | Wird noch bekannt gegeben | Effiziente Mistral-Variante       |

## GPU-Empfehlungen

{% hint style="warning" %}
**Multi-GPU-Rigs der 80GB-Klasse sind auf dem Clore.ai-Marktplatz nicht gelistet.** Die größten heute gelisteten Systeme sind 4× RTX PRO 6000 Blackwell (je 96 GB, 380 GB gesamt) und 8–11× RTX 5090 (je 32 GB). Kapazitäten für A100 / H200 / B200 werden als [Bare Metal](https://clore.ai/bare-metal) auf Anfrage verkauft. Prüfe [GPU-Preise & Verfügbarkeit](/guides/guides_v2-de/erste-schritte/pricing.md) bevor du eine Bereitstellung dimensionierst.
{% endhint %}

| Modellgröße | Minimale GPU   | Empfohlen            | Clore.ai-Kosten      |
| ----------- | -------------- | -------------------- | -------------------- |
| 7B (Q4)     | RTX 3060 12 GB | RTX 3090             | 0,03–0,21 $/Std.     |
| 13B (Q4)    | RTX 3090 24 GB | RTX 4090             | $0.07–0.42/Std.      |
| 27B (Q4)    | RTX 3090 24 GB | RTX 4090 / 5090      | 0,07–0,77 $/Std.     |
| 34B (Q4)    | 2× RTX 3090    | 1× RTX 5090          | 0,14–0,77 $/Std.     |
| 70B (Q4)    | 2× RTX 5090    | 1× RTX PRO 6000 96GB | $0.50–1.54/Std.      |
| 120B+ (Q2)  | 2× RTX 4090    | 2× RTX 5090          | 0,28–1,54 $/Std.     |
| 400B+ (Q2)  | 8× RTX 5090    | 4× RTX PRO 6000      | ca. 2,00–5,00 $/Std. |

## Quantisierungsleitfaden

| Format   | VRAM-Nutzung | Qualität     | Geschwindigkeit |
| -------- | ------------ | ------------ | --------------- |
| Q2\_K    | Niedrigste   | Schlecht     | Schnellste      |
| Q4\_K\_M | Niedrig      | Gut          | Schnell         |
| Q5\_K\_M | Mittel       | Großartig    | Mittel          |
| Q8\_0    | Hoch         | Hervorragend | Langsamer       |
| FP16     | Am höchsten  | Am besten    | Langsamste      |

## Siehe auch

* [Training & Feinabstimmung](/guides/guides_v2-de/training/training.md)
* [Vision-Language-Modelle](/guides/guides_v2-de/computer-vision-modelle/vision-models.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/language-models.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
