> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/kimi-k3.md).

# Kimi K3 (2,8T — Was es braucht)

Was Kimi K3 ist, was es wirklich braucht, um 2,8 Billionen Parameter selbst zu hosten, und was stattdessen auf Clore.ai ausgeführt werden sollte

{% hint style="danger" %}
**Die kurze Antwort vorweg: Kimi K3 passt auf keinen der auf dem Clore.ai-Marktplatz gelisteten Server, bei keiner Quantisierung.** Der kleinste Community-Build ist ein 1-Bit-GGUF mit **466 GB**; das größte Rig auf dem Marktplatz hat 380 GB. Diese Seite erklärt, was K3 ist, was das Hosten tatsächlich erfordert und welche Modelle dir heute auf mietbarer Hardware die meiste Leistung bieten.
{% endhint %}

Moonshot AI veröffentlichte **Kimi K3** Gewichte am **27. Juli 2026** — das erste offene Modell in der 3-Billionen-Parameter-Klasse. Es ist ein echter Meilenstein und zugleich ein nützlicher Kalibrierpunkt dafür, wie weit sich „offene Gewichte“ und „selbst hostbar“ auseinanderentwickelt haben.

### Wichtige Spezifikationen

| Eigenschaft      | Wert                                                                                                                               |
| ---------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Parameter        | 2,8 T gesamt, 16 von 896 Experten pro Token aktiv                                                                                  |
| Architektur      | Kimi Delta Attention (KDA) + Attention Residuals, Stable LatentMoE, 69 KDA + 24 Gated-MLA-Layer                                    |
| Modalität        | Text, Bild, Video rein → Text raus                                                                                                 |
| Kontext          | 1.000.000 Tokens                                                                                                                   |
| Lizenz           | Kimi-K3-Lizenz (benutzerdefiniert — kommerzielle Inferenz oberhalb von etwa 20 Mio. $ Jahresumsatz erfordert separate Bedingungen) |
| Veröffentlichung | Gehostet am 16. Juli 2026, Gewichte am 27. Juli 2026                                                                               |
| Gewichte         | Q8 1.561 GB · Q4\_K\_XL 1.509 GB · Q2\_K\_XL 861 GB · IQ1\_S 594 GB · Q1\_0 466 GB                                                 |

Moonshot gibt an, etwa **2,5× bessere Skalierungseffizienz als Kimi K2** durch das spärlichere MoE und den KDA-Attention-Stack, mit nativem multimodalem Training statt eines nachträglich aufgesetzten Vision-Towers.

## Was das Hosten tatsächlich erfordert

| Build                    | Größe    | Benötigte Hardware            |
| ------------------------ | -------- | ----------------------------- |
| `UD-Q8_K_XL`             | 1.561 GB | \~20× H100 80 GB              |
| `UD-Q4_K_XL`             | 1.509 GB | \~20× H100 80 GB              |
| `UD-Q2_K_XL`             | 861 GB   | \~11× H100 80 GB oder 8× H200 |
| `UD-IQ1_S`               | 594 GB   | \~8× H100 80 GB               |
| `UD-Q1_0`                | 466 GB   | \~6× H100 80 GB               |
| `mgoin/Kimi-K3-pruned75` | 475 GB   | \~6× H100 80 GB               |

Zum Vergleich: Die größten Rigs auf dem Clore.ai-Marktplatz sind 4× RTX PRO 6000 Blackwell (380 GB) und 11× RTX 5090 (341 GB). Selbst der 1-Bit-Build übersteigt sie, und eine 1-Bit-Quantisierung eines Frontier-Modells ist ohnehin kein ernstzunehmendes Produktionsziel.

{% hint style="info" %}
**Wenn du K3 wirklich bereitstellen musst**ist das [Bare Metal](https://clore.ai/bare-metal) eine andere Liga — dedizierte H200- oder B200-Nodes, auf Anfrage bereitgestellt statt minutenweise gemietet. Prüfe zuerst die Lizenz: Die Bedingungen beschränken kommerzielle Inferenz oberhalb von etwa 20 Mio. $ Jahresumsatz.
{% endhint %}

## Was stattdessen auf Clore.ai ausführen

| Wenn du K3 für … wolltest           | Dann nimm das hier                                                                                           | Passt auf                      |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------ | ------------------------------ |
| Spitzenklasse-Coding und Agenten    | [GLM-5.2](/guides/guides_v2-de/sprachmodelle/glm-5-2.md) — MIT, 1 Mio. Kontext, Top-Werte bei offenem Coding | 10× RTX 5090 Rig, ca. 5,00 $/h |
| Multimodal mit langem Kontext       | [MiniMax M3](/guides/guides_v2-de/sprachmodelle/minimax-m3.md) — 1 Mio. Kontext, native Videoverarbeitung    | 8× RTX 5090, ca. 2,00–3,50 $/h |
| Offenes Reasoning im großen Maßstab | [Nemotron 3 Ultra](/guides/guides_v2-de/sprachmodelle/nemotron-3-ultra.md) — 550B, OpenMDW-Lizenz            | 4× RTX PRO 6000, ca. 5,00 $/h  |
| Ein Modell auf einer Karte          | [Qwen3.8-27B](/guides/guides_v2-de/sprachmodelle/qwen38-27b.md) — Apache 2.0, Vision, 262K Kontext           | 1× RTX 4090, 0,14–0,42 $/h     |
| Beste Qualität pro gemietetem GB    | [Mistral Small 4](/guides/guides_v2-de/sprachmodelle/mistral-small4.md) — 119B / 6,5B aktiv                  | 2× RTX 4090, 0,28–0,84 $/h     |

Die ehrliche Zusammenfassung: Seit Juni 2026 hat sich die offene Frontier aufgespalten. Modelle wie GLM-5.2 und Nemotron 3 Ultra liegen am Rand dessen, was ein großer Consumer-Rig aufnehmen kann; K3 liegt darüber. Der praktische Abstand zwischen ihnen schrumpft größtenteils durch ein gutes Harness und einen langen Kontext — beides liefern dir die Modelle oben.

## K3 nutzen, ohne es zu hosten

Moonshot stellt K3 über die Kimi-App, Kimi Code und seine API bereit. Wenn du auf Clore.ai baust, ist ein gängiges Muster, die teuren Frontier-Aufrufe über eine API laufen zu lassen und den Pfad mit hohem Volumen lokal auf einer gemieteten GPU auszuführen — siehe den Kostenvergleich in [Gemini 3.1 Flash Lite](/guides/guides_v2-de/sprachmodelle/gemini-3-1-flash-lite.md)— der durchgeht, wann Self-Hosting tatsächlich gewinnt.

## Nächste Schritte

* [GLM-5.2](/guides/guides_v2-de/sprachmodelle/glm-5-2.md) — das größte Modell, das du hier realistisch mieten kannst
* [GPU-Preise & Verfügbarkeit](/guides/guides_v2-de/erste-schritte/pricing.md) — wo die Obergrenze derzeit liegt
* [Qwen3.8-27B](/guides/guides_v2-de/sprachmodelle/qwen38-27b.md) — das Arbeitstier für eine einzelne Karte
* [Kimi K2.5](/guides/guides_v2-de/sprachmodelle/kimi-k2.md) — die vorherige Generation, deutlich besser handhabbar

### Links

* [Kimi K3 auf Hugging Face](https://huggingface.co/moonshotai/Kimi-K3) · [GGUF-Builds](https://huggingface.co/unsloth/Kimi-K3-GGUF)
* [Kimi-K3-Lizenz](https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE)
* **Eine GPU mieten:** [Marktplatz](https://clore.ai/marketplace) · [Bare Metal](https://clore.ai/bare-metal)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/kimi-k3.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
