> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle.md).

# Sprachmodelle

- [Überblick](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/language-models.md)
- [Ollama](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/ollama.md): Führe LLMs lokal mit Ollama auf Clore.ai-GPUs aus
- [Open WebUI](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/open-webui.md): ChatGPT-ähnliche Oberfläche zum Ausführen von LLMs auf Clore.ai-GPUs
- [vLLM](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/vllm.md): LLM-Inferenz mit hohem Durchsatz mit vLLM auf Clore.ai-GPUs
- [Llama.cpp Server](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/llamacpp-server.md): Effiziente LLM-Inferenz mit llama.cpp-Server auf Clore.ai-GPUs
- [Text Generation WebUI](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/text-generation-webui.md): Führe text-generation-webui für LLM-Inferenz auf Clore.ai-GPUs aus
- [ExLlamaV2](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/exllamav2-fast.md): Maximale Geschwindigkeit bei der LLM-Inferenz mit ExLlamaV2 auf Clore.ai-GPUs
- [LocalAI](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/localai-openai-compatible.md): Self-gehostete, mit OpenAI kompatible API mit LocalAI auf Clore.ai
- [Llama 3.3 70B](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/llama33.md): Führe Metas Llama-3.3-70B-Modell auf Clore.ai-GPUs aus
- [Mistral & Mixtral](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-mixtral.md): Führe Mistral- und Mixtral-Modelle auf Clore.ai-GPUs aus
- [DeepSeek Coder](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/deepseek-coder.md): Best-in-class Codegenerierung mit DeepSeek Coder auf Clore.ai
- [DeepSeek-V3](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/deepseek-v3.md): Führe DeepSeek-V3 mit außergewöhnlichem Reasoning auf Clore.ai-GPUs aus
- [DeepSeek-R1-Reasoning-Modell](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/deepseek-r1.md): Führe das Open-Source-Reasoning-Modell DeepSeek-R1 auf Clore.ai-GPUs aus
- [Qwen2.5](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/qwen25.md): Führe Alibabas Qwen2.5-Mehrsprach-LLMs auf Clore.ai-GPUs aus
- [CodeLlama](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/codellama.md): Code mit CodeLlama generieren, vervollständigen und erklären auf Clore.ai
- [Gemma 2](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/gemma2.md): Führe Googles Gemma-2-Modelle effizient auf Clore.ai-GPUs aus
- [Phi-4](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/phi4.md): Führe Microsofts kleines Sprachmodell Phi-4 auf Clore.ai-GPUs aus
- [Llama 4 (Scout & Maverick)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/llama4.md): Führe Metas Llama-4-Scout-&-Maverick-MoE-Modelle auf Clore.ai-GPUs aus
- [Gemma 3](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/gemma3.md): Führe Googles multimodale Gemma-3-Modelle auf Clore.ai aus — schlägt Llama-405B bei 15x kleinerer Größe
- [Gemma 4 (26B MoE, 4B aktiv)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/gemma4.md): Deploye Gemma 4 (26B MoE, 4B aktiv) von Google auf Clore.ai — das Open-Weight-Modell, veröffentlicht im April 2026, das aufstieg zu
- [Mistral Small 3.1](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-small.md): Deploye Mistral Small 3.1 (24B) auf Clore.ai — das ideale Single-GPU-Produktionsmodell
- [Qwen3.5](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/qwen35.md): Führe Alibaba Qwen3.5 auf Clore.ai aus — das frischeste Frontier-Modell (Feb. 2026)
- [Qwen3.5-Omni (multimodal)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/qwen35-omni.md)
- [GLM-5](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/glm5.md): Deploye GLM-5 (744B MoE) von Zhipu AI auf Clore.ai — API-Zugriff und Self-Hosting mit vLLM
- [GLM-4.7-Flash](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/glm-47-flash.md): Deploye GLM-4.7-Flash (30B MoE) von Zhipu AI auf Clore.ai — effizientes Sprachmodell mit 59,2 % SWE-bench-Performance
- [Kimi K2.5](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/kimi-k2.md): Deploye Kimi K2.5 (1T MoE multimodal) von Moonshot AI auf Clore.ai-GPUs
- [Mistral Large 3 (675B MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-large3.md): Führe Mistral Large 3 aus — ein 675B MoE-Frontier-Modell mit 41B aktiven Parametern auf Clore.ai-GPUs
- [Mistral Medium 3.5 (128B dicht, 256K)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-medium35.md): Deploye Mistral Medium 3.5 auf Clore.ai — 128B dicht, 256K Kontext, dual-modales Reasoning, veröffentlicht im April 2026. Produktionsreife vLLM/SGLang-Einrichtung auf 4× H100 oder 2× H200.
- [MiMo-V2-Flash](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mimo-v2-flash.md): Deploye MiMo-V2-Flash (309B MoE) mit spekulativer Decodierung auf Clore.ai — extrem schnelle Inferenz mit 150+ Token/s
- [Ling-2.5-1T (1 Billion Parameter)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/ling25.md): Führe Ling-2.5-1T aus — das Open-Source-LLM mit 1 Billion Parametern von Ant Group mit hybrider linearer Attention auf Clore.ai-GPUs
- [LFM2-24B-A2B](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/lfm2-24b.md): Deploye LFM2-24B-A2B von Liquid AI auf Clore.ai — hybride SSM+Attention-Architektur mit 24B Gesamt-/2B aktiven Parametern
- [DeepSeek V4 (1.6T MoE, multimodal)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/deepseek-v4.md): Deploye DeepSeek V4 auf Clore.ai — das MIT-lizenzierte Frontier-MoE, aktualisiert als Flash-0731 und Pro-0813
- [GLM-5.1 (744B MoE, #1 SWE-Bench Pro)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/glm-5-1.md): Deploye GLM-5.1 (744B MoE, 40B aktiv) von Z.ai auf Clore.ai — das Open-Weight-Modell, das im April 2026 SWE-Bench Pro anführte
- [NVIDIA Nemotron 3 Super (120B MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/nvidia-nemotron-3-super.md)
- [Gemini 3.1 Flash Lite](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/gemini-3-1-flash-lite.md)
- [Hy3-Vorschau (Tencent Hunyuan 3, 295B MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/hy3-preview.md): Deploye Tencents Hy3-Vorschau (295B MoE, 21B aktiv, 256K ctx) auf Clore.ai — das erste Modell aus Tencents neu aufgebautem Hunyuan-Trainingsstack, abgestimmt auf langfristiges Reasoning und agentisches Programmieren
- [MiMo-V2.5-Pro (Xiaomi 1T MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mimo-v25-pro.md): Deploye MiMo-V2.5-Pro (1,02T MoE, 42B aktiv, 1M Kontext) von Xiaomi auf Clore.ai — die erste Open-Weight-Pro-Stufe vom MiMo-Team, FP8-nativ, hybride Attention
- [MiniMax M2.7 (229B MoE Coding)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/minimax-m27.md): Deploye MiniMax M2.7 (229B MoE) auf Clore.ai — die Open-Weight-Self-Hosted-Version hinter MiniMax' Vorstoß mit dem Coding-Agenten, mit FP8-Single-Node-Deployment auf H100/H200
- [Ling-2.6-flash (Ant Group 104B MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/ling-26-flash.md): Deploye Ling-2.6-flash (104B MoE, 7,4B aktiv) von Ant Group auf Clore.ai — das auf Agenten abgestimmte Flash-Geschwistermodell, das auf eine einzelne RTX 4090 passt
- [Qwen3.6-27B (dicht, Single-GPU)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/qwen36-27b.md): Deploye Qwen3.6-27B von Alibaba auf Clore.ai — ein dichtes 27B-Modell, das auf eine einzelne RTX 4090 passt und mit 262K nativer Kontextlänge geliefert wird
- [Qwen3.8-27B (Dense VLM, eine Karte)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/qwen38-27b.md): Deploye Qwen3.8-27B auf Clore.ai — Alibabas dichtes 27B Vision-Language-Modell, das auf einer einzelnen RTX 4090 mit Q4 und einer einzelnen RTX 5090 mit FP8 läuft
- [GLM-5.2 (MIT, 1M Kontext)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/glm-5-2.md): Führe GLM-5.2 aus, das MIT-lizenzierte 1M-Kontext-Coding-Flaggschiff von Z.ai, auf einem großen Multi-GPU-Rig aus dem Clore.ai-Marktplatz aus
- [MiniMax M3 (428B multimodal)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/minimax-m3.md): Deploye MiniMax M3, ein nativ multimodales 428B MoE mit 1M-Token-Kontext, auf Multi-GPU-Rigs von Clore.ai
- [NVIDIA Nemotron 3 Ultra (550B Mamba-MoE)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/nemotron-3-ultra.md): Führe NVIDIA Nemotron 3 Ultra, den offenen 550B Mamba-MoE-Hybrid, auf Blackwell-Rigs aus dem Clore.ai-Marktplatz aus
- [Mistral Small 4 (119B MoE, 6,5B aktiv)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-small4.md): Deploye Mistral Small 4 (119B MoE, 6,5B aktiv) auf zwei Consumer-GPUs aus dem Clore.ai-Marktplatz — Apache 2.0, 256K Kontext, Vision und Reasoning in einem Modell
- [Kimi K3 (2,8T — Was es braucht)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/kimi-k3.md): Was Kimi K3 ist, was es wirklich braucht, um 2,8 Billionen Parameter selbst zu hosten, und was stattdessen auf Clore.ai ausgeführt werden sollte
- [TGI (Text Generation Inference)](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/tgi.md): Führe HuggingFace Text Generation Inference (TGI) für den produktiven LLM-Betrieb auf Clore.ai-GPUs aus
- [SGLang](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/sglang.md): Deploye SGLang für leistungsstarken LLM-Betrieb mit RadixAttention auf Clore.ai-GPUs
- [Aphrodite Engine](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/aphrodite-engine.md): Führe Aphrodite Engine für LLM-Inferenz auf älteren und modernen GPUs auf Clore.ai aus
- [LiteLLM KI-Gateway](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/litellm.md): Deploye LiteLLM als KI-Gateway-Proxy für 100+ LLMs auf Clore.ai-GPUs
- [MLC-LLM](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mlc-llm.md)
- [PowerInfer](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/powerinfer.md)
- [LMDeploy](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/lmdeploy.md)
- [Mistral.rs](https://docs.clore.ai/guides/guides_v2-de/sprachmodelle/mistral-rs.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-de/sprachmodelle.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
