> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-fr/modeles-de-langage/language-models.md).

# Aperçu

Exécutez de grands modèles de langage (LLM) sur les GPU CLORE.AI pour l'inférence et les applications de chat.

## Outils populaires

| Outil                                                                                     | Cas d’utilisation                                     | Difficulté |
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------- | ---------- |
| [Ollama](/guides/guides_v2-fr/modeles-de-langage/ollama.md)                               | Configuration LLM la plus simple                      | Débutant   |
| [Open WebUI](/guides/guides_v2-fr/modeles-de-langage/open-webui.md)                       | Interface de type ChatGPT                             | Débutant   |
| [vLLM](/guides/guides_v2-fr/modeles-de-langage/vllm.md)                                   | Service de production à haut débit                    | Moyen      |
| [Llama.cpp Server](/guides/guides_v2-fr/modeles-de-langage/llamacpp-server.md)            | Inférence GGUF efficace                               | Facile     |
| [Text Generation WebUI](/guides/guides_v2-fr/modeles-de-langage/text-generation-webui.md) | Interface de chat complète                            | Facile     |
| [ExLlamaV2](/guides/guides_v2-fr/modeles-de-langage/exllamav2-fast.md)                    | Inférence EXL2 la plus rapide                         | Moyen      |
| [LocalAI](/guides/guides_v2-fr/modeles-de-langage/localai-openai-compatible.md)           | API compatible avec OpenAI                            | Moyen      |
| [SGLang](/guides/guides_v2-fr/modeles-de-langage/sglang.md)                               | Génération structurée rapide                          | Moyen      |
| [Text Generation Inference (TGI)](/guides/guides_v2-fr/modeles-de-langage/tgi.md)         | Solution de service HuggingFace                       | Moyen      |
| [LMDeploy](/guides/guides_v2-fr/modeles-de-langage/lmdeploy.md)                           | Boîte à outils de service MMlab                       | Moyen      |
| [Aphrodite Engine](/guides/guides_v2-fr/modeles-de-langage/aphrodite-engine.md)           | Fork de vLLM avec des fonctionnalités supplémentaires | Moyen      |
| [MLC-LLM](/guides/guides_v2-fr/modeles-de-langage/mlc-llm.md)                             | Compilation pour l'apprentissage automatique          | Difficile  |
| [LiteLLM](/guides/guides_v2-fr/modeles-de-langage/litellm.md)                             | Proxy d'API unifiée                                   | Moyen      |
| [PowerInfer](/guides/guides_v2-fr/modeles-de-langage/powerinfer.md)                       | Inférence de modèle creux                             | Difficile  |
| [Mistral.rs](/guides/guides_v2-fr/modeles-de-langage/mistral-rs.md)                       | Moteur d'inférence basé sur Rust                      | Moyen      |

## Guides des modèles

### Modèles les plus récents et les meilleurs

| Modèle                                                                          | Paramètres              | Licence           | Tient sur                   |
| ------------------------------------------------------------------------------- | ----------------------- | ----------------- | --------------------------- |
| [Qwen3.8-27B](/guides/guides_v2-fr/modeles-de-langage/qwen38-27b.md)            | 27B dense               | Apache 2.0        | 1× RTX 4090 (Q4)            |
| [Mistral Small 4](/guides/guides_v2-fr/modeles-de-langage/mistral-small4.md)    | 119B / 6.5B actifs      | Apache 2.0        | 2× RTX 4090 (Q2)            |
| [GLM-5.2](/guides/guides_v2-fr/modeles-de-langage/glm-5-2.md)                   | \~750B / 40B actifs     | MIT               | 10× RTX 5090 (Q2)           |
| [DeepSeek V4](/guides/guides_v2-fr/modeles-de-langage/deepseek-v4.md)           | Flash 167GB · Pro 893GB | MIT               | 6–8× RTX 5090 (Flash)       |
| [MiniMax M3](/guides/guides_v2-fr/modeles-de-langage/minimax-m3.md)             | 428B / 23B actifs       | minimax-community | 6–8× RTX 5090 (Q2–Q3)       |
| [Nemotron 3 Ultra](/guides/guides_v2-fr/modeles-de-langage/nemotron-3-ultra.md) | 550B / 55B actifs       | OpenMDW-1.1       | 4× RTX PRO 6000 (NVFP4)     |
| [Kimi K3](/guides/guides_v2-fr/modeles-de-langage/kimi-k3.md)                   | 2.8T                    | Licence Kimi K3   | Rien sur la place de marché |
| [GLM-5.1](/guides/guides_v2-fr/modeles-de-langage/glm-5-1.md)                   | 744B / 40B actifs       | MIT               | Grosses configurations      |
| [Llama 4](/guides/guides_v2-fr/modeles-de-langage/llama4.md)                    | Scout et Maverick       | Communauté Llama  | 1× RTX 4090 (Scout Q4)      |
| [DeepSeek-R1](/guides/guides_v2-fr/modeles-de-langage/deepseek-r1.md)           | 671B MoE                | MIT               | Grosses configurations      |
| [Qwen2.5](/guides/guides_v2-fr/modeles-de-langage/qwen25.md)                    | 0.5B–72B                | Apache 2.0        | N’importe quelle carte      |

### Modèles spécialisés

| Modèle                                                                      | Paramètres | Idéal pour                 |
| --------------------------------------------------------------------------- | ---------- | -------------------------- |
| [DeepSeek Coder](/guides/guides_v2-fr/modeles-de-langage/deepseek-coder.md) | 6.7B-33B   | Génération de code         |
| [CodeLlama](/guides/guides_v2-fr/modeles-de-langage/codellama.md)           | 7B-34B     | Complétion de code         |
| [GLM-4.7-Flash](/guides/guides_v2-fr/modeles-de-langage/glm-47-flash.md)    | 4.7B       | Rapide en chinois/anglais  |
| [GLM-5](/guides/guides_v2-fr/modeles-de-langage/glm5.md)                    | À venir    | Dernier modèle de Zhipu AI |
| [Kimi K2.5](/guides/guides_v2-fr/modeles-de-langage/kimi-k2.md)             | À venir    | Modèle de Moonshot AI      |
| [Ling-2.5-1T](/guides/guides_v2-fr/modeles-de-langage/ling25.md)            | 1T         | LLM open source massif     |
| [LFM2-24B](/guides/guides_v2-fr/modeles-de-langage/lfm2-24b.md)             | 24B        | Modèle Liquid AI           |
| [MiMo-V2-Flash](/guides/guides_v2-fr/modeles-de-langage/mimo-v2-flash.md)   | À venir    | Modèle d’inférence rapide  |

### Modèles efficaces

| Modèle                                                                        | Paramètres | Idéal pour                       |
| ----------------------------------------------------------------------------- | ---------- | -------------------------------- |
| [Gemma 2](/guides/guides_v2-fr/modeles-de-langage/gemma2.md)                  | 2B-27B     | Inférence efficace               |
| [Gemma 3](/guides/guides_v2-fr/modeles-de-langage/gemma3.md)                  | À venir    | Dernier modèle compact de Google |
| [Phi-4](/guides/guides_v2-fr/modeles-de-langage/phi4.md)                      | 14B        | Petit mais performant            |
| [Mistral/Mixtral](/guides/guides_v2-fr/modeles-de-langage/mistral-mixtral.md) | 7B / 8x7B  | Usage général                    |
| [Mistral Large 3](/guides/guides_v2-fr/modeles-de-langage/mistral-large3.md)  | MoE 675B   | De niveau entreprise             |
| [Mistral Small 3.1](/guides/guides_v2-fr/modeles-de-langage/mistral-small.md) | À venir    | Variante Mistral efficace        |

## Recommandations GPU

{% hint style="warning" %}
**Les configurations multi-GPU de classe 80 Go ne sont pas répertoriées sur la place de marché Clore.ai.** Les plus grosses machines répertoriées aujourd’hui sont 4× RTX PRO 6000 Blackwell (96 Go chacune, 380 Go au total) et 8–11× RTX 5090 (32 Go chacune). La capacité A100 / H200 / B200 est proposée en bare metal sur demande. Consultez [bare metal](https://clore.ai/bare-metal) sur demande. Consultez [Tarifs et disponibilité des GPU](/guides/guides_v2-fr/premiers-pas/pricing.md) avant de dimensionner un déploiement.
{% endhint %}

| Taille du modèle | GPU minimum    | Recommandé           | Coût Clore.ai   |
| ---------------- | -------------- | -------------------- | --------------- |
| 7B (Q4)          | RTX 3060 12 Go | RTX 3090             | 0,03–0,21 $/h   |
| 13B (Q4)         | RTX 3090 24 Go | RTX 4090             | 0,07–0,42 $/h   |
| 27B (Q4)         | RTX 3090 24 Go | RTX 4090 / 5090      | 0,07–0,77 $/h   |
| 34B (Q4)         | 2× RTX 3090    | 1× RTX 5090          | 0,14–0,77 $/h   |
| 70B (Q4)         | 2× RTX 5090    | 1× RTX PRO 6000 96GB | 0,50–1,54 $/h   |
| 120B+ (Q2)       | 2× RTX 4090    | 2× RTX 5090          | 0,28–1,54 $/h   |
| 400B+ (Q2)       | 8× RTX 5090    | 4× RTX PRO 6000      | \~2,00–5,00 $/h |

## Guide de quantification

| Format   | Utilisation de la VRAM | Qualité    | Vitesse        |
| -------- | ---------------------- | ---------- | -------------- |
| Q2\_K    | La plus faible         | Médiocre   | La plus rapide |
| Q4\_K\_M | Faible                 | Bon        | Rapide         |
| Q5\_K\_M | Moyen                  | Excellent  | Moyen          |
| Q8\_0    | Élevée                 | Excellente | Plus lent      |
| FP16     | La plus élevée         | Meilleur   | Le plus lent   |

## Voir aussi

* [Entraînement et fine-tuning](/guides/guides_v2-fr/entrainement/training.md)
* [Modèles vision-langage](/guides/guides_v2-fr/modeles-de-vision/vision-models.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-fr/modeles-de-langage/language-models.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
