> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-hi/language-models.md).

# भाषा मॉडल

- [अवलोकन](https://docs.clore.ai/guides/guides_v2-hi/language-models/language-models.md)
- [Ollama](https://docs.clore.ai/guides/guides_v2-hi/language-models/ollama.md): Clore.ai GPUs पर Ollama के साथ स्थानीय रूप से LLMs चलाएँ
- [Open WebUI](https://docs.clore.ai/guides/guides_v2-hi/language-models/open-webui.md): Clore.ai GPUs पर LLMs चलाने के लिए ChatGPT जैसी interface
- [vLLM](https://docs.clore.ai/guides/guides_v2-hi/language-models/vllm.md): Clore.ai GPUs पर vLLM के साथ high-throughput LLM inference
- [Llama.cpp Server](https://docs.clore.ai/guides/guides_v2-hi/language-models/llamacpp-server.md): Clore.ai GPUs पर llama.cpp server के साथ कुशल LLM inference
- [Text Generation WebUI](https://docs.clore.ai/guides/guides_v2-hi/language-models/text-generation-webui.md): Clore.ai GPUs पर LLM inference के लिए text-generation-webui चलाएँ
- [ExLlamaV2](https://docs.clore.ai/guides/guides_v2-hi/language-models/exllamav2-fast.md): Clore.ai GPUs पर ExLlamaV2 के साथ अधिकतम गति वाला LLM inference
- [LocalAI](https://docs.clore.ai/guides/guides_v2-hi/language-models/localai-openai-compatible.md): Clore.ai पर LocalAI के साथ self-hosted OpenAI-compatible API
- [Llama 3.3 70B](https://docs.clore.ai/guides/guides_v2-hi/language-models/llama33.md): Clore.ai GPUs पर Meta के Llama 3.3 70B model चलाएँ
- [Mistral और Mixtral](https://docs.clore.ai/guides/guides_v2-hi/language-models/mistral-mixtral.md): Clore.ai GPUs पर Mistral और Mixtral models चलाएँ
- [DeepSeek Coder](https://docs.clore.ai/guides/guides_v2-hi/language-models/deepseek-coder.md): Clore.ai पर DeepSeek Coder के साथ class-श्रेष्ठ code generation
- [DeepSeek-V3](https://docs.clore.ai/guides/guides_v2-hi/language-models/deepseek-v3.md): असाधारण reasoning के साथ DeepSeek-V3 को Clore.ai GPUs पर चलाएँ
- [DeepSeek-R1 Reasoning Model](https://docs.clore.ai/guides/guides_v2-hi/language-models/deepseek-r1.md): Clore.ai GPUs पर DeepSeek-R1 open-source reasoning model चलाएँ
- [Qwen2.5](https://docs.clore.ai/guides/guides_v2-hi/language-models/qwen25.md): Alibaba के Qwen2.5 multilingual LLMs को Clore.ai GPUs पर चलाएँ
- [CodeLlama](https://docs.clore.ai/guides/guides_v2-hi/language-models/codellama.md): Clore.ai पर CodeLlama के साथ code generate, complete, और explain करें
- [Gemma 2](https://docs.clore.ai/guides/guides_v2-hi/language-models/gemma2.md): Clore.ai GPUs पर Google के Gemma 2 models कुशलतापूर्वक चलाएँ
- [Phi-4](https://docs.clore.ai/guides/guides_v2-hi/language-models/phi4.md): Microsoft के Phi-4 छोटे भाषा मॉडल को Clore.ai GPUs पर चलाएँ
- [Llama 4 (Scout & Maverick)](https://docs.clore.ai/guides/guides_v2-hi/language-models/llama4.md): Meta Llama 4 Scout & Maverick MoE models को Clore.ai GPUs पर चलाएँ
- [Gemma 3](https://docs.clore.ai/guides/guides_v2-hi/language-models/gemma3.md): Google Gemma 3 multimodal models को Clore.ai पर चलाएँ — 15x छोटे आकार में Llama-405B को पछाड़ता है
- [Gemma 4 (26B MoE, 4B सक्रिय)](https://docs.clore.ai/guides/guides_v2-hi/language-models/gemma4.md): Gemma 4 (26B MoE, 4B सक्रिय) को Google द्वारा Clore.ai पर डिप्लॉय करें — अप्रैल 2026 में जारी किया गया ओपन-वेट मॉडल जो पहुँच गया
- [Mistral Small 3.1](https://docs.clore.ai/guides/guides_v2-hi/language-models/mistral-small.md): Mistral Small 3.1 (24B) को Clore.ai पर डिप्लॉय करें — आदर्श single-GPU production model
- [Qwen3.5](https://docs.clore.ai/guides/guides_v2-hi/language-models/qwen35.md): Alibaba Qwen3.5 को Clore.ai पर चलाएँ — सबसे ताज़ा frontier model (फ़रवरी 2026)
- [Qwen3.5-Omni (मल्टीमोडल)](https://docs.clore.ai/guides/guides_v2-hi/language-models/qwen35-omni.md)
- [GLM-5](https://docs.clore.ai/guides/guides_v2-hi/language-models/glm5.md): GLM-5 (744B MoE) को Zhipu AI द्वारा Clore.ai पर डिप्लॉय करें — vLLM के साथ API access और self-hosting
- [GLM-4.7-Flash](https://docs.clore.ai/guides/guides_v2-hi/language-models/glm-47-flash.md): GLM-4.7-Flash (30B MoE) को Zhipu AI द्वारा Clore.ai पर डिप्लॉय करें — 59.2% SWE-bench प्रदर्शन वाला कुशल भाषा मॉडल
- [Kimi K2.5](https://docs.clore.ai/guides/guides_v2-hi/language-models/kimi-k2.md): Kimi K2.5 (1T MoE multimodal) को Moonshot AI द्वारा Clore.ai GPUs पर डिप्लॉय करें
- [Mistral Large 3 (675B MoE)](https://docs.clore.ai/guides/guides_v2-hi/language-models/mistral-large3.md): Mistral Large 3 चलाएँ — 41B active parameters वाला 675B MoE frontier model Clore.ai GPUs पर
- [Mistral Medium 3.5 (128B Dense, 256K)](https://docs.clore.ai/guides/guides_v2-hi/language-models/mistral-medium35.md): Mistral Medium 3.5 को Clore.ai पर डिप्लॉय करें — 128B dense, 256K context, dual-mode reasoning, अप्रैल 2026 में जारी। 4× H100 या 2× H200 पर production vLLM/SGLang सेटअप।
- [MiMo-V2-Flash](https://docs.clore.ai/guides/guides_v2-hi/language-models/mimo-v2-flash.md): MiMo-V2-Flash (309B MoE) को speculative decoding के साथ Clore.ai पर डिप्लॉय करें — 150+ tok/s के साथ ultra-fast inference
- [Ling-2.5-1T (1 ट्रिलियन पैरामीटर)](https://docs.clore.ai/guides/guides_v2-hi/language-models/ling25.md): Ling-2.5-1T चलाएँ — hybrid linear attention के साथ Ant Group का 1 ट्रिलियन parameter open-source LLM Clore.ai GPUs पर
- [LFM2-24B-A2B](https://docs.clore.ai/guides/guides_v2-hi/language-models/lfm2-24b.md): LFM2-24B-A2B को Liquid AI द्वारा Clore.ai पर डिप्लॉय करें — hybrid SSM+Attention architecture, 24B total / 2B active parameters के साथ
- [DeepSeek V4 (1.6T MoE, मल्टीमोडल)](https://docs.clore.ai/guides/guides_v2-hi/language-models/deepseek-v4.md): DeepSeek V4 (1.6T-param Pro और 284B Flash) को Clore.ai पर डिप्लॉय करें — अप्रैल 22, 2026 को जारी open-weight frontier MoE
- [GLM-5.1 (744B MoE, #1 SWE-Bench Pro)](https://docs.clore.ai/guides/guides_v2-hi/language-models/glm-5-1.md): GLM-5.1 (744B MoE, 40B सक्रिय) को Z.ai द्वारा Clore.ai पर डिप्लॉय करें — अप्रैल 2026 में SWE-Bench Pro में शीर्ष पर रहने वाला open-weight मॉडल
- [NVIDIA Nemotron 3 Super (120B MoE)](https://docs.clore.ai/guides/guides_v2-hi/language-models/nvidia-nemotron-3-super.md)
- [Gemini 3.1 Flash Lite](https://docs.clore.ai/guides/guides_v2-hi/language-models/gemini-3-1-flash-lite.md)
- [Hy3 Preview (Tencent Hunyuan 3, 295B MoE)](https://docs.clore.ai/guides/guides_v2-hi/language-models/hy3-preview.md): Tencent के Hy3 Preview (295B MoE, 21B सक्रिय, 256K ctx) को Clore.ai पर डिप्लॉय करें — Tencent Hunyuan के पुनर्निर्मित ट्रेनिंग स्टैक का पहला मॉडल, long-horizon reasoning और agentic coding के लिए ट्यून किया गया
- [MiMo-V2.5-Pro (Xiaomi 1T MoE)](https://docs.clore.ai/guides/guides_v2-hi/language-models/mimo-v25-pro.md): MiMo-V2.5-Pro (1.02T MoE, 42B सक्रिय, 1M context) को Xiaomi द्वारा Clore.ai पर डिप्लॉय करें — MiMo टीम का पहला open-weight Pro tier, FP8 native, hybrid attention
- [MiniMax M2.7 (229B MoE Coding)](https://docs.clore.ai/guides/guides_v2-hi/language-models/minimax-m27.md): MiniMax M2.7 (229B MoE) को Clore.ai पर डिप्लॉय करें — MiniMax के coding agent push के पीछे का open-weight self-hosted रिलीज़, H100/H200 पर FP8 single-node deployment के साथ
- [Ling-2.6-flash (Ant Group 104B MoE)](https://docs.clore.ai/guides/guides_v2-hi/language-models/ling-26-flash.md): Ling-2.6-flash (104B MoE, 7.4B सक्रिय) को Ant Group द्वारा Clore.ai पर डिप्लॉय करें — agent-tuned flash sibling जो एक सिंगल RTX 4090 पर फिट बैठता है
- [Qwen3.6-27B (Dense, Single-GPU)](https://docs.clore.ai/guides/guides_v2-hi/language-models/qwen36-27b.md): Qwen3.6-27B को Alibaba द्वारा Clore.ai पर डिप्लॉय करें — एक dense 27B जो एक RTX 4090 पर फिट होता है और 262K native context के साथ आता है
- [TGI (Text Generation Inference)](https://docs.clore.ai/guides/guides_v2-hi/language-models/tgi.md): Clore.ai GPUs पर production LLM serving के लिए HuggingFace Text Generation Inference (TGI) चलाएँ
- [SGLang](https://docs.clore.ai/guides/guides_v2-hi/language-models/sglang.md): Clore.ai GPUs पर RadixAttention के साथ high-performance LLM serving के लिए SGLang डिप्लॉय करें
- [Aphrodite Engine](https://docs.clore.ai/guides/guides_v2-hi/language-models/aphrodite-engine.md): Clore.ai पर legacy और modern GPUs पर LLM inference के लिए Aphrodite Engine चलाएँ
- [LiteLLM AI Gateway](https://docs.clore.ai/guides/guides_v2-hi/language-models/litellm.md): Clore.ai GPUs पर 100+ LLMs के लिए AI Gateway proxy के रूप में LiteLLM डिप्लॉय करें
- [MLC-LLM](https://docs.clore.ai/guides/guides_v2-hi/language-models/mlc-llm.md)
- [PowerInfer](https://docs.clore.ai/guides/guides_v2-hi/language-models/powerinfer.md)
- [LMDeploy](https://docs.clore.ai/guides/guides_v2-hi/language-models/lmdeploy.md)
- [Mistral.rs](https://docs.clore.ai/guides/guides_v2-hi/language-models/mistral-rs.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-hi/language-models.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
