> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing.md).

# 语言模型

- [概览](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/language-models.md)
- [Ollama](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/ollama.md): 使用 Ollama 在 Clore.ai GPU 上本地运行 LLM
- [Open WebUI](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/open-webui.md): 在 Clore.ai GPU 上运行 LLM 的 ChatGPT 式界面
- [vLLM](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/vllm.md): 在 Clore.ai GPU 上使用 vLLM 进行高吞吐量 LLM 推理
- [Llama.cpp 服务器](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/llamacpp-server.md): 在 Clore.ai GPU 上使用 llama.cpp server 高效推理 LLM
- [Text Generation WebUI](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/text-generation-webui.md): 在 Clore.ai GPU 上运行 text-generation-webui 进行 LLM 推理
- [ExLlamaV2](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/exllamav2-fast.md): 在 Clore.ai GPU 上使用 ExLlamaV2 实现最快速度的 LLM 推理
- [LocalAI](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/localai-openai-compatible.md): 在 Clore.ai 上使用 LocalAI 自托管兼容 OpenAI 的 API
- [Llama 3.3 70B](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/llama33.md): 在 Clore.ai GPU 上运行 Meta 的 Llama 3.3 70B 模型
- [Mistral 与 Mixtral](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-mixtral.md): 在 Clore.ai GPU 上运行 Mistral 和 Mixtral 模型
- [DeepSeek Coder](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/deepseek-coder.md): 在 Clore.ai 上使用 DeepSeek Coder 实现一流代码生成
- [DeepSeek-V3](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/deepseek-v3.md): 在 Clore.ai GPU 上运行具有出色推理能力的 DeepSeek-V3
- [DeepSeek-R1 推理模型](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/deepseek-r1.md): 在 Clore.ai GPU 上运行 DeepSeek-R1 开源推理模型
- [Qwen2.5](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/qwen25.md): 在 Clore.ai GPU 上运行阿里巴巴的 Qwen2.5 多语言 LLM
- [CodeLlama](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/codellama.md): 在 Clore.ai 上使用 CodeLlama 生成、补全并解释代码
- [Gemma 2](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/gemma2.md): 在 Clore.ai GPU 上高效运行 Google 的 Gemma 2 模型
- [Phi-4](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/phi4.md): 在 Clore.ai GPU 上运行微软的 Phi-4 小型语言模型
- [Llama 4（Scout 与 Maverick）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/llama4.md): 在 Clore.ai GPU 上运行 Meta Llama 4 Scout 与 Maverick MoE 模型
- [Gemma 3](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/gemma3.md): 在 Clore.ai 上运行 Google Gemma 3 多模态模型——比 Llama-405B 小 15 倍却表现更强
- [Gemma 4（26B MoE，4B 活跃）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/gemma4.md): 在 Clore.ai 上部署 Google 的 Gemma 4（26B MoE，4B 活跃）——2026 年 4 月发布的开源权重模型，迅速攀升至
- [Mistral Small 3.1](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-small.md): 在 Clore.ai 上部署 Mistral Small 3.1（24B）——理想的单 GPU 生产级模型
- [Qwen3.5](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/qwen35.md): 在 Clore.ai 上运行阿里巴巴 Qwen3.5——最新前沿模型（2026 年 2 月）
- [Qwen3.5-Omni（多模态）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/qwen35-omni.md)
- [GLM-5](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/glm5.md): 在 Clore.ai 上部署 Zhipu AI 的 GLM-5（744B MoE）——通过 vLLM 提供 API 访问和自托管
- [GLM-4.7-Flash](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/glm-47-flash.md): 在 Clore.ai 上部署 Zhipu AI 的 GLM-4.7-Flash（30B MoE）——高效语言模型，SWE-bench 表现达 59.2%
- [Kimi K2.5](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/kimi-k2.md): 在 Clore.ai GPU 上部署 Moonshot AI 的 Kimi K2.5（1T MoE 多模态）
- [Mistral Large 3（675B MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-large3.md): 在 Clore.ai GPU 上运行 Mistral Large 3——一款拥有 410 亿活跃参数的 675B MoE 前沿模型
- [Mistral Medium 3.5（128B 稠密，256K）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-medium35.md): 在 Clore.ai 上部署 Mistral Medium 3.5——128B 稠密、256K 上下文、双模式推理，2026 年 4 月发布。在 4× H100 或 2× H200 上的生产级 vLLM/SGLang 配置。
- [MiMo-V2-Flash](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mimo-v2-flash.md): 在 Clore.ai 上使用投机解码部署 MiMo-V2-Flash（309B MoE）——超快推理，速度达 150+ tok/s
- [Ling-2.5-1T（1 万亿参数）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/ling25.md): 在 Clore.ai GPU 上运行 Ling-2.5-1T——蚂蚁集团的 1 万亿参数开源 LLM，采用混合线性注意力
- [LFM2-24B-A2B](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/lfm2-24b.md): 在 Clore.ai 上部署 Liquid AI 的 LFM2-24B-A2B——混合 SSM+Attention 架构，总参数 24B / 活跃参数 2B
- [DeepSeek V4（1.6T MoE，多模态）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/deepseek-v4.md): 在 Clore.ai 上部署 DeepSeek V4——MIT 许可的前沿 MoE，更新为 Flash-0731 和 Pro-0813
- [GLM-5.1（744B MoE，SWE-Bench Pro 第一）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/glm-5-1.md): 在 Clore.ai 上部署 Z.ai 的 GLM-5.1（744B MoE，40B 活跃）——2026 年 4 月登顶 SWE-Bench Pro 的开源权重模型
- [NVIDIA Nemotron 3 Super（120B MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/nvidia-nemotron-3-super.md)
- [Gemini 3.1 Flash Lite](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/gemini-3-1-flash-lite.md)
- [Hy3 预览版（腾讯混元 3，295B MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/hy3-preview.md): 在 Clore.ai 上部署腾讯的 Hy3 预览版（295B MoE，21B 活跃，256K ctx）——腾讯混元重建训练栈的首个模型，针对长程推理和智能体编码进行了优化
- [MiMo-V2.5-Pro（小米 1T MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mimo-v25-pro.md): 在 Clore.ai 上部署小米的 MiMo-V2.5-Pro（1.02T MoE，42B 活跃，100 万上下文）——MiMo 团队首个开源权重 Pro 级模型，原生 FP8，混合注意力
- [MiniMax M2.7（229B MoE 编码）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/minimax-m27.md): 在 Clore.ai 上部署 MiniMax M2.7（229B MoE）——MiniMax 编程智能体推进背后的开源自托管版本，支持在 H100/H200 上进行 FP8 单节点部署
- [Ling-2.6-flash（蚂蚁集团 104B MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/ling-26-flash.md): 在 Clore.ai 上部署蚂蚁集团的 Ling-2.6-flash（104B MoE，7.4B 活跃）——面向智能体优化的 flash 兄弟模型，可运行在单张 RTX 4090 上
- [Qwen3.6-27B（稠密，单 GPU）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/qwen36-27b.md): 在 Clore.ai 上部署阿里巴巴的 Qwen3.6-27B——一款可运行于单张 RTX 4090 的稠密 27B 模型，附带原生 262K 上下文
- [Qwen3.8-27B（稠密 VLM，单卡）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/qwen38-27b.md): 在 Clore.ai 上部署 Qwen3.8-27B——阿里巴巴的稠密 27B 视觉语言模型，在 Q4 下可于单张 RTX 4090 上运行，在 FP8 下可于单张 RTX 5090 上运行
- [GLM-5.2（MIT，100 万上下文）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/glm-5-2.md): 在 Clore.ai 市场上的大型多 GPU 机器上运行 GLM-5.2，Z.ai 的 MIT 许可、100 万上下文的代码旗舰模型
- [MiniMax M3（428B 多模态）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/minimax-m3.md): 在 Clore.ai 的多 GPU 机器上部署 MiniMax M3，这是一款原生多模态的 428B MoE，拥有 100 万 token 上下文
- [NVIDIA Nemotron 3 Ultra（550B Mamba-MoE）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/nemotron-3-ultra.md): 在 Clore.ai 市场的 Blackwell 机器上运行 NVIDIA Nemotron 3 Ultra，这款 550B 开源 Mamba-MoE 混合模型
- [Mistral Small 4（119B MoE，6.5B 活跃）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-small4.md): 在 Clore.ai 市场上的两块消费级 GPU 上部署 Mistral Small 4（119B MoE，6.5B 活跃）——Apache 2.0 许可、256K 上下文、视觉与推理合一的模型
- [Kimi K3（2.8T——所需条件）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/kimi-k3.md): Kimi K3 是什么、要自托管 2.8 万亿参数实际需要什么，以及在 Clore.ai 上该运行什么来替代
- [TGI（文本生成推理）](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/tgi.md): 在 Clore.ai GPU 上运行 HuggingFace Text Generation Inference（TGI）用于生产级 LLM 服务
- [SGLang](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/sglang.md): 在 Clore.ai GPU 上部署 SGLang，借助 RadixAttention 实现高性能 LLM 服务
- [Aphrodite 引擎](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/aphrodite-engine.md): 在 Clore.ai 上使用 Aphrodite Engine 在旧款和现代 GPU 上进行 LLM 推理
- [LiteLLM AI 网关](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/litellm.md): 在 Clore.ai GPU 上将 LiteLLM 部署为支持 100+ LLM 的 AI 网关代理
- [MLC-LLM](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mlc-llm.md)
- [PowerInfer](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/powerinfer.md)
- [LMDeploy](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/lmdeploy.md)
- [Mistral.rs](https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing/mistral-rs.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-zh/yu-yan-mo-xing.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
