# Overview

**Clore.ai** is a decentralized GPU marketplace where anyone can rent high-performance computing power or earn by sharing their hardware. Pay with crypto, get instant access to GPUs worldwide.

{% embed url="<https://youtu.be/hA3fO5VWCAk>" %}

***

## Quick Start

### I want to rent GPUs

1. [Create an account](https://clore.ai/login) and [deposit funds](/getting-started/deposit-withdrawal)
2. Browse the [Marketplace](https://clore.ai/marketplace)
3. Choose [On-Demand or Spot](/for-renters/on-demand-vs-spot) rental
4. [Connect to your server](/for-renters/how-to-connect)

{% content-ref url="/pages/cEHgDv7NGBquDGHeLyQ4" %}
[Marketplace Overview](/for-renters/overview)
{% endcontent-ref %}

### I want to host GPUs

1. Install [Clore Hosting Software](/for-hosts/installing-clore-hosting/installing-clore-hosting-software)
2. Configure [Server Settings](/for-hosts/server-settings)
3. Set your [Pricing](/for-hosts/server-settings/how-to-price-your-servers-for-rental)
4. Boost earnings with [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics)

{% content-ref url="/pages/Osggs5K7eGeYr9UurR4i" %}
[Getting Started](/for-hosts/installing-clore-hosting)
{% endcontent-ref %}

***

## Why Clore.ai?

| Feature                 | Benefit                                  |
| ----------------------- | ---------------------------------------- |
| **Peer-to-peer**        | No middlemen, lower costs                |
| **Global network**      | Thousands of GPUs available 24/7         |
| **Flexible pricing**    | On-Demand (guaranteed) or Spot (cheaper) |
| **Crypto payments**     | CLORE, BTC, USDT, USDC                   |
| **Instant access**      | Start using GPUs in minutes              |
| **Earn passive income** | Monetize idle GPU hardware               |

***

## Rental Options

### On-Demand

* **Guaranteed access** — no one can interrupt your rental
* **Fee:** 10% (5% with PoH)
* **Best for:** AI training, rendering, production workloads

### Spot

* **Lower cost** — but can be outbid by others
* **Fee:** 2.5% (1.25% with max PoH)
* **Best for:** Mining, batch jobs, testing

{% content-ref url="/pages/E7V0uUw9wjm5QB2hYGx2" %}
[On-Demand vs Spot](/for-renters/on-demand-vs-spot)
{% endcontent-ref %}

{% content-ref url="/pages/3Fy2rz08NzsQhx7vBl4g" %}
[Fee Structure](/for-renters/fee-structure)
{% endcontent-ref %}

***

## CLORE Token

CLORE is an ERC-20 token on Ethereum — the native currency of the Clore.ai ecosystem.

| Property   | Value                                        |
| ---------- | -------------------------------------------- |
| Network    | Ethereum (ERC-20)                            |
| Contract   | `0xe60201989b8628f43dc0605f585a72bcf1f1e977` |
| Max Supply | 1,000,000,000                                |

**Use CLORE to:**

* Pay for GPU rentals
* Reduce fees via Proof of Holding
* Earn marketplace rewards

{% content-ref url="/pages/JjjojIyKxaiBPzzLwvtg" %}
[Clore coin](/clore-coin)
{% endcontent-ref %}

{% content-ref url="/pages/eoZkyCjSDfQ9akLAkyHP" %}
[Tokenomics](/main/tokenomics)
{% endcontent-ref %}

{% content-ref url="/pages/CchE2DBqhNn0f2i6Y9sI" %}
[Wallet](/wallet)
{% endcontent-ref %}

***

## Proof of Holding (PoH)

Hold CLORE tokens to unlock benefits:

* **Lower fees** — up to 50% reduction on marketplace fees
* **Earn rewards** — daily CLORE distribution to PoH participants
* **No lockup** — withdraw anytime, no penalties

{% content-ref url="/pages/Db4Ut4ssjHaKEmdt6plm" %}
[Overview](/proof-of-holding/overview)
{% endcontent-ref %}

{% content-ref url="/pages/ZXxfwYY21Yz9PCv9X7ja" %}
[PoH Marketplace](/proof-of-holding/poh-marketplace)
{% endcontent-ref %}

***

## Use Cases

| Use Case                 | Description                                           |
| ------------------------ | ----------------------------------------------------- |
| **AI/ML Training**       | Train models on powerful GPUs without buying hardware |
| **Inference**            | Run LLMs, Stable Diffusion, and other AI models       |
| **3D Rendering**         | Render Blender, Maya, and other 3D projects           |
| **Scientific Computing** | Run simulations and data analysis                     |
| **Crypto Mining**        | Mine during idle time or as primary use               |

{% content-ref url="/pages/ROrd9ptTMk5ui0NBaIMI" %}
[Use cases](/for-renters/use-cases)
{% endcontent-ref %}

{% content-ref url="/pages/RwYSmhxRmzWwTgq3QPBq" %}
[Available Docker Images](/for-renters/docker-images)
{% endcontent-ref %}

***

## Ecosystem

### Bare Metal

Dedicated GPU servers and clusters of up to 1,000 GPUs, reserved for fixed terms with self-serve checkout.

{% content-ref url="/pages/4X44XlkrJDlxy86TsAzu" %}
[Bare Metal Rental](/for-renters/bare-metal-rental)
{% endcontent-ref %}

### Clore VPN

Privacy-focused VPN service, paid with CLORE.

{% content-ref url="/pages/c56twejKnklc3SYyRjLX" %}
[Clore VPN](/main/clore-vpn)
{% endcontent-ref %}

***

## Resources

| Resource                  | Link                                                 |
| ------------------------- | ---------------------------------------------------- |
| Marketplace               | [clore.ai/marketplace](https://clore.ai/marketplace) |
| Roadmap                   | [clore.ai/roadmap](https://clore.ai/roadmap)         |
| API Documentation         | [API Guide](/for-hosts/api)                          |
| Interactive API Reference | [clore.ai/api-docs](https://clore.ai/api-docs)       |
| FAQ                       | [Frequently Asked Questions](/help/faq)              |
| Support                   | [Support & Contact](/help/support)                   |

### Contact & Support

* **Support Portal:** [clore.ai/support](https://clore.ai/support)
* **Email:** <support@clore.ai>

### Community

* **Telegram:** [t.me/clorechat](https://t.me/clorechat)
* **Discord:** [discord.gg/clore-ai](https://discord.gg/clore-ai)
* **Twitter:** [@clore\_ai](https://twitter.com/clore_ai)


# Info

Introduction to Clore.ai

<div data-full-width="true"><figure><img src="/files/381CcY773OgaYY2dPjgV" alt=""><figcaption></figcaption></figure></div>

Clore.ai is a revolutionary platform designed to disrupt the traditional GPU leasing industry. By offering a global, peer-to-peer marketplace, Clore.ai enables users to rent high-performance GPUs on demand, providing accessible and affordable computing power for a wide range of applications, including artificial intelligence development, scientific research, video rendering, and cryptocurrency mining.

The platform is built on principles of transparency and fairness. Clore.ai enables GPU owners to monetize their idle resources by leasing them out to users worldwide, while renters gain flexible access to the computing power they need without the high costs typically associated with traditional cloud providers. The platform operates using Clore Coin (CLORE), its native cryptocurrency, which facilitates transactions, rewards participants, and supports the platform’s sustainable growth.<br>

## Core Objectives of Clore.ai

#### 1. Democratize Access to Advanced Computing Power

Clore.ai aims to make high-performance GPU resources accessible to everyone, from individual users and small startups to large enterprises. By providing a cost-effective and flexible marketplace, Clore.ai reduces the barriers to entry, allowing users to access powerful computational resources without the need for significant upfront investments in hardware.

#### 2. Encourage Sustainable Growth

Clore.ai is designed with long-term growth and scalability in mind. Through continuous innovation and community engagement, the platform fosters sustainable development. Its robust economic model, supported by the fixed supply of 1.0 billion Clore Coins (CLORE), promotes stability by gradually reducing rewards over time. This ensures the long-term value of the ecosystem and incentivizes users to actively participate through the Proof of Holding (PoH) system.

#### 3. Foster a Vibrant Ecosystem

Clore.ai encourages active participation through the integration of its native cryptocurrency, Clore Coin. By holding and using CLORE coins, users can reduce fees and increase their earnings. The Proof of Holding (PoH) system rewards users for long-term engagement, creating a vibrant and loyal community that supports the platform’s ongoing growth and success.

#### 4. Expand Platform Offerings

Clore.ai is committed to continuous innovation and keeps expanding its services: Clore VPN is already live, and an AI and ML model marketplace is planned. These additions will further enhance the platform’s value and utility for its global user base, positioning Clore.ai as a comprehensive solution for advanced computing needs.

#### 5. Maximize Utilization of Idle Resources

Clore.ai not only provides a marketplace for leasing GPUs but also ensures that unused GPUs remain productive. If a GPU is not being rented out, owners can switch to cryptocurrency mining through HiveOS, ensuring that GPU resources are fully utilized and profitable.

#### 6. Support Diverse Use Cases

The platform is designed to cater to a wide range of applications, including AI development, machine learning, scientific research, 3D rendering, and cryptocurrency mining. Clore.ai enables users from various industries to leverage its GPU marketplace for both short-term and long-term computational needs.


# Tokenomics

## CLORE — ERC-20 on Ethereum

CLORE is the native token of the Clore.ai GPU marketplace. It powers payments, rewards, and fee tiers across the platform.

* **Network:** Ethereum (ERC-20)
* **Total supply:** 1,000,000,000 CLORE (hard-capped, cannot increase)
* **Contract:** [0xe60201989b8628f43dc0605f585a72bcf1f1e977](https://etherscan.io/token/0xe60201989b8628f43dc0605f585a72bcf1f1e977)

{% embed url="<https://etherscan.io/token/0xe60201989b8628f43dc0605f585a72bcf1f1e977>" %}

***

### Token utility

CLORE is the core payment and incentive asset in the Clore ecosystem:

* **Payments on the GPU marketplace** — CLORE is used as the payment token for renting GPUs and related services on clore.ai.
* **Payments for Clore VPN** (where available) — CLORE is used as the payment token for accessing Clore VPN.
* **PoH & marketplace rewards / fee tiers** — CLORE is used in our Proof-of-Holding and marketplace reward system (fee reductions and rewards for active participants).

<figure><img src="/files/HFlBnoXM4gLquXchyM77" alt=""><figcaption></figcaption></figure>

***

### Supply distribution

<figure><img src="/files/tkuUjXpUr8ywtMZIBq8r" alt=""><figcaption></figcaption></figure>

| Allocation                           | Amount      | Notes                                                                            |
| ------------------------------------ | ----------- | -------------------------------------------------------------------------------- |
| L1 holder claim pool (Main multisig) | 630,000,000 | Reserved for 1:1 conversion of legacy PoW holders. Claim window closed Feb 2026. |
| Marketplace Rewards Vesting          | 143,000,000 | Rewards pool for PoH + marketplace incentives.                                   |
| Marketing Vesting                    | 87,000,000  | 2-year vesting.                                                                  |
| Team Vesting                         | 40,000,000  | 5-year vesting, 6-month cliff.                                                   |
| Treasury Safe                        | 25,000,000  |                                                                                  |
| POL (Protocol-Owned Liquidity)       | 25,000,000  |                                                                                  |
| Ecosystem Wallet                     | 25,000,000  |                                                                                  |
| Growth Wallet                        | 25,000,000  |                                                                                  |

#### <mark style="color:red;">**Fully diluted supply is hard-capped at 1.0B CLORE and cannot increase.**</mark>

<figure><img src="/files/DqlSih4q7LQZMryTLKx0" alt=""><figcaption></figcaption></figure>

***

### Marketplace rewards (PoH + marketplace)

Rewards for PoH and marketplace activity are funded from a dedicated rewards pool.

* Initial daily emission: 164,000 CLORE/day\
  (82,000 PoH + 82,000 marketplace)
* Decay schedule: -30% per year implemented via a monthly factor\
  r = 0.7^(1/12)

***

### Key benefits of ERC-20 CLORE

* **Broad wallet support:** MetaMask, Trust Wallet, Ledger, Trezor, and any standard ERC-20 wallet.
* **DEX liquidity:** Uniswap and other DEXs available.
* **Team allocation reduced:** 10% → 4%, with a **5-year vesting** schedule.
* **Faster integrations + new listings:** exchanges and custodians support ERC-20 flows, enabling quicker rollouts and more exchange listings.
* **Fast deposits:** deposit confirmations on the Clore marketplace and exchanges take **under 3 minutes**.
* **No miners:** miner sell pressure removed, cleaner long-term supply dynamic.
* **Reduced supply:** supply hard-capped at 1B with burn mechanics applied at migration.
* **Better ecosystem tooling:** explorers, analytics, custody solutions, treasury operations, and DeFi infrastructure.
* **Bigger reach + stronger tech foundation:** easier onboarding for a wider audience and more flexibility for future product + DeFi integrations.
* **Cross-chain expansion (planned):** bridges to BSC / Polygon / Solana and other networks (use official bridges only).
* **More openness & decentralization:** ERC-20 standards make CLORE easier to integrate, verify, and use across the broader crypto ecosystem.

***

### Vesting contracts (Ethereum mainnet)

* Team Vesting (5 years, 6-month cliff):\
  0xAA2d16f6134B7f0D861785B2785C90f6E9914f24
* Marketing (2 years vesting):\
  0xd903325a3f9def246629ab1c92ba42fa7c90d14f
* Marketplace Rewards (2 years vesting):\
  0xc4c910324a2e95a975d22394c0277b290c68cbb0

### Multisigs and wallets (Ethereum mainnet)

* Main multisig (L1 holder claim pool):\
  0x7E12737906b92944227f5867B6141290Ea9C775a
* Treasury Safe:\
  0x80eB6B2d648dF461F828af6d66Acee0dCDfcf9E6
* Liquidity / POL multisig:\
  0x5151eb844f9f79611be28dda4220debe44617314
* Ecosystem multisig:\
  0xFe79D578C02A98bd8AE272d6826eEA2092049538
* Growth multisig:\
  0xafdD92B6130c4B13366EFc878fA6B7E4E1D415C5
* Team wallet:\
  0x6195C21Effa5f0a3C7f542616858c2C01C34077b
* Marketing wallet:\
  0xb2362cf645a494145881855a21b82b5b3f27a7dc
* Marketplace wallet / rewards operator:\
  0x7A235C00841639868322eE478b605dD70BD37b97

***

## Legacy: PoW to ERC-20 Migration

{% hint style="info" %}
**Migration complete.** CLORE successfully migrated from the legacy PoW chain to Ethereum (ERC-20) in December 2025. The claim window closed in February 2026. All marketplace operations now run on ERC-20 CLORE.
{% endhint %}

### About the migration

CLORE was migrated from the legacy PoW chain to a new CLORE token on Ethereum mainnet (ERC-20). The migration was 1:1: every CLORE held on the PoW network at the time of the snapshot was eligible to be redeemed for exactly 1 CLORE (ERC-20).

### Why we stopped PoS and chose an ERC-20 migration

The biggest structural issue with PoW for CLORE is constant sell pressure from miners. Mining rewards are regularly sold to cover electricity and operating costs, which creates a built-in, ongoing headwind that is difficult to outgrow.

We spent over 1.5 years working toward PoS. We tested, iterated, and invested significant effort because we believed PoS could be the "next stage" of the Clore chain. This decision was not easy — we truly tried. But as the work progressed, it became clear that finishing PoS properly would require **at least another full year of focused, full-team development**, which would significantly distract us from other critical work across the project.

And even if we committed that time, PoS would still not solve our core objective. **PoS does not remove sell pressure — it shifts it.** Our main motivation for leaving PoW is reducing structural sell pressure. With PoS, sell pressure still exists: it comes from **stakers harvesting yield** rather than miners selling emissions.

After deep internal evaluation, we decided to pause the PoS transition and proceed with an ERC-20 migration for these reasons:

* Maintaining a custom wallet stack adds significant complexity.
* Exchange integrations are slower and more difficult.
* Infrastructure overhead is higher.
* Composability with the broader crypto ecosystem is lower.

The ERC-20 path gives us a cleaner foundation to build a healthier long-term model around real product usage.

### Migration timeline (completed)

| Event                               | Date (UTC)             | Status   |
| ----------------------------------- | ---------------------- | -------- |
| Migration announced                 | 16 Dec 2025            | ✅ Done   |
| Legacy deposits/withdrawals closed  | 20 Dec 2025            | ✅ Done   |
| Snapshot taken                      | 21 Dec 2025, 00:00 UTC | ✅ Done   |
| ERC-20 deposits/withdrawals resumed | 22 Dec 2025            | ✅ Done   |
| Claim window opened                 | 21 Dec 2025            | ✅ Done   |
| Claim window closed                 | Feb 2026               | ✅ Closed |

**Current state:**

* The PoW chain is permanently unsupported — no deposits, withdrawals, or operational support.
* All exchanges and clore.ai operate exclusively on the ERC-20 network.
* The claim portal is closed. Unclaimed legacy CLORE is permanently unrecoverable.


# Clore VPN

<figure><img src="/files/UvSS0s9R7D9Qcacg5Y8e" alt=""><figcaption></figcaption></figure>

We’re excited to announce the launch of **Clore VPN Beta**! You can now connect to our VPN and enjoy enhanced privacy and secure internet access.

### How to Connect to Clore VPN

Follow these simple steps to get started:

1. Visit [**vpn.clore.ai**](https://vpn.clore.ai) and purchase a VPN subscription.
2. After completing the payment, you will receive **credentials** for VPN access.
3. Use any VPN client that supports the VLESS protocol. Recommended clients for different platforms:
   * **iOS**: [Streisand](https://apps.apple.com/ru/app/streisand/id6450534064)
   * **Android**: [V2RayNG](https://play.google.com/store/apps/details?id=com.v2ray.ang)
   * **macOS**: [Streisand](https://apps.apple.com/us/app/streisand/id6450534064)
   * **Windows**: [Hiddify for Windows](https://github.com/hiddify/hiddify-next/releases/download/v2.0.5/Hiddify-Windows-Setup-x64.exe)
4. Insert the provided credentials into your VPN client settings.

You’re now ready to use Clore VPN!

### How to Join the Clore VPN Network

We are actively building the server network for Clore VPN. If you own a server and want to contribute, follow these steps:

1. Go to your server settings on the Clore platform.
2. Click the **Apply for VPN** button.

We are particularly looking for servers located in the following countries:

* **Japan**,
* **Netherlands**,
* **Hungary**,
* **Finland**.

> In the future, we plan to expand the list of available locations and add new countries as needed.


# Clore coin

<div data-full-width="true"><figure><img src="/files/aljH17707cESOe9z928b" alt=""><figcaption></figcaption></figure></div>

**What is CLORE?**

CLORE is the native token of the Clore.ai platform. It is used for payments on the GPU marketplace, fee reductions through Proof of Holding (PoH), and ecosystem rewards.

**Token Details**

| Property   | Value                                        |
| ---------- | -------------------------------------------- |
| Network    | Ethereum (ERC-20)                            |
| Contract   | `0xe60201989b8628f43dc0605f585a72bcf1f1e977` |
| Symbol     | CLORE                                        |
| Decimals   | 18                                           |
| Max Supply | 1,000,000,000 (1 billion)                    |

**Token Utility**

* **Payments** - Pay for GPU rentals on the Clore.ai marketplace
* **Fee Reductions** - Hold CLORE in PoH to reduce marketplace fees
* **Rewards** - Earn CLORE through PoH staking and marketplace activity
* **VPN Payments** - Pay for Clore VPN service (where available)

**Migration to ERC-20**

CLORE was migrated from its original PoW blockchain to Ethereum (ERC-20) in December 2025. This migration brought several benefits:

* Immediate wallet support (MetaMask, Trust Wallet, hardware wallets)
* DEX liquidity (Uniswap and others)
* Faster exchange integrations
* Reduced supply (hard-capped at 1B)
* Better ecosystem tooling and DeFi compatibility

{% content-ref url="/pages/eoZkyCjSDfQ9akLAkyHP" %}
[Tokenomics](/main/tokenomics)
{% endcontent-ref %}


# Wallet

CLORE is an ERC-20 token on the Ethereum network. You can store and manage your CLORE tokens using any wallet that supports Ethereum and ERC-20 tokens.

## CLORE Token Details

| Property | Value                                        |
| -------- | -------------------------------------------- |
| Network  | Ethereum (ERC-20)                            |
| Contract | `0xe60201989b8628f43dc0605f585a72bcf1f1e977` |
| Symbol   | CLORE                                        |
| Decimals | 18                                           |

## Compatible Wallets

Since CLORE is an ERC-20 token, you can use any Ethereum-compatible wallet.

### Browser Extension Wallets

| Wallet              | Description                                          | Link                                                   |
| ------------------- | ---------------------------------------------------- | ------------------------------------------------------ |
| **MetaMask**        | The most popular Web3 wallet with 30M+ monthly users | [metamask.io](https://metamask.io)                     |
| **Rabby Wallet**    | Security-focused with clear transaction previews     | [rabby.io](https://rabby.io)                           |
| **Coinbase Wallet** | Self-custody wallet by Coinbase with easy onboarding | [coinbase.com/wallet](https://www.coinbase.com/wallet) |
| **Phantom**         | Clean interface with built-in swap and NFT features  | [phantom.app](https://phantom.app)                     |
| **Frame**           | Privacy-focused desktop wallet with hardware support | [frame.sh](https://frame.sh)                           |
| **OKX Wallet**      | Multi-chain wallet supporting 80+ networks           | [okx.com/wallet](https://www.okx.com/wallet)           |
| **Rainbow**         | Beautiful wallet designed for NFT collectors         | [rainbow.me](https://rainbow.me)                       |

### Mobile Wallets

| Wallet              | Description                                             | Link                                                   |
| ------------------- | ------------------------------------------------------- | ------------------------------------------------------ |
| **MetaMask**        | Mobile version with browser extension sync              | [metamask.io](https://metamask.io/download/)           |
| **Trust Wallet**    | Multi-chain wallet supporting 100+ blockchains          | [trustwallet.com](https://trustwallet.com)             |
| **Rainbow**         | Simple interface with L2 bridging and cross-chain swaps | [rainbow.me](https://rainbow.me)                       |
| **Zerion**          | DeFi-focused with comprehensive portfolio tracking      | [zerion.io](https://zerion.io)                         |
| **Coinbase Wallet** | Mobile wallet with easy access to DeFi and NFTs         | [coinbase.com/wallet](https://www.coinbase.com/wallet) |

### Hardware Wallets (Cold Storage)

For maximum security, use a hardware wallet:

| Wallet       | Description                                                 | Link                                   |
| ------------ | ----------------------------------------------------------- | -------------------------------------- |
| **Ledger**   | CC EAL5+ certified, supports 5,500+ cryptocurrencies        | [ledger.com](https://www.ledger.com)   |
| **Trezor**   | Open-source with touchscreen interface                      | [trezor.io](https://trezor.io)         |
| **Tangem**   | Card-shaped wallet, no charging needed, 25+ year durability | [tangem.com](https://tangem.com)       |
| **SafePal**  | Air-gapped with QR code signing, backed by Binance          | [safepal.com](https://www.safepal.com) |
| **Keystone** | Air-gapped with large touchscreen, open-source firmware     | [keyst.one](https://keyst.one)         |

### Desktop Wallets

| Wallet            | Description                                       | Link                                       |
| ----------------- | ------------------------------------------------- | ------------------------------------------ |
| **Exodus**        | Multi-platform with built-in exchange and staking | [exodus.com](https://www.exodus.com)       |
| **Atomic Wallet** | Decentralized with built-in atomic swap exchange  | [atomicwallet.io](https://atomicwallet.io) |

***

## Adding CLORE to Your Wallet

### MetaMask / Browser Wallets

1. Open your wallet
2. Click **Import Token** or **Add Token**
3. Select **Custom Token**
4. Enter contract address: `0xe60201989b8628f43dc0605f585a72bcf1f1e977`
5. Symbol and decimals should auto-fill (CLORE, 18)
6. Click **Add**

### Trust Wallet / Mobile

1. Open Trust Wallet
2. Tap the toggle icon (top right)
3. Search for "CLORE" or scroll to find it
4. If not found, tap **Add Custom Token**
5. Select **Ethereum** network
6. Paste contract: `0xe60201989b8628f43dc0605f585a72bcf1f1e977`
7. Save

***

## Troubleshooting

### Token Not Showing

1. Make sure you're on the Ethereum network (not BSC or other networks)
2. Verify the contract address is correct
3. Try adding as a custom token manually
4. Refresh your wallet or restart the app

### Wrong Network

CLORE is an **Ethereum (ERC-20)** token. Make sure you're sending to an Ethereum address, not BSC, Polygon, or other networks. Sending to the wrong network may result in lost funds.


# Conclusion

<div data-full-width="true"><figure><img src="/files/9W5Kev7CYji4knQwWWkH" alt=""><figcaption></figcaption></figure></div>

## Conclusion

The Clore.ai platform represents a transformative leap in the GPU leasing industry, offering a fair, transparent, and efficient solution for accessing GPU power. Through its unique Proof of Holding (PoH) system and the utility of Clore Coin ($CLORE), Clore.ai democratizes access to high-performance computing while rewarding users for their participation. The platform’s commitment to regulatory compliance, security, and sustainability ensures a reliable and trustworthy environment for users.

## Key Points:

1\. Fair Access to GPU Power: Clore.ai eliminates intermediaries by providing a direct and cost-effective way for users to lease GPU power through its peer-to-peer marketplace.

2\. Proof of Holding (PoH): This system offers immediate rewards for holding Clore coins, promoting user engagement and network stability.

3\. Compliance and Security: Clore.ai adheres to strict regulatory standards, safeguarding user data and transactions with advanced security protocols.

4\. Sustainability and Growth: A fixed supply of 1.3 billion Clore coins and a structured reward reduction mechanism ensure long-term value and ecosystem health.

<br>

Join the Clore.ai revolution today by becoming part of our innovative GPU marketplace. Whether you’re looking to harness the power of advanced computing or monetize your idle GPU resources, Clore.ai offers a secure, transparent, and rewarding platform to achieve your goals. Visit our website, explore the Clore Coin opportunities, and start maximizing your potential with Clore.ai!


# Company

<div data-full-width="true"><figure><img src="/files/IfVuNMao4kFOiWMONSDR" alt=""><figcaption></figcaption></figure></div>

Clore.ai was conceived out of a vision to transform the GPU leasing industry by addressing the inefficiencies and high costs associated with traditional cloud computing solutions. The founder recognized a significant opportunity in the underutilized power of GPUs globally—often left idle by their owners due to high costs and limited access to practical applications. This realization led to the birth of Clore.ai, a platform designed to unlock the potential of these GPUs by creating a peer-to-peer marketplace where computing power could be shared, monetized, and accessed more affordably.

## Commitment to Fairness and Accessibility

From the outset, Clore.ai’s development was driven by a commitment to fairness, accessibility, and inclusivity. The platform’s vision was to democratize access to high-performance computing resources, ensuring that individuals, startups, and large enterprises alike could benefit from powerful computational tools without needing significant upfront investments. By eliminating intermediaries, Clore.ai reduced costs for users and created a more accessible and transparent marketplace for all.

## Focus on Regulatory Compliance

Operating as a registered company in the Czech Republic, Clore.ai has placed a strong emphasis on regulatory compliance and legal adherence. The platform aligns itself with stringent European Union regulations, ensuring that all aspects of its operations meet the highest standards of legal scrutiny. This includes robust policies for data protection, cybersecurity, financial regulations, and user rights, all of which are integral to fostering trust within the user base.

## Legal Framework and User Protection

The company’s compliance framework is designed to provide users with confidence in the integrity of their transactions and the legal standing of their activities on the platform. Clore.ai takes security and user protection seriously, employing advanced cybersecurity protocols to safeguard sensitive information and financial transactions. By prioritizing regulatory adherence, the platform can continue to scale globally while ensuring that its services meet the expectations of a rapidly evolving technological landscape.

## Evolution and Future Growth

Clore.ai’s journey from a visionary concept to a leading platform in the GPU leasing industry is a testament to the founders’ dedication to innovation and the empowerment of its global user base. As the industry evolves, Clore.ai continues to adapt and refine its platform to meet the growing demands of its users, ensuring that it remains at the forefront of advanced computing solutions.

The company’s governance and operational strategies continue to evolve, driven by a clear understanding of the industry’s challenges and a commitment to providing innovative solutions. This ongoing adaptation ensures that Clore.ai can meet the needs of current users while anticipating the future demands of the computing landscape.

## Building a Sustainable Future

By staying true to its founding principles of fairness, accessibility, and regulatory compliance, Clore.ai has created a platform that not only addresses the current needs of its users but also positions itself as a leader in the future of GPU leasing. The company’s long-term vision focuses on sustainable growth, regulatory integrity, and user empowerment, ensuring that Clore.ai remains a trusted and innovative platform in a rapidly changing industry.


# Marketplace Overview

Clore.ai GPU Leasing Product

<div data-full-width="true"><figure><img src="/files/8tsneTS6sbJ5hPxd3hEW" alt=""><figcaption></figcaption></figure></div>

## Overview of Clore.ai GPU Leasing

Clore.ai offers a platform that revolutionizes the GPU leasing industry by providing users with flexible, cost-effective access to high-performance computing power. Whether you are an AI researcher, cryptocurrency miner, or a creative professional, Clore.ai enables you to rent GPUs on-demand through a peer-to-peer network, ensuring you get the computational resources you need without the high costs associated with traditional cloud providers.

## Leasing Options:

Clore.ai offers two types of leasing options to cater to different user needs:

1\. On-Demand Leasing: This option comes with a 10% base fee, which can be reduced by up to 50% through the Proof of Holding (PoH) system (proportional to your CLORE holdings, max reduction at 2M CLORE). PoH fee reduction applies to both renters and hosters. On-demand leases are non-interruptible, meaning no one can overbid your offer, and the price is final. This is ideal for tasks that require uninterrupted computing power, such as AI development.

2\. Spot Leasing: This option has a lower base fee of 2.5%, which can be reduced by up to 50% through the PoH system (proportional to your CLORE holdings, max reduction at 2M CLORE). Spot leases are interruptible, allowing other users to overbid your offer. This type of leasing is particularly suitable for tasks like cryptocurrency mining, where continuous processing is less critical.

## More Ways to Rent

* **Bare Metal**: reserve dedicated servers and clusters of up to 1,000 GPUs for fixed terms, from two weeks up to five years, with self-serve checkout. See [Bare Metal Rental](/for-renters/bare-metal-rental).
* **Partial GPU rental**: on supported multi-GPU servers you can rent only some of the GPUs and pay proportionally. See [Partial GPU Rental](/for-renters/partial-gpu-rental).

## Live USD-Based Pricing

You can now list your servers in **USD** while the system automatically adjusts the **CLORE** price **in real time** during the rental, based on the current exchange rate. No more worries about coin fluctuations — your pricing stays accurate and fair in USD terms throughout the rental.\\

**Example**

* Target price: **$10/day**
* If **1 CLORE = $0.05**, the rental price is **200 CLORE**
* If **1 CLORE = $0.10** during the rental, the price adjusts to **100 CLORE** — automatically, in real time

**Notes**

* Currently works **only for On-Demand** orders.
* For **Spot** orders, the legacy **Calculate Based on USD** option still applies: the CLORE amount is calculated at the moment the rental starts and remains fixed for the duration.
* If you encounter any issues, please report them via the [support portal](https://clore.ai/support).

## Key Features and Benefits

Clore.ai connects GPU owners and renters through a marketplace, allowing GPU owners to monetize their unused resources while providing renters with affordable access to high-performance GPUs.

The platform supports a wide range of applications, including artificial intelligence (AI) development, machine learning, scientific research, video rendering, and more.<br>

### Monetize Unused GPU Capacity:

GPU owners can list their idle GPUs on the Clore marketplace, turning unused resources into a revenue stream. By participating in the marketplace, owners can optimize the use of their hardware and generate income.

### Idle GPU Mining with HiveOS:

When a GPU is not rented out, Clore.ai allows owners to utilize HiveOS’s flightsheet to mine any cryptocurrency of their choice. This ensures that GPU resources are always productive, generating income even when not in use by a renter.

### CLORE Token and Host Rewards:

CLORE is an ERC-20 token on Ethereum. Clore.ai distributes **164,000 CLORE daily** (initial rate) to the ecosystem:

* **82,000 CLORE** → [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) rewards (for hosts)
* **82,000 CLORE** → PoH Staking rewards

Emissions decrease by 30% annually.

This unique reward system means hosts can earn **more than the rental fee itself**. For example, if a user rents a GPU for $10, competitors typically take around 20% as a fee, resulting in the host receiving just $8. On Clore.ai, the user pays $10, the platform takes a fee between 1.25% and 10%, and thanks to the reward distribution system, the host could receive an additional CLORE reward on top. This means the host can end up receiving significantly more than the renter pays.

Rewards are distributed based on:

* **Server price** (15%)
* **MFP** (15%) — Maximum Fair Price, Clore's internal quality score for each server
* **Server price × PoH holdings** (25%)
* **MFP × PoH holdings** (45%)

{% content-ref url="<https://github.com/defiocean/clore-docs/tree/main/gpu-marketplace/clore.ai-reward-distribution-system.md>" %}
<https://github.com/defiocean/clore-docs/tree/main/gpu-marketplace/clore.ai-reward-distribution-system.md>
{% endcontent-ref %}

### Competitive Pricing:

Clore.ai offers highly competitive pricing compared to traditional cloud providers. The platform’s operational efficiencies allow for lower costs, which are passed on to users in the form of reduced fees.

The platform also supports spot and on-demand orders, allowing users to choose the pricing model that best suits their needs.

## How Clore.ai GPU Leasing Works

1\) Browse and Select GPUs: users can browse the available servers with desirable GPUs on the marketplace and select the ones that meet their computational requirements. Filters such as GPU model, price, disk size, CPU cores, network speed, server rating, and many more are available to help narrow down the options; an exact Server ID search shows the matching server regardless of other filters. Switch between card and table views (the table has a configurable column picker), pin servers, save favourites, and keep private notes while you compare. Right-click a card (long-press on mobile) for quick actions, and open a server's uptime chart or overclocking profiles straight from the card.

2\) Set Job: create a task for this server: it could be mining a specific cryptocurrency or training an AI model.

3\) Monitor and Manage: clore.ai provides tools for monitoring and managing rented GPUs. Users can track performance, adjust settings, and manage their rentals through an intuitive dashboard.

4\) Idle GPU Mining with HiveOS: for GPU owners, Clore.ai offers the ability to mine any cryptocurrency through HiveOS’s flightsheet if the GPU is not currently rented out. This feature ensures that the GPU remains productive, generating income even when not in use by a renter.

## Who Can Benefit from Clore.ai GPU Leasing?

• AI Developers and Researchers: Access powerful GPUs for training machine learning models, running simulations, and conducting research.

• Cryptocurrency Miners: Leverage Clore.ai’s network to mine cryptocurrencies efficiently with high-performance GPUs.

• Creative Professionals: Utilize GPUs for rendering 3D graphics, editing videos, and other intensive creative tasks.

• Businesses and Startups: Scale computing resources on-demand to meet project needs without investing in expensive hardware.

***

## 📚 Guides & Tutorials

Looking for step-by-step instructions? Check out our full guides library:

👉 [**docs.clore.ai/guides**](https://docs.clore.ai/guides) — walkthroughs for renters, hosts, and everything in between.

***

## Getting Started (Renters)

New to Clore.ai? Follow this path:

1. [**Deposit funds**](/getting-started/deposit-withdrawal) — Add CLORE, BTC, USDT, or USDC
2. [**Understand fees**](/for-renters/fee-structure) — Learn about On-Demand vs Spot pricing
3. [**Choose rental type**](/for-renters/on-demand-vs-spot) — Pick On-Demand or Spot
4. [**Connect to your server**](/for-renters/how-to-connect) — SSH, Jupyter, or web terminal
5. [**Manage your order**](/for-renters/order-management) — Extend, stop, or monitor

{% content-ref url="/pages/8fTxdLzWSsDxoko7LjtC" %}
[New to Clore?](/getting-started/deposit-withdrawal)
{% endcontent-ref %}

***

> 📝 Read more on our blog: [blog.clore.ai](https://blog.clore.ai)


# Fee Structure

## Supported Payment Currencies

Clore.ai supports multiple currencies for marketplace payments:

| Currency          | Deposit | Withdraw | Pay for Rentals | Extra Hoster Fee |
| ----------------- | ------- | -------- | --------------- | ---------------- |
| **CLORE**         | ✅       | ✅        | ✅               | 0%               |
| **Bitcoin (BTC)** | ✅       | ✅        | ✅               | 15%              |
| **USDT/USDC**     | ✅       | ✅        | ✅               | 15%              |

> **Tip:** Paying with CLORE avoids extra currency fees entirely.

***

## Marketplace Fees

### Base Fees

| Order Type | Base Fee |
| ---------- | -------- |
| On-Demand  | 10%      |
| Spot       | 2.5%     |

The base fee is **split 50/50** between renter and host.

### Extra Fees (Non-CLORE Currencies)

When paying with non-CLORE currencies (BTC, USDT, or USDC), an additional **extra hoster fee** is applied:

| Currency  | Extra Renter Fee | Extra Hoster Fee |
| --------- | ---------------- | ---------------- |
| CLORE     | 0%               | 0%               |
| Bitcoin   | 0%               | 15%              |
| USDT/USDC | 0%               | 15%              |

* **Extra renter fees** are not reducible by any mechanism
* **Extra hoster fees** can be reduced to 0% through [MFP Lock](#mfp-extra-fee-reduction) tiers

> **Important:** The extra hoster fee applies when renting with any non-CLORE currency (BTC, USDT, or USDC). Hosts are encouraged to lock MFP to offset this fee.

***

## Fee Reduction Mechanisms

Clore.ai provides two ways to reduce fees:

### 1. PoH (Proof of Holding) Fee Reduction

Hold CLORE tokens in Proof of Holding to reduce **base fees** for both renters and hosters.

**Formula:**

```
fee_reduction = round(min(poh_amount / 2,000,000, 1) × 50)%
```

The reduction scales **linearly** from 0% to a maximum of **50%** at 2,000,000 CLORE:

| CLORE in PoH | Fee Reduction |
| ------------ | ------------- |
| 0            | 0%            |
| 200,000      | 5%            |
| 500,000      | 13%           |
| 1,000,000    | 25%           |
| 1,500,000    | 38%           |
| 2,000,000+   | 50% (maximum) |

**Key points:**

* Applies to **both renter and hoster** base fees
* Works for **all currencies** (CLORE, BTC, etc.)
* Does **NOT** reduce extra fees
* Calculated per-transaction based on current PoH holdings
* Result is always rounded to a whole percentage

### 2. MFP Extra Fee Reduction

Lock CLORE in [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) tiers to reduce **extra hoster fees** on non-CLORE currency orders.

| MFP Tier           | Max Reduction | Capacity          | Unlock Period |
| ------------------ | ------------- | ----------------- | ------------- |
| Tier 1             | 30%           | MFP × 1,000 CLORE | 7 days        |
| Tier 2             | 30%           | MFP × 4,000 CLORE | 14 days       |
| Tier 3             | 40%           | MFP × 2,000 CLORE | 28 days       |
| **All Tiers Full** | **100%**      | MFP × 7,000 CLORE | —             |

Reductions are **additive** and **proportional** to how full each tier is. When all three tiers are fully locked, the extra hoster fee is reduced to **0%**.

**Key points:**

* Only reduces **hoster extra fees** (not renter fees, not base fees)
* Each tier contributes proportionally: `round(fill_ratio × max_reduction%)`
* Only matters for non-CLORE currencies (CLORE has 0% extra fee)
* Calculated dynamically based on current lock balance

***

## Fee Examples

### Example 1: CLORE Payment, No PoH

On-demand order, 100 CLORE/day settlement:

| Component            | Amount        |
| -------------------- | ------------- |
| Renter base fee (5%) | 5 CLORE       |
| Renter extra fee     | 0 CLORE       |
| **Renter pays**      | **105 CLORE** |
| Hoster base fee (5%) | 5 CLORE       |
| Hoster extra fee     | 0 CLORE       |
| **Hoster receives**  | **95 CLORE**  |

### Example 2: Bitcoin Payment, No PoH, No MFP Lock

On-demand order, 100 CLORE equivalent/day:

| Component              | Amount        |
| ---------------------- | ------------- |
| Renter base fee (5%)   | 5 CLORE       |
| Renter extra fee (0%)  | 0 CLORE       |
| **Renter pays**        | **105 CLORE** |
| Hoster base fee (5%)   | 5 CLORE       |
| Hoster extra fee (15%) | 15 CLORE      |
| **Hoster receives**    | **80 CLORE**  |

### Example 3: Bitcoin Payment, 1M CLORE PoH (Both Sides), Full MFP Lock

On-demand order, 100 CLORE equivalent/day:

| Component           | Calculation     | Amount           |
| ------------------- | --------------- | ---------------- |
| Renter base fee     | 5 × (1 - 25%)   | 3.75 CLORE       |
| Renter extra fee    | 0               | 0 CLORE          |
| **Renter pays**     |                 | **103.75 CLORE** |
| Hoster base fee     | 5 × (1 - 25%)   | 3.75 CLORE       |
| Hoster extra fee    | 15 × (1 - 100%) | 0 CLORE          |
| **Hoster receives** |                 | **96.25 CLORE**  |

> With full MFP Lock and 1M CLORE in PoH, Bitcoin payments become nearly as cost-effective as CLORE payments.

***

## Order Creation Fee

A small one-time fee is charged when creating a new order:

| Order Type | Bitcoin       | USDT/USDC | CLORE      |
| ---------- | ------------- | --------- | ---------- |
| On-Demand  | 0.0000003 BTC | $0.10     | 0.5 CLORE  |
| Spot       | 0.0000001 BTC | $0.005    | 0.05 CLORE |

***

## Minimum Balance Requirements

To create an order, you need a minimum balance:

| Order Type | Bitcoin     | USDT/USDC | CLORE   |
| ---------- | ----------- | --------- | ------- |
| On-Demand  | 0.00001 BTC | $1.00     | 5 CLORE |
| Spot       | 0.00001 BTC | $1.00     | 5 CLORE |

***

## Rental Duration Limits

| Parameter      | Value                 |
| -------------- | --------------------- |
| Minimum rental | 6 hours               |
| Maximum rental | 3000 hours (125 days) |

***

## Price Display

The price shown on the marketplace is the **host's base price** (without fees).

Your actual cost as a renter = **listed price + your base fee + your extra fee (if any)**

***

## Related Pages

{% content-ref url="/pages/Db4Ut4ssjHaKEmdt6plm" %}
[Overview](/proof-of-holding/overview)
{% endcontent-ref %}

{% content-ref url="/pages/E7V0uUw9wjm5QB2hYGx2" %}
[On-Demand vs Spot](/for-renters/on-demand-vs-spot)
{% endcontent-ref %}

{% content-ref url="/pages/oEJbVy1ap9ut7R2BNoXV" %}
[MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics)
{% endcontent-ref %}

{% content-ref url="/pages/ZXxfwYY21Yz9PCv9X7ja" %}
[PoH Marketplace](/proof-of-holding/poh-marketplace)
{% endcontent-ref %}

{% content-ref url="/pages/gvc9GYipFrOMhGqP530g" %}
[Host Fee Guide](/for-hosts/host-fees)
{% endcontent-ref %}


# On-Demand vs Spot

Clore.ai offers two rental types with different trade-offs. This guide helps you choose the right one.

## Quick Comparison

| Feature                         | On-Demand           | Spot               |
| ------------------------------- | ------------------- | ------------------ |
| **Guaranteed access**           | ✅ Yes               | ❌ No               |
| **Can be interrupted**          | ❌ No                | ✅ Yes (outbid)     |
| **Base fee**                    | 10%                 | 2.5%               |
| **Fee with max PoH (2M CLORE)** | 5%                  | 1.25%              |
| **Best for**                    | Training, rendering | Mining, batch jobs |

## On-Demand Rental

### How It Works

* You pay a fixed price for guaranteed GPU access
* No one can take over your rental
* Runs for your specified duration (6h - 3000h)

### Best Use Cases

* **ML/AI training** - Long training runs that can't be interrupted
* **3D rendering** - Projects with deadlines
* **Development** - When you need reliable access
* **Production workloads** - Critical tasks

### Pros

* Guaranteed uptime
* Predictable costs
* No risk of interruption

### Cons

* Higher fees (10% vs 2.5%)
* Less flexible pricing

***

## Spot Rental

### How It Works

* You bid a price for GPU access
* Other users can outbid you at any time
* If outbid, your rental ends and you stop paying

### Best Use Cases

* **Cryptocurrency mining** - Can handle interruptions
* **Batch processing** - Jobs that can be restarted
* **Testing** - Quick experiments
* **Fault-tolerant workloads** - With checkpoint/resume

### Pros

* Much lower fees (2.5% vs 10%)
* Cost-effective for flexible workloads
* Good for price-sensitive users

### Cons

* Can be interrupted anytime
* Not suitable for critical tasks
* Need to handle interruptions in your workflow

***

## When to Use Each

### Choose On-Demand when:

* ✅ Your task cannot be interrupted
* ✅ You have a deadline
* ✅ Training a model that doesn't checkpoint well
* ✅ Running production services
* ✅ You need guaranteed availability

### Choose Spot when:

* ✅ Your workload can handle interruptions
* ✅ You're mining cryptocurrency
* ✅ Running batch jobs with checkpointing
* ✅ Cost is your primary concern
* ✅ You're doing quick tests or experiments

***

## Cost Example

Renting a GPU at 10 CLORE/day for 7 days:

| Type      | Base Cost | Fee               | Total           |
| --------- | --------- | ----------------- | --------------- |
| On-Demand | 70 CLORE  | 7 CLORE (10%)     | **77 CLORE**    |
| Spot      | 70 CLORE  | 1.75 CLORE (2.5%) | **71.75 CLORE** |

**Savings with Spot:** 5.25 CLORE (6.8%)

With max PoH (2M+ CLORE in PoH):

| Type      | Base Cost | Fee                 | Total            |
| --------- | --------- | ------------------- | ---------------- |
| On-Demand | 70 CLORE  | 3.5 CLORE (5%)      | **73.5 CLORE**   |
| Spot      | 70 CLORE  | 0.875 CLORE (1.25%) | **70.875 CLORE** |

> **Note on BTC and USDT/USDC payments:** When an order is paid in BTC or USDT/USDC, the host faces an additional 15% extra hoster fee, which can be reduced to 0% through [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics). CLORE payments carry no extra fee. See [Fee Structure](/for-renters/fee-structure) for details.

***

## Tips

1. **Start with Spot** for testing, switch to On-Demand for production
2. **Use checkpointing** if using Spot for training
3. **Monitor your Spot rentals** - set up alerts for when you're outbid
4. **Hold CLORE in PoH** to reduce fees on both types (up to 50% at 2M CLORE)
5. **Pay with CLORE** when possible to avoid extra currency fees on hosts
6. **Need only part of a rig?** Some multi-GPU servers support [Partial GPU Rental](/for-renters/partial-gpu-rental) (on-demand only)

{% content-ref url="/pages/xMpwOnJVzJOgugChTxXj" %}
[Spot Market](/spot/spot-market)
{% endcontent-ref %}


# Bare Metal Rental

Reserve dedicated GPU servers and clusters of up to 1,000 GPUs with self-serve checkout and terms from two weeks to five years.

Bare Metal gives you entire physical servers instead of containers. No virtualization, no resource sharing, full root access, and hardware reserved exclusively for you for the whole rental term.

Use it when you need guaranteed capacity for production workloads: multi-node training, inference clusters, rendering farms, HPC.

## What You Can Rent

|                  |                                                                                                                                                                                                                            |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **GPU models**   | GB200, B300, B200, GH200, H200 (SXM and PCIe), H100, H800, A100, MI300X, MI325X, L40S, L4, A40, A10, V100, T4, RTX PRO 6000 Blackwell, RTX 6000 Ada, RTX A6000, RTX 5090, RTX 5080, RTX 5070, RTX 4090, RTX 3090, RTX 3080 |
| **Regions**      | USA, Japan, EU, France, Slovenia, Iceland, UK, India, Thailand, Hong Kong, China                                                                                                                                           |
| **Cluster size** | Up to 1,000 GPUs per order (quantity options depend on the SKU)                                                                                                                                                            |
| **Rental terms** | From 14 or 30 days (depending on configuration) up to 5 years                                                                                                                                                              |
| **Networking**   | Up to 3.2 Tbps InfiniBand or RoCE fabrics on flagship clusters                                                                                                                                                             |

The catalog changes as providers join, so check the live configurator for current availability and rates.

## Pricing

* Prices are quoted in **USD per GPU per hour** and shown live in the configurator.
* Longer terms cost less per hour: a 12-month contract is significantly cheaper than a 1-month one.
* The total for your selected term is displayed before checkout. No hidden fees.

## How to Order

1. Open [clore.ai/bare-metal](https://clore.ai/bare-metal).
2. Pick a region, GPU configuration, quantity, and rental term. The per-hour and total prices update as you go.
3. Press **Proceed to checkout**. If you are not signed in, you can create an account or sign in right inside the checkout modal.
4. Add your SSH public key (optional at checkout; if you skip it, send it to the team afterwards to get access provisioned).
5. Pay with **Bitcoin** or **USDT/USDC** from your Clore balance. You can top up in the same modal.

After payment the cluster is provisioned and access details are delivered to you. Typical delivery is within a few hours, up to 48 hours for large or custom configurations.

Your active bare metal orders are listed at the top of the same page.

The term is fixed: the hardware is reserved for you for the whole period and released when it ends (disks are sanitized between rentals). To continue beyond the term, contact the team in advance.

## Bare Metal vs Marketplace Rental

|              | Marketplace             | Bare Metal                      |
| ------------ | ----------------------- | ------------------------------- |
| Billing      | Per minute              | Fixed term (14 days to 5 years) |
| What you get | Container on a host     | Entire physical server(s)       |
| Scale        | Single server           | Clusters up to 1,000 GPUs       |
| Start time   | Seconds                 | Hours (up to 48h)               |
| Best for     | Experiments, short jobs | Production, sustained workloads |

## Need Something Custom?

For configurations not listed in the catalog (specific interconnects, storage, compliance requirements), contact the team at <marketing@clore.ai>.

Want to supply Bare Metal servers instead? See the provider requirements:

{% content-ref url="/pages/SMsesMSN0VMEseunbMpg" %}
[Bare Metal](/for-hosts/advanced/bare-metal)
{% endcontent-ref %}


# Partial GPU Rental

Rent only the GPUs you need on a multi-GPU server and pay a proportional price.

Partial rental lets you rent a subset of the GPUs on a multi-GPU server instead of the whole rig. If a host offers an 8x RTX 4090 machine and you need two cards, you rent two and pay for two.

{% hint style="info" %}
Partial rental is rolling out gradually. It appears only on servers whose hosts run the latest hosting software and have not opted out, so at any given moment only part of the marketplace offers it. Coverage grows as hosts update.
{% endhint %}

## How It Works (Renters)

1. Turn on the **Partial GPU rental** filter (there is also a **Min. available GPUs** filter), or look for the **N/M GPUs free** badge on supported server cards.
2. In the rent dialog, pick the GPUs you want in the **GPU allocation** step. The price scales with your share: renting 2 of 8 GPUs costs 2/8 of the full-server price.
3. Complete checkout as usual. Your order runs in an isolated container with exactly the GPUs assigned to you.

Good to know:

* **On-Demand only.** Spot bidding always takes the whole server.
* **Homogeneous rigs only.** Partial rental is available only on servers where all GPUs are the same model. Mixed-GPU rigs can be rented whole-rig only.
* Several renters can share one machine, each with their own isolated container and dedicated GPUs. A single GPU is never shared between two orders.
* GPUs are dedicated to your order, while the machine's CPU, RAM, and disk are shared with any other orders running on the same server.
* Billing stays per minute, and the regular fee structure applies.

## For Hosts

* Partial rental is enabled by default on eligible servers (homogeneous GPUs, up-to-date hosting software). You can opt out per server in your server settings.
* While a server is partially rented, overclocking changes are blocked (Apply shows a warning).
* Pricing follows your regular per-server settings; renters pay proportionally per GPU.

## API

Marketplace listings expose availability in the `partial_gpu_rental` field: `false` when unavailable, or an object with `total_gpus`, `available_gpus`, and `free_indices`. To create a partial order, pass `gpu_count` to `create_order`, and optionally `gpu_indices` to pick exact slots. When `gpu_indices` is omitted, free GPUs are assigned automatically.

{% content-ref url="/pages/VsUJqQ2pg4ZUv2TbJaMy" %}
[API](/for-hosts/api)
{% endcontent-ref %}


# How to Connect

After renting a GPU server on Clore.ai, you can connect via SSH or access pre-installed services like Jupyter Notebook.

## Connection Information

Once your order is active, you'll find connection details in:

1. **My Orders** page → Click on your active order
2. **Order details** will show: IP address, SSH port, and credentials

## SSH Connection

### Linux / macOS

Open Terminal and run:

```bash
ssh root@<IP_ADDRESS> -p <PORT>
```

Example:

```bash
ssh root@185.123.45.67 -p 22022
```

### Windows

**Using PowerShell:**

```powershell
ssh root@<IP_ADDRESS> -p <PORT>
```

**Using PuTTY:**

1. Download [PuTTY](https://putty.org)
2. Enter the IP address in "Host Name"
3. Enter the port number in "Port"
4. Click "Open"
5. Login as `root` with the provided password

## SSH Keys (Recommended)

For passwordless login, add your SSH public key:

1. Go to **Account** → **SSH Keys**
2. Click **Add Key**
3. Paste your public key (usually `~/.ssh/id_rsa.pub`)
4. Save

Your key will be automatically deployed to new rentals.

### Generate SSH Key (if you don't have one)

```bash
ssh-keygen -t rsa -b 4096
```

## Jupyter Notebook Access

Many images come with Jupyter pre-installed:

1. Check your order details for the Jupyter URL
2. Open the URL in your browser (usually `http://<IP>:<JUPYTER_PORT>`)
3. Enter the token shown in your order details

## File Transfer

### Using SCP

Upload file to server:

```bash
scp -P <PORT> /local/file.txt root@<IP>:/remote/path/
```

Download file from server:

```bash
scp -P <PORT> root@<IP>:/remote/file.txt /local/path/
```

### Using rsync

For larger transfers or syncing directories:

```bash
rsync -avz -e "ssh -p <PORT>" /local/folder/ root@<IP>:/remote/folder/
```

## Port Forwarding

To access services running on the server locally:

```bash
ssh -L 8080:localhost:8080 root@<IP> -p <PORT>
```

This forwards server port 8080 to your local port 8080.

## Common Issues

| Issue              | Solution                                                                  |
| ------------------ | ------------------------------------------------------------------------- |
| Connection refused | Check if the server is fully booted (wait 2-3 minutes after order starts) |
| Permission denied  | Verify password/SSH key is correct                                        |
| Connection timeout | Check if the port is correct                                              |
| Host key changed   | Remove old key: `ssh-keygen -R [<IP>]:<PORT>`                             |

## Useful Commands After Connecting

Check GPU status:

```bash
nvidia-smi
```

Check available disk space:

```bash
df -h
```

Check system resources:

```bash
htop
```


# Available Docker Images

When renting a GPU server on Clore.ai, you can choose from pre-configured Docker images or use your own custom images.

## Clore Official Images

Pre-built images maintained by Clore.ai, optimized for the decentralized GPU marketplace.

### General Purpose

| Docker Image                              | Description                       | Ports    | Updated     |
| ----------------------------------------- | --------------------------------- | -------- | ----------- |
| `cloreai/jupyter:ubuntu24.04-v2`          | Jupyter Lab + SSH on Ubuntu 24.04 | 22, 8888 | Jan 2025 ✅  |
| `cloreai/ml-tools:0.1`                    | Jupyter Lab + VS Code Web server  | 22, 8888 | Jul 2023 ⚠️ |
| `cloreai/ubuntu20.04-jupyter:latest`      | Ubuntu 20.04 + SSH + Jupyter      | 22, 8888 | Nov 2022 ⚠️ |
| `cloreai/ubuntu-20.04-remote-desktop:1.2` | Ubuntu with remote desktop GUI    | 22, 3389 | May 2023 ⚠️ |
| `cloreai/torch:2.0.1`                     | PyTorch 2.0.1 + CUDA              | 22       | Jul 2023 ⚠️ |

### AI/ML Inference

| Docker Image                            | Description                 | Ports | Updated     |
| --------------------------------------- | --------------------------- | ----- | ----------- |
| `cloreai/deepseek-r1-8b:latest`         | DeepSeek R1 8B ready to run | 8000  | Jan 2025 ✅  |
| `cloreai/stable-diffusion-webui:latest` | AUTOMATIC1111 SD WebUI      | 7860  | Sep 2023 ⚠️ |
| `cloreai/oobabooga:1.5`                 | Text Generation WebUI       | 7860  | Aug 2023 ⚠️ |

### Infrastructure & Mining

| Docker Image                | Description                 | Updated     |
| --------------------------- | --------------------------- | ----------- |
| `cloreai/monitoring:0.7`    | Server monitoring agent     | Sep 2024 ✅  |
| `cloreai/hiveos:0.4`        | HiveOS integration          | Feb 2025 ✅  |
| `cloreai/openvpn-proxy:0.2` | VPN/proxy forwarding        | Feb 2025 ✅  |
| `cloreai/proxy:0.2`         | Port forwarding system      | Jan 2024    |
| `cloreai/automining:0.1`    | Auto mining setup           | Jun 2023 ⚠️ |
| `cloreai/kuzco:latest`      | Kuzco distributed inference | Jun 2025 ✅  |
| `cloreai/partner:nastya-01` | Partner integration tools   | Apr 2025 ✅  |

> ⚠️ Images marked with ⚠️ haven't been updated in over a year. For ML workloads, consider using the **community images** below which offer newer CUDA and framework versions.

## Recommended Community Images

Battle-tested images from the broader AI/ML community with active maintenance and recent versions.

### Deep Learning Frameworks

| Image                   | Tag                              | Description                   | Use Case                         | Ports              |
| ----------------------- | -------------------------------- | ----------------------------- | -------------------------------- | ------------------ |
| `pytorch/pytorch`       | `2.10.0-cuda13.0-cudnn9-runtime` | Latest PyTorch with CUDA 13.0 | Deep learning training/inference | 8888 (Jupyter)     |
| `nvidia/cuda`           | `13.1.1-runtime-ubuntu22.04`     | NVIDIA CUDA runtime           | Custom CUDA applications         | -                  |
| `nvidia/cuda`           | `13.1.1-devel-ubuntu22.04`       | CUDA development tools        | Building CUDA projects           | -                  |
| `tensorflow/tensorflow` | `2.15.0-gpu`                     | TensorFlow GPU support        | TensorFlow workflows             | 8888 (TensorBoard) |

### LLM Inference Engines

| Image                                           | Tag      | Description                  | Use Case                  | Ports |
| ----------------------------------------------- | -------- | ---------------------------- | ------------------------- | ----- |
| `vllm/vllm-openai`                              | `latest` | High-performance LLM serving | Production LLM APIs       | 8000  |
| `ghcr.io/huggingface/text-generation-inference` | `3.3.5`  | Hugging Face TGI             | Enterprise LLM serving    | 80    |
| `ollama/ollama`                                 | `latest` | Local LLM runner             | Local/edge LLM deployment | 11434 |

### Image Generation

| Image                             | Tag      | Description            | Use Case                 | Ports |
| --------------------------------- | -------- | ---------------------- | ------------------------ | ----- |
| `goolashe/automatic1111-sd-webui` | `latest` | Stable Diffusion WebUI | AI art generation        | 7860  |
| `sinfallas/comfyui`               | `latest` | ComfyUI node-based SD  | Advanced image workflows | 8188  |

### Development Environments

| Image                         | Tag                            | Description          | Use Case               | Ports |
| ----------------------------- | ------------------------------ | -------------------- | ---------------------- | ----- |
| `jupyter/tensorflow-notebook` | `latest`                       | Jupyter + TensorFlow | ML experimentation     | 8888  |
| `jupyter/pytorch-notebook`    | `latest`                       | Jupyter + PyTorch    | Deep learning research | 8888  |
| `cschranz/gpu-jupyter`        | `v1.10_cuda-12.9_ubuntu-24.04` | GPU-enabled Jupyter  | GPU computing          | 8888  |

## Selecting an Image

### Via Web Interface

1. Go to **Marketplace** → Find a server → **Rent**
2. In the order form, select **Docker Image** from dropdown
3. Choose from pre-configured options or enter custom image
4. Configure exposed ports (comma-separated: `8888,7860,8000`)
5. Add environment variables if needed
6. Submit order

### Via API

```json
{
  "image": "pytorch/pytorch:2.10.0-cuda13.0-cudnn9-runtime",
  "ports": [8888, 8000],
  "env": {
    "JUPYTER_ENABLE_LAB": "yes"
  }
}
```

## Custom Docker Images

### Supported Registries

* **Docker Hub**: `username/image:tag`
* **GitHub Container Registry**: `ghcr.io/user/image:tag`
* **NVIDIA NGC**: `nvcr.io/nvidia/image:tag`
* **Google Container Registry**: `gcr.io/project/image:tag`

### Examples

```bash
# PyTorch latest with CUDA 12.1
pytorch/pytorch:2.1.0-cuda12.1-cudnn8-runtime

# NVIDIA CUDA base
nvidia/cuda:12.2.0-base-ubuntu22.04

# Hugging Face Transformers
huggingface/transformers-pytorch-gpu:4.35.0

# Custom image
your-username/my-ai-model:v1.0
```

### Requirements for Custom Images

* ✅ Must be publicly accessible
* ✅ Should include NVIDIA GPU support for GPU instances
* ✅ Base on CUDA-enabled images for GPU acceleration
* ✅ Include necessary drivers and libraries
* ⚠️ Large images (>10GB) may take longer to download

## Port Configuration

Expose ports for your applications to make them accessible from outside:

| Port  | Common Use           | Framework             |
| ----- | -------------------- | --------------------- |
| 22    | SSH access           | System                |
| 8888  | Jupyter Notebook/Lab | Jupyter               |
| 7860  | Gradio interfaces    | SD WebUI, Gradio apps |
| 8000  | API servers          | vLLM, FastAPI         |
| 3000  | Web applications     | React, Node.js        |
| 8080  | HTTP services        | General web services  |
| 11434 | Ollama API           | Ollama                |
| 8188  | ComfyUI interface    | ComfyUI               |

### Setting Ports in Order Form

```
8888,7860,8000
```

## Environment Variables

Pass configuration to your containers:

### Common Examples

```bash
# Hugging Face token
HF_TOKEN=hf_your_token_here

# Model configuration
MODEL_NAME=meta-llama/Llama-2-7b-chat-hf
CUDA_VISIBLE_DEVICES=0

# Jupyter configuration
JUPYTER_ENABLE_LAB=yes
JUPYTER_TOKEN=your_secure_token

# vLLM configuration
VLLM_WORKER_MULTIPROC_METHOD=spawn
```

### Security Notes

* ❌ Don't put secrets in environment variables
* ✅ Use temporary tokens or API keys
* ✅ Mount secrets as volumes when possible

## Persistent Storage

### Storage Locations

* `/workspace` - Usually persistent during rental period
* `/root` - May be reset on container restart
* `/tmp` - Temporary storage, not persistent

### Best Practices

* Store important data in `/workspace`
* Use external storage for backups (S3, GCS, etc.)
* Download models to persistent directories
* Use volume mounts for large datasets

## Best Practices

### Image Selection

1. **Use recent tags** - Avoid `latest` in production, prefer versioned tags
2. **Choose appropriate base** - Match CUDA version with your framework
3. **Consider image size** - Smaller images download faster
4. **Check maintenance** - Prefer actively maintained images

### Security

1. **Expose minimal ports** - Only expose ports you need
2. **Use authentication** - Set tokens for Jupyter/web interfaces
3. **Update regularly** - Use recent image versions
4. **Monitor access** - Check who connects to your services

### Performance

1. **GPU compatibility** - Verify CUDA version matches your needs
2. **Pre-download models** - Include models in custom images for faster startup
3. **Optimize containers** - Use multi-stage builds, minimize layers
4. **Cache management** - Leverage Docker layer caching

## Troubleshooting

### Image Won't Start

```bash
# Check if image exists and is public
docker pull your-image:tag

# Verify CUDA compatibility
nvidia-smi
nvcc --version
```

**Common Issues:**

* Image doesn't exist or is private
* Incompatible CUDA version
* Insufficient disk space
* Wrong architecture (arm64 vs x86\_64)

### GPU Not Accessible

```bash
# Check GPU availability
nvidia-smi

# Test CUDA in container
python -c "import torch; print(torch.cuda.is_available())"
```

**Solutions:**

* Use GPU-compatible base images
* Ensure NVIDIA Container Toolkit is available
* Check CUDA driver compatibility

### Can't Access Exposed Ports

1. Wait for container to fully start (check logs)
2. Verify service is running inside container: `netstat -tlnp`
3. Check if service binds to 0.0.0.0, not 127.0.0.1
4. Confirm port is exposed in order form

### Performance Issues

* Use local SSD storage for model weights
* Optimize batch sizes for available GPU memory
* Monitor GPU utilization: `nvidia-smi -l 1`
* Check CPU/memory usage: `htop`

## Quick Start Examples

### Deploy Jupyter with PyTorch

```bash
Image: pytorch/pytorch:2.10.0-cuda13.0-cudnn9-runtime
Ports: 8888
Command: jupyter lab --ip=0.0.0.0 --allow-root --no-browser
```

### Deploy vLLM API Server

```bash
Image: vllm/vllm-openai:latest
Ports: 8000
Env: MODEL_NAME=microsoft/DialoGPT-medium
Command: python -m vllm.entrypoints.openai.api_server --model $MODEL_NAME --host 0.0.0.0
```

### Deploy Stable Diffusion WebUI

```bash
Image: goolashe/automatic1111-sd-webui:latest
Ports: 7860
Command: --listen --api --xformers
```

### Deploy Ollama

```bash
Image: ollama/ollama:latest
Ports: 11434
Command: ollama serve
# Then: ollama run llama2 (inside container)
```


# Order Management

How to create and manage your GPU rentals on Clore.ai.

## Creating an Order

1. Pick a server on the [marketplace](https://clore.ai/marketplace) and press **Rent**.
2. The rent dialog walks you through three steps: **choose a Docker image** (or one of your saved templates), **configure it** (ports, environment variables, SSH key, Jupyter token), and **Confirm & pay**. At the last step you either rent on-demand at the listed price or switch to **Rent on spot** and place your bid.
3. Confirm to pay from your balance (available currencies depend on which coins the server accepts). If you are not signed in or short on funds, you can sign in and top up right inside the dialog, no need to leave checkout.

> On servers showing an **N/M GPUs free** badge you can rent a subset of the GPUs and pick them in the GPU allocation step. See [Partial GPU Rental](/for-renters/partial-gpu-rental).

## Viewing Your Orders

1. Go to **My Orders** in the navigation
2. See all your active and past orders
3. Click on any order for details

### Order Statuses

| Status         | Meaning                               |
| -------------- | ------------------------------------- |
| **Pending**    | Order is being set up                 |
| **Running**    | Server is active and accessible       |
| **Completed**  | Rental period ended normally          |
| **Terminated** | Ended early (by you, host, or system) |
| **Outbid**     | Spot order - someone outbid you       |

## Order Details

Each order shows:

* Server specifications (GPU, RAM, storage)
* Connection info (IP, port, credentials)
* Start/end time
* Cost breakdown
* Logs and history

## Extending a Rental

To extend an active On-Demand order:

1. Go to **My Orders**
2. Click on your active order
3. Click **Extend**
4. Select additional hours (up to max 3000h total)
5. Confirm payment

> **Note:** Spot orders cannot be extended - create a new order instead.

## Canceling an Order

### On-Demand Orders

1. Go to **My Orders**
2. Click on the active order
3. Click **Terminate**
4. Confirm termination

You stop being charged immediately upon termination.

### Spot Orders

Same process - but remember, with Spot you can also be terminated if outbid.

## Order History

View your complete rental history:

1. Go to **My Orders**
2. Filter by status or date
3. Click on any order for full details

### Exporting History

Order history can be viewed in your account for tracking expenses and usage patterns.

## Billing & Costs

### How Billing Works

* **On-Demand:** Charged per minute while the order runs; terminate anytime to stop the charges
* **Spot:** Charged per minute at your bid price while your offer holds the server

### Cost Breakdown

Each order shows:

* Base rental cost
* Platform fee
* Total paid

> **Creation fee:** a small one-time fee applies per order (see [Fee Structure](/for-renters/fee-structure)). It is charged only if the order stays active longer than 10 minutes, so quickly cancelled orders pay none.

### Refunds

* Early termination: You stop paying immediately
* Server issues: Contact support for potential refunds

## Communication with Host

### In-Order Chat

1. Open your active order
2. Use the **Chat** feature
3. Communicate directly with the host

### When to Contact Host

* Server not responding
* Specs don't match listing
* Need specific configuration
* Technical issues

## Order Logs

View what happened during your rental:

1. Open order details
2. Go to **Logs** section
3. See events like:
   * Order created
   * Server started
   * Connection established
   * Termination reason

## Tips for Managing Orders

1. **Monitor active orders** - Check regularly for any issues
2. **Set budget alerts** - Track your spending
3. **Save server favorites** - For easy re-rental
4. **Review before extending** - Make sure you still need the server
5. **Terminate promptly** - Don't pay for unused time

## Common Questions

### Can I pause a rental?

No, rentals run continuously. Terminate and create a new order if needed.

### What happens if I run out of balance?

The order will be terminated after a grace period (\~5 minutes).

### Can I transfer an order to another user?

No, orders are tied to your account.

### How long until I can access a new order?

Usually 2-5 minutes after order creation.


# How to use templates (Mass rent)

{% hint style="info" %}
Template management has moved since these screenshots were taken: you now add, edit, and delete templates from the marketplace sidebar, pick saved ones from the **My templates** tab at the image step of the rent dialog, and quick-rent with the selected template. The general idea below still applies.
{% endhint %}

1\. Navigate to the [Clore Marketplace](https://clore.ai/marketplace).

2\. Select the preferred server from which you would like to start your rental.

<figure><img src="/files/6n1YbXo0bro0TWSPIivP" alt=""><figcaption></figcaption></figure>

3\. Click the “Rent” button next to the chosen server.

4\. Configure your server image. For this example, select Ubuntu Jupiter.

5\. Once the configuration is complete, click “Save Template.”

<figure><img src="/files/SF4ILXKx4PLCj5tMD39Z" alt="" width="563"><figcaption></figcaption></figure>

6\. Assign a name to your template and click “Create.”

<figure><img src="/files/lR46WX67ctqKIFZkX3UW" alt="" width="375"><figcaption></figcaption></figure>

7\. Return to the [Clore Marketplace](https://clore.ai/marketplace), and you will now see a template icon.

<figure><img src="/files/4E7Yo98yZjLphiV6mGuE" alt="" width="340"><figcaption></figcaption></figure>

8\. Click “Use Templates,” then pick your template and hit “Use for Rent.”

<figure><img src="/files/UDlO7X8B9Oq9U7Vq6ioz" alt="" width="375"><figcaption></figcaption></figure>

9\. Click on every desired server. The rental process will automatically begin.

<figure><img src="/files/RrxsnGbRujVEMRuWufo5" alt="" width="348"><figcaption></figcaption></figure>


# Renter Troubleshooting

Common issues and solutions when renting GPU servers on Clore.ai.

## Server Connection Issues

### Cannot Connect via SSH

**Symptoms:** Connection refused, timeout, or permission denied.

**Solutions:**

1. **Wait for server to boot**
   * New orders take 2-5 minutes to fully initialize
   * Check order status in "My Orders" - should show "Running"
2. **Verify connection details**
   * Double-check IP address and port from order details
   * Use the exact command format: `ssh root@IP -p PORT`
3. **Check your network**
   * Try from a different network (some ISPs block certain ports)
   * Use Clore VPN if available in your region
4. **SSH key issues**
   * Ensure your public key is added in Account → SSH Keys
   * Try password authentication if keys don't work

### Server Not Responding

**Symptoms:** Server was working but now unresponsive.

**Solutions:**

1. **Check server status** in My Orders
2. **Wait 5-10 minutes** - server may be restarting
3. **Contact host** via chat if issue persists
4. **Request refund** if server is completely unavailable

## GPU Issues

### GPU Not Detected

**Symptoms:** `nvidia-smi` returns error or shows no GPUs.

**Solutions:**

```bash
# Check if NVIDIA driver is loaded
lsmod | grep nvidia

# Restart NVIDIA services
sudo systemctl restart nvidia-persistenced

# Check for driver issues
dmesg | grep -i nvidia
```

If GPU still not detected:

1. Contact the host via order chat
2. Request server restart through support
3. Consider ending order and finding different server

### CUDA Errors

**Symptoms:** CUDA out of memory, version mismatch errors.

**Solutions:**

1. **Check CUDA version compatibility**

   ```bash
   nvcc --version
   nvidia-smi
   ```
2. **Free GPU memory**

   ```bash
   # Find processes using GPU
   nvidia-smi

   # Kill specific process
   kill -9 <PID>
   ```
3. **Use appropriate Docker image** with matching CUDA version

### GPU Performance Lower Than Expected

**Solutions:**

1. Check if other processes are using GPU: `nvidia-smi`
2. Verify GPU model matches listing
3. Monitor GPU utilization during your workload
4. Contact host if specs don't match advertised

## Docker / Container Issues

### Container Won't Start

**Solutions:**

1. **Check Docker status**

   ```bash
   docker ps -a
   docker logs <container_id>
   ```
2. **Restart Docker service**

   ```bash
   sudo systemctl restart docker
   ```
3. **Check disk space**

   ```bash
   df -h
   ```

### Missing Dependencies

**Solutions:**

1. Install required packages:

   ```bash
   apt-get update && apt-get install -y <package>
   ```
2. Use pip for Python packages:

   ```bash
   pip install <package>
   ```
3. Consider using a different Docker image with pre-installed tools

## Network Issues on Server

### Cannot Access Internet from Server

**Solutions:**

1. **Check DNS**

   ```bash
   cat /etc/resolv.conf
   # Try adding Google DNS
   echo "nameserver 8.8.8.8" >> /etc/resolv.conf
   ```
2. **Test connectivity**

   ```bash
   ping 8.8.8.8
   curl -I https://google.com
   ```

### Ports Not Accessible

Some servers have firewall restrictions:

1. Check which ports are open in order details
2. Use SSH port forwarding for services
3. Contact host if you need specific ports opened

## Billing Issues

### Charged More Than Expected

1. Check order history for actual usage duration
2. Review fee structure (On-Demand vs Spot rates)
3. Account for creation fees
4. Contact support with order ID if discrepancy exists

### Order Terminated Unexpectedly

**For Spot orders:**

* Another user may have outbid you
* This is normal for Spot rentals
* Use On-Demand for guaranteed access

**For On-Demand orders:**

* Check if balance was sufficient
* Review order logs for termination reason
* Contact support if unclear

## Getting Help

### Contact Host

* Use the chat feature in your order details
* Hosts typically respond within a few hours

### Clore.ai Support

* Discord: [discord.gg/clore-ai](https://discord.gg/clore-ai)
* Use #support channel for technical issues

### Request Refund

If server doesn't meet advertised specs:

1. Document the issue (screenshots, logs)
2. Contact host first
3. If unresolved, contact Clore.ai support with evidence


# Use cases

## **Case 1**

**Leasing GPUs for AI Development**

<figure><img src="/files/PctPkb4FFh0Kgsa6c4wc" alt=""><figcaption></figcaption></figure>

**Step 1: Platform Access**

Users access the Clore.ai platform via the web interface

They log in using their credentials or create a new account if they are first-time users.

#### **Step 2:** Browsing and Selecting Servers with Desirable GPUs

In the GPU marketplace, users select servers equipped with their desired GPUs rather than selecting just the GPU alone. The platform offers filters to search by GPU model, GPU count, CPU cores, RAM, reliability, network speed (upload/download), CUDA version, country, PCI width, and version. Users can choose between:

• Mainline Market: Offers GPUs without power limits, ideal for AI workloads.

• Efficient Market: Provides power-limited (PL enabled) GPUs for more energy-efficient tasks.

#### **Step 3:** Payment Process

Once a server is selected, users can see the daily rental cost upfront. Payments are made from an internal wallet integrated into the platform, allowing users to pay using Clore Coin (CLORE), Bitcoin (BTC), or USD stablecoins (USDT/USDC).

#### **Step 4:** Lease Management

After payment, users can choose from pre-configured system images suited for AI work, such as “Stable Diffusion” for machine learning, or a basic “Ubuntu Jupiter” image for general computing tasks. This helps users get started quickly without the need for manual software setup.

<figure><img src="/files/Nk3NfAjAWWGic0wQ4GJt" alt=""><figcaption></figcaption></figure>

#### **Step 5:** Utilizing the Server

Once the payment is confirmed, users gain access to the server, enabling them to run AI models, simulations, or other computational workloads.

> 📚 See also: [How to Train Your AI Model for Under $1/Hour](https://blog.clore.ai/how-to-train-your-ai-model-for-under-1hour-complete-guide-with-cloreai/)

## Case 2

**Renting GPUs for Cryptocurrency Mining**

#### **Step 1:** Selecting Mining-Optimized Servers

Miners enter the marketplace and filter listings to find servers suitable for cryptocurrency mining based on their hardware requirements (e.g., GPU model, count, network speed).

#### **Step 2:** Leasing and Setup in HiveOS

Once a server is selected, miners proceed with the leasing process, similar to the AI use case. After leasing, they can integrate the machine into their HiveOS account by providing their HiveOS credentials. All mining configurations, software installations, and monitoring are managed within HiveOS.

**Step 3:** Mining Earnings and Payments

Any cryptocurrency mined during the lease is credited directly to the miner’s wallet that was set in HiveOS.


# Getting Started

## Welcome to the Clore.ai server hosting guide!<br>

If you’re looking to start earning by hosting your servers on Clore.ai, you’ve come to the right place! This manual will guide you through all the key steps—from initial setup to revenue optimization. You’ll learn how to properly connect your server, set up rentals, and ensure its stable performance.

## Here’s what you’ll find in this guide:

1\. Setting Up Your Server for Rentals: A step-by-step walkthrough on connecting your server to the platform, making it available for rent, including HiveOS integration and overclocking profile setup.

2\. Calculating and Setting Rental Prices: Practical guidance on determining rental rates based on mining profitability or use in complex tasks, such as AI processing. We’ll explain how to set an effective price and configure automatic updates.

3\. Managing Rentals: Learn how to set up a Telegram bot for real-time rental status notifications and server monitoring.

4\. Reinstallation and Saving Settings: Simple and reliable methods for reinstalling drivers and retaining previous tokens if you need to update the system.

5\. Earning Bonuses on Clore: An overview of the POH and MFP systems—key tools for boosting profitability and earning bonuses, including how to set up your servers for extra earnings.<br>

## Clore.ai opens new possibilities for GPU server owners: our guide will help you connect, configure, and maximize your server’s potential for a stable income.

Ready to get started? Follow the steps in our guide and take the first step toward earning with Clore.ai!


# Installing on Ubuntu

## Server Requirements

The server (or rig – these terms are nearly interchangeable in this context) must be equipped with NVIDIA GPUs, as AMD is currently not supported. The minimum required disk space is 32 GB; for reliability, it's recommended to use an SSD instead of a flash drive. A minimum of 8 GB of RAM is required, but 16 GB will provide greater stability. As for the CPU, the system can work with a Celeron on a 1151 socket, but for more efficient performance, consider using a CPU like the i7-6700.

Before proceeding, it is highly recommended to disable any overclocking, including the Power Limit (PL), and reset the GPUs to factory settings. Afterward, stress-test the system for stability by, for example, testing the GPUs using the kawpow algorithm and loading the CPU. Monitor temperatures and ensure everything is running stably.

If the system operates stably and temperatures are within a safe range, continue to the next step in the instructions. If temperatures are too high or errors occur, address these issues first – for example, by improving cooling or troubleshooting – and ensure stable operation before proceeding.

## Recommended OS & Drivers

### Operating System

* **Ubuntu 22.04 LTS** — recommended, best GPU driver compatibility
* **Ubuntu 24.04 LTS** — supported, but kernels 6.16+ may have issues with older driver branches (R550 and below)

### NVIDIA Drivers

| Branch      | Version    | CUDA Support    | Recommended For                |
| ----------- | ---------- | --------------- | ------------------------------ |
| R580 (LTSB) | 580.126.18 | Up to CUDA 12.8 | Most GPUs, long-term stability |
| R590        | 590.48.01  | Up to CUDA 13.1 | RTX 50 series, latest features |

Install the recommended driver:

```bash
sudo apt install nvidia-driver-580
```

Or for RTX 50 series:

```bash
sudo apt install nvidia-driver-590
```

### CUDA Toolkit (for ML/AI workloads)

Renters running ML workloads expect CUDA. Recommended versions:

| CUDA Version | Min Driver | Status                          |
| ------------ | ---------- | ------------------------------- |
| CUDA 12.8    | R570+      | Stable, wide ecosystem support  |
| CUDA 13.1    | R590+      | Latest, RTX 50 series optimized |

Most Docker images include their own CUDA runtime, so hosts don't always need to install the CUDA Toolkit system-wide. However, having compatible drivers is essential.

## Registering and Adding the Server

### 1. Go to the [website](http://clore.ai/), register, log in, and navigate to the marketplace:

<figure><img src="https://img1.teletype.in/files/0e/86/0e86de72-544d-48d8-8d82-cf120e516a81.png" alt=""><figcaption></figcaption></figure>

### 2. **Adding a Server:** There are two ways to add a server:

**Method 1:** Go to the "My Servers" section and click the "+Add Server" button. Enter the server's name and click "Next."

<figure><img src="https://img4.teletype.in/files/f7/8e/f78e0a46-06fa-4a5d-b429-f21b78eafb6c.png" alt=""><figcaption></figcaption></figure>

After adding, the server will be marked with a red circle, meaning it is inactive. We will activate it later, but for now, click on the created server to obtain a key – you will need it later.

<figure><img src="https://img4.teletype.in/files/36/ae/36aeeab8-98e0-4fea-81e9-d731d5211df2.png" alt=""><figcaption></figcaption></figure>

### 3. **Run Updates in Sequence:**

```bash
sudo apt update && sudo apt upgrade -y
```

### 4. Install dependencies:

```
sudo apt install -y curl git gnupg lsb-release
```

### 5. Switch to superuser mode:

```bash
sudo -i
```

### 6. **Install the Software:**

```bash
bash <(curl -s https://gitlab.com/cloreai-public/hosting/-/raw/main/install.sh)
```

If the system reports that `git` is missing, install it with:

```bash
apt install -y git
```

Then retry the installation.

If you encounter a `gpg` error, use:

<figure><img src="https://telegra.ph/file/e2ef8c5760193ad523e20.png" alt=""><figcaption></figcaption></figure>

```bash
apt install gpg -y --allow-downgrades
```

<figure><img src="https://img3.teletype.in/files/66/1c/661c9073-cc8e-4734-aa85-cff08902d4d6.png" alt=""><figcaption></figcaption></figure>

Afterward, rerun the installation.

```
bash <(curl -s https://gitlab.com/cloreai-public/hosting/-/raw/main/install.sh)
```

### 7. **Activate the Server:**

```bash
/opt/clore-hosting/clore.sh --init-token <token>
```

Replace `<token>` with the key obtained earlier.

If an error indicates a missing folder or file, the installation likely didn’t complete correctly, and the `clore-hosting` folder was not created. In this case, repeat the installation.

### 8. **Finalize:**

Restart the rig, wait a moment, and refresh the marketplace page. If everything was set up correctly, the server will be marked with a green circle.

```
sudo reboot
```

<figure><img src="https://img2.teletype.in/files/98/9c/989c1cbd-2670-4568-b784-020af71451be.png" alt=""><figcaption></figcaption></figure>

## How to Disable All Installed Services

If you need to disable everything previously installed:

1. Disable the services:

   ```bash
   systemctl disable clore-hosting.service
   systemctl disable docker.service
   systemctl disable docker.socket
   ```
2. Reboot the system:

   ```bash
   reboot
   ```

## How to Re-enable Services

To re-enable the services:

1. Enable the services:

   ```bash
   systemctl enable clore-hosting.service
   systemctl enable docker.service
   systemctl enable docker.socket
   ```
2. Reboot the system:

   ```bash
   reboot
   ```

## Removing the Previously Installed Token

To delete the token, use the command:

```bash
/opt/clore-hosting/clore.sh --reset
```

The file containing the token is located at:

```
/opt/clore-hosting/client/auth
```


# Installing Software

## Server Requirements

The server (or rig – these terms are nearly interchangeable in this context) must be equipped with NVIDIA GPUs, as AMD is currently not supported. The minimum required disk space is 32 GB; for reliability, it's recommended to use an SSD instead of a flash drive. A minimum of 8 GB of RAM is required, but 16 GB will provide greater stability. As for the CPU, the system can work with a Celeron on a 1151 socket, but for more efficient performance, consider using a CPU like the i7-6700.

Before proceeding, it is highly recommended to disable any overclocking, including the Power Limit (PL), and reset the GPUs to factory settings. Afterward, stress-test the system for stability by, for example, testing the GPUs using the kawpow algorithm and loading the CPU. Monitor temperatures and ensure everything is running stably.

If the system operates stably and temperatures are within a safe range, continue to the next step in the instructions. If temperatures are too high or errors occur, address these issues first – for example, by improving cooling or troubleshooting – and ensure stable operation before proceeding.

## Recommended Drivers & CUDA (HiveOS)

HiveOS includes its own driver management via the `nvidia-driver-update` command. For best compatibility with Clore.ai workloads (especially ML/AI), use the following recommended versions:

### NVIDIA Drivers

| Branch          | Version    | CUDA Support    | Recommended For                                      |
| --------------- | ---------- | --------------- | ---------------------------------------------------- |
| **R580 (LTSB)** | 580.126.18 | Up to CUDA 12.8 | Most GPUs — stable, long-term support until Aug 2028 |
| **R590**        | 590.48.01  | Up to CUDA 13.1 | RTX 50 series (5090/5080), latest features           |

To install a specific version in HiveOS:

```bash
nvidia-driver-update 580.126.18 --force
```

For RTX 50 series GPUs:

```bash
nvidia-driver-update 590.48.01 --force
```

> **Important:** Do not use `nvidia-driver-update --force` without specifying a version — it may install an older default driver that doesn't support modern CUDA workloads.

### CUDA Toolkit Compatibility

Most renters use Docker images that include their own CUDA runtime, so hosts typically don't need to install the CUDA Toolkit manually. However, **the host's NVIDIA driver must support the CUDA version required by the renter's workload.**

| CUDA Version | Minimum Driver | Status                              |
| ------------ | -------------- | ----------------------------------- |
| CUDA 12.4    | R550+          | Widely used in ML ecosystem         |
| CUDA 12.8    | R570+          | Latest stable 12.x branch           |
| CUDA 13.1    | R590+          | Latest, optimized for RTX 50 series |

**Recommendation:** Install R580 LTSB (580.126.18) for broad compatibility with CUDA 12.x workloads. If you host RTX 50 series GPUs, use R590 (590.48.01) for full CUDA 13.x support.

## Registering and Adding the Server

### 1. Go to the [website](http://clore.ai/), register, log in, and navigate to the marketplace:

<figure><img src="https://img1.teletype.in/files/0e/86/0e86de72-544d-48d8-8d82-cf120e516a81.png" alt=""><figcaption></figcaption></figure>

### 2. **Adding a Server:** There are two ways to add a server:

**Method 1:** Go to the "My Servers" section and click the "+Add Server" button. Enter the server's name and click "Next."

<figure><img src="https://img4.teletype.in/files/f7/8e/f78e0a46-06fa-4a5d-b429-f21b78eafb6c.png" alt=""><figcaption></figcaption></figure>

After adding, the server will be marked with a red circle, meaning it is inactive. We will activate it later, but for now, click on the created server to obtain a key – you will need it later.

<figure><img src="https://img4.teletype.in/files/36/ae/36aeeab8-98e0-4fea-81e9-d731d5211df2.png" alt=""><figcaption></figcaption></figure>

### 3. HiveOS Setup:

Choose the rig and open Shell. For those who rarely use HiveOS, images have been added below for clarity.

<figure><img src="https://img1.teletype.in/files/45/06/4506318a-02cf-4de5-b5c8-bcf44df412ea.png" alt=""><figcaption></figcaption></figure>

<figure><img src="https://img3.teletype.in/files/e7/8e/e78e68e8-04da-4f84-89ab-546426d5f761.png" alt=""><figcaption></figcaption></figure>

<figure><img src="https://telegra.ph/file/49b76dd27191faca74a44.png" alt=""><figcaption></figcaption></figure>

### 4. **HiveOS Update:** Run the command:

```bash
hive-replace -y --stable
```

#### **If HiveOS Disk Space Issues Arise:** If there's less free space on the disk than expected after installation or update (e.g., only 20 GB free on a 512 GB disk), execute the following:

* **For M.2:**

  ```bash
  growpart /dev/nvme0n1 4
  resize2fs /dev/nvme0n1p4
  ```
* **For SATA:**

  ```bash
  growpart /dev/sda 4
  resize2fs /dev/sda4
  ```

### 5. **Run Updates in Sequence:**

```bash
selfupgrade --force
apt update
apt upgrade
apt autoremove
```

### 6. **Update Necessary Drivers:**

```bash
nvidia-driver-update --force
```

> **Tip:** To install a specific recommended driver version, use:
>
> ```bash
> nvidia-driver-update 580.126.18 --force
> ```
>
> For RTX 50 series GPUs, use version `590.48.01` or later.

### 7. **Reboot the Rig:**

```bash
reboot
```

### 8. **Switch to Superuser Mode:**

```bash
sudo -i
```

### 9. **Install the Software:**

> **Tip: installer channels.** The My Servers onboarding page generates ready-to-copy one-line install and register commands with a choice of **Stable** or **Beta** channel (plus a legacy fallback). If you use those generated commands, you can skip the manual install and activation steps below.

```bash
bash <(curl -s https://gitlab.com/cloreai-public/hosting/-/raw/main/install.sh)
```

If the system reports that `git` is missing, install it with:

```bash
apt install -y git
```

Then retry the installation.

If you encounter a `gpg` error, use:

<figure><img src="https://telegra.ph/file/e2ef8c5760193ad523e20.png" alt=""><figcaption></figcaption></figure>

```bash
apt install gpg -y --allow-downgrades
```

<figure><img src="https://img3.teletype.in/files/66/1c/661c9073-cc8e-4734-aa85-cff08902d4d6.png" alt=""><figcaption></figcaption></figure>

Afterward, rerun the installation.

```
bash <(curl -s https://gitlab.com/cloreai-public/hosting/-/raw/main/install.sh)
```

### 10. **Activate the Server:**

```bash
/opt/clore-hosting/clore.sh --init-token <token>
```

Replace `<token>` with the key obtained earlier.

If an error indicates a missing folder or file, the installation likely didn’t complete correctly, and the `clore-hosting` folder was not created. In this case, repeat the installation.

### 11. **Finalize reboot:**

Restart the rig, wait a moment, and refresh the marketplace page. If everything was set up correctly, the server will be marked with a green circle.

```
reboot
```

<figure><img src="https://img2.teletype.in/files/98/9c/989c1cbd-2670-4568-b784-020af71451be.png" alt=""><figcaption></figcaption></figure>

## How to Disable All Installed Services

If you need to disable everything previously installed:

1. Disable the services:

   ```bash
   systemctl disable clore-hosting.service
   systemctl disable docker.service
   systemctl disable docker.socket
   ```
2. Reboot the system:

   ```bash
   reboot
   ```

## How to Re-enable Services

To re-enable the services:

1. Enable the services:

   ```bash
   systemctl enable clore-hosting.service
   systemctl enable docker.service
   systemctl enable docker.socket
   ```
2. Reboot the system:

   ```bash
   reboot
   ```

## Removing the Previously Installed Token

To delete the token, use the command:

```bash
/opt/clore-hosting/clore.sh --reset
```

The file containing the token is located at:

```
/opt/clore-hosting/client/auth
```


# Network Requirements

To successfully host GPU servers on Clore.ai, your network must meet these requirements.

## Minimum Requirements

| Parameter      | Requirement                          |
| -------------- | ------------------------------------ |
| Download speed | 100 Mbps minimum                     |
| Upload speed   | 100 Mbps minimum                     |
| Latency        | < 100ms to major regions             |
| IP type        | Static or Dynamic (static preferred) |

> **Note:** Higher bandwidth results in better server ratings and more rentals.

## Required Ports

The following ports must be accessible from the internet:

| Port      | Protocol | Purpose                          |
| --------- | -------- | -------------------------------- |
| 22        | TCP      | SSH access (or custom SSH port)  |
| 8080      | TCP      | Jupyter Notebook (if enabled)    |
| 3000-4000 | TCP      | Application ports (configurable) |
| Custom    | TCP/UDP  | As defined in server settings    |

## Firewall Configuration

### UFW (Ubuntu)

```bash
# Allow SSH
sudo ufw allow 22/tcp

# Allow Jupyter
sudo ufw allow 8080/tcp

# Allow port range for applications
sudo ufw allow 3000:4000/tcp

# Enable firewall
sudo ufw enable
```

### iptables

```bash
# Allow SSH
iptables -A INPUT -p tcp --dport 22 -j ACCEPT

# Allow Jupyter
iptables -A INPUT -p tcp --dport 8080 -j ACCEPT

# Allow port range
iptables -A INPUT -p tcp --dport 3000:4000 -j ACCEPT

# Save rules
iptables-save > /etc/iptables/rules.v4
```

## NAT / Port Forwarding

If your server is behind a router:

1. **Access router admin panel** (usually 192.168.1.1)
2. **Find Port Forwarding section**
3. **Forward required ports** to your server's internal IP
4. **Set static internal IP** for your server

### Example Port Forwarding Rules

| External Port | Internal IP   | Internal Port | Protocol |
| ------------- | ------------- | ------------- | -------- |
| 22022         | 192.168.1.100 | 22            | TCP      |
| 8080          | 192.168.1.100 | 8080          | TCP      |

## Static vs Dynamic IP

### Static IP (Recommended)

* Consistent connection for renters
* Better for DNS and bookmarks
* Higher server reliability rating

### Dynamic IP

* Works but requires DDNS service
* IP changes may briefly interrupt rentals
* Configure DDNS client on your server:

```bash
# Example with ddclient
sudo apt install ddclient
sudo nano /etc/ddclient.conf
```

## Bandwidth Considerations

### Impact on Earnings

| Speed    | Impact                                |
| -------- | ------------------------------------- |
| 100 Mbps | Minimum - basic rentals               |
| 500 Mbps | Good - suitable for most ML workloads |
| 1 Gbps+  | Excellent - attracts premium renters  |

### Monitoring Bandwidth

```bash
# Install monitoring tool
sudo apt install iftop

# Monitor in real-time
sudo iftop -i eth0
```

## Network Testing

### Speed Test

```bash
# Install speedtest
sudo apt install speedtest-cli

# Run test
speedtest-cli
```

### Latency Test

```bash
# Test to common regions
ping -c 10 8.8.8.8        # Google DNS
ping -c 10 1.1.1.1        # Cloudflare
```

### Port Accessibility Test

From outside your network, verify ports are open:

```bash
# Using nmap from another machine
nmap -p 22,8080 YOUR_PUBLIC_IP

# Or use online port checker
# https://www.yougetsignal.com/tools/open-ports/
```

## Troubleshooting

### Ports Not Accessible

1. Check firewall rules on server
2. Verify router port forwarding
3. Contact ISP - some block hosting
4. Try different port numbers

### Slow Connection

1. Run speed test
2. Check for bandwidth throttling
3. Consider upgrading internet plan
4. Optimize server network settings:

```bash
# Increase network buffer sizes
sudo sysctl -w net.core.rmem_max=16777216
sudo sysctl -w net.core.wmem_max=16777216
```

### Connection Drops

1. Check for IP address changes
2. Verify router stability
3. Monitor system logs: `dmesg | grep -i network`
4. Consider using ethernet instead of WiFi


# Server Management

Let’s look at the server settings page. By clicking on a created server, you will be taken to a page where you can view the server information and its settings:

<figure><img src="/files/rMA2zUdG2jmDdj9BPvCl" alt=""><figcaption></figcaption></figure>

The page displays the server status, rental status (whether it’s rented or not), remaining rental time (if the server is rented), Backend version, and the MFP value (more on this below).

Server statuses can be as follows: “Operating Normally,” “CUDA Error,” “Docker Error.” There may also be a “Starting Up” status, which temporarily indicates that the system is determining which of the primary statuses is applicable.

* CUDA Error indicates problems with drivers, GPU disconnections, often caused by riser faults or insufficient power supply.
* Docker Error is associated with SSD issues or too slow an internet connection, where Docker fails to deploy within the set time.

For both errors (CUDA and Docker), a simple reboot sometimes helps. If this doesn’t work, it may be necessary to check and replace risers, the power supply, or the SSD. In case of a Docker Error, it is recommended to reinstall the system, starting with Hive OS. If the problem persists, test the SSD with specialized software, reflash it, or replace it with another.

## Rental Settings (Price, Duration)

Below the server information are the rental settings. Here you can enable or disable rental. When enabled, the server becomes available on the marketplace; when disabled, it disappears from the rental list, which is convenient for maintenance or updates. You can also set the maximum rental duration and set the rental price (the rental price is specified for the entire server per 24 hours):

<figure><img src="/files/QvGMatzpGujgNkgtWxMq" alt=""><figcaption></figcaption></figure>

### There are four types of pricing available for rental:

1. On-demand (rental price) in BTC,
2. Spot in BTC,
3. Rental price in CLORE,
4. Spot in CLORE.

On-demand rental allows the renter to use the server for the entire specified maximum period, as long as they have funds or do not end the rental themselves. Spot rental allows the rental to be overtaken by another user who offers a higher bid. In this case, the server can also be rented at the rental price by another user.

There are also two checkboxes:

* Enable BTC — allows renting the server for BTC. When disabled, BTC will be unavailable.
* **E**nable CLORE — allows renting the server for CLORE. When disabled, CLORE will be unavailable.

It is recommended to set the desired rental parameters (duration and prices, even in fields not activated by the checkbox) before enabling the rental (Available to rent), and then enable rental and click Apply. Be sure to refresh the page and check all settings.

There is also a checkbox, Calculate based on USD. When enabled, you can set the rental and Spot price in USD, and the BTC and CLORE prices will be automatically recalculated according to the current exchange rate (every 2 hours). This keeps the rental prices up-to-date with the USD exchange rate.

<figure><img src="/files/t8p54IlOnn5m1ValWEaB" alt=""><figcaption></figcaption></figure>

## Additional Price Setting Methods

On the server list page, there is a Set GPU pricing button, which allows you to specify prices separately for each GPU. The server price will be calculated automatically based on the number of GPUs.

<figure><img src="/files/vmNVmglvfL0AMd3xc0lE" alt=""><figcaption></figcaption></figure>

You can set On-demand and Spot prices in BTC and CLORE or enable Calculate based on USD to automatically recalculate prices in BTC and CLORE at the current rate.

Example: You have a server with 10 3070 GPUs, and you set the price for one GPU at $1. The total server price will automatically be set to $10 (10 GPUs \* $1 = $10).

<figure><img src="/files/5tIXBU3D0A1SlZHjvLOW" alt=""><figcaption></figcaption></figure>

Another way to set the rental price is described in the following section.

## Rental Graph, Server Rating, and Auto-Price

If you scroll further down the server settings page, you’ll find the server rating displayed as an average rating (with the number of ratings in brackets). Stars visually represent the server’s average rating.

<figure><img src="/files/yF3bUto7Dwjgz7unkQlk" alt=""><figcaption></figcaption></figure>

There is also a graph that shows the server’s rental duration at different periods, including rental prices. Various rental methods (Spot BTC, Rent BTC, Spot CLORE, and Rent CLORE) are marked with different colors, which can be seen in the legend below the graph.

Please note that the rental method designations below the graph are clickable. Clicking them toggles the graph display. Sometimes, the server is “rented” but does not display on the graph because the previous rental was, for example, in Rent CLORE, and the current one is in Spot BTC. In this case, to see the current status, switch the graph to the desired mode.

Another rental pricing method can also be activated here, as mentioned earlier. This feature is enabled by checking the box, after which you can set a multiplier for Rent and Spot prices. Then the hashrate and idle mining coin (when the server is not rented) from Hive OS will be applied. The system will assess your idle mining profitability, apply the specified multiplier, and set the resulting price for rental.

<figure><img src="/files/m195WiGn745PFrcNoqTS" alt=""><figcaption></figcaption></figure>

Example: If idle mining profitability is $4, and you set a multiplier of 1.2 for the Rent price, the final Rent price in BTC and CLORE will be $4.8 (1.2 \* $4 = $4.8) according to the current rates.

This method works under the following conditions: the server must mine a coin through Hive OS, and this coin must be known and available in mining calculators. For example, if mining Qubic in idle mode, this method will not work since Qubic is not available in calculators.

<figure><img src="/files/O9kwKqKKLqEcgxCvB1Tp" alt=""><figcaption></figcaption></figure>

Important: if you enable the auto-price function and click Apply, the rental will be automatically enabled, even if it was previously disabled, and the server will become available on the marketplace. The reverse is also true: if profitability cannot be assessed for some reason, the server will be hidden from the marketplace and will be unavailable for rental.

## Partial GPU Rental (multi-GPU rigs)

If your server has several identical GPUs and runs an up-to-date hosting backend, renters can rent only part of it: they pick a GPU count and pay a proportional share of your server price. Several renters can then share the machine, each in an isolated container with dedicated GPUs.

* Partial rental is **enabled by default** on eligible servers. You can opt out per server with the **Allow partial GPU rental** checkbox in the rental settings.
* While at least one partial order is active, overclocking changes are blocked (Apply warns instead of failing silently).
* Mixed-model rigs are not eligible and are always rented as a whole.

{% content-ref url="/pages/0zyrYnLQ5gtl9Y3fO1G5" %}
[Partial GPU Rental](/for-renters/partial-gpu-rental)
{% endcontent-ref %}

## Live stats

The My Servers list includes a **Live stats** column (enabled by default): a compact real-time view of per-GPU activity plus CPU and RAM load for each server, so you can spot problems without opening every server page.

The list columns are configurable and your selection persists. Rented servers show the current rental price and currency directly in the list and in the server header. Server details also include a status guide and an uptime chart for diagnostics.

## Overclocking

<figure><img src="/files/sYcdoHdvTPsKXURwYI1v" alt=""><figcaption></figcaption></figure>

Here, you can see the default overclocking profile, which can be edited. Note: any changes to the default overclocking profile will move the server to the Power Efficient marketplace. Currently, the marketplace is divided into two segments: Mainline and Power Efficient.

<figure><img src="/files/TdnRqTAMrxl9mgyRUyN1" alt=""><figcaption></figcaption></figure>

If you modify the default profile and want to return the server to the Mainline marketplace, click the RESET button (only this button, do not click Apply). After a while, the server will reappear in the Mainline marketplace.

<figure><img src="/files/eLHhLAOi2v5VlbJB6sZq" alt=""><figcaption></figcaption></figure>

You can also add additional overclocking profiles and customize them as you wish. For each profile, you can specify an additional fee for its use or leave it blank to avoid additional charges for the renter. If the renter does not select any available profile (or none are available), the server will run with the default overclocking profile.

Important: specify offset values half as large as in Hive OS for both memory and core.

## Internet Speed Test

If the internet speed is displayed lower than actual after adding the server, or if you have upgraded to a higher provider plan, you can test the server’s speed. This feature is available once every two weeks.

<figure><img src="/files/rBpslEv5e81xqDI7ns3Q" alt=""><figcaption></figcaption></figure>

## Idle Mining (background job)

To avoid server idle time between rentals, you can configure background mining. Three setup options are available:

Option 1: Select Docker image: cloreai/mining, set the algorithm, overclocking profile, and fill in the necessary data fields.

<figure><img src="/files/aX2X0veCsA5MB5W4GhZB" alt=""><figcaption></figcaption></figure>

Please note that the selection of available algorithms is limited, and only the BZ miner is supported.

Option 2: You can configure any other miner and algorithm using a script. Enable the corresponding switch and paste the script text in the field. Scripts for launching other miners and information about them are available

<figure><img src="/files/c6ddl7vNBvBHkw8a7e2M" alt=""><figcaption></figcaption></figure>

Option 3: Launch a Flightsheet through Hive OS.

<figure><img src="/files/Uq6RCLrmWYFMLMmzCq2V" alt=""><figcaption></figcaption></figure>

For this to work, be sure to select Docker image: HiveOS Flightsheet — without this, no flightsheet will start.

***

## Related Pages

{% content-ref url="/pages/YqdpYakM2IqKMLL4kuAR" %}
[How to Price Your Servers](/for-hosts/server-settings/how-to-price-your-servers-for-rental)
{% endcontent-ref %}

{% content-ref url="/pages/oEJbVy1ap9ut7R2BNoXV" %}
[MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics)
{% endcontent-ref %}

{% content-ref url="<https://github.com/defiocean/clore-docs/tree/main/gpu-marketplace/clore.ai-reward-distribution-system.md>" %}
<https://github.com/defiocean/clore-docs/tree/main/gpu-marketplace/clore.ai-reward-distribution-system.md>
{% endcontent-ref %}


# How to Price Your Servers

The minimum profitability that your GPUs can bring is the mining income, calculated using a mining calculator, such as hashrate.no. This serves as a baseline for setting a rental price. Let’s consider an example for a mining rig with 8 x 3070 GPUs.

<figure><img src="/files/p71R6QBAMiCyk5JR0V1M" alt=""><figcaption></figcaption></figure>

#### Determining Rig Profitability:

Go to the hashrate.no website and find the profitability for a single 3070 card, for example, $0.37. For a rig with 8 cards, the profitability would be:

```
0,37 * 8 = 2,96$
```

This is the current profitability for an 8 x 3070 rig, which can be used as a starting point when setting the rental price.

### Setting the Price on the Server:

As mentioned earlier, you can directly specify a price in dollars in the server settings by enabling the corresponding feature. The price in BTC and CLORE will then be automatically recalculated and updated based on the exchange rate. If you prefer to set fixed values for BTC and CLORE, manually convert the calculated dollar amount.

### Configuring in the Clore Dashboard:

Log in to your dashboard on clore.ai, and insert the calculated values into the On-demand pricing field. Set the Spot price slightly lower, then enable the rental.

Note: You can set the price above or below the baseline, depending on demand. Mining profitability in this case serves as a guideline.

> **Tip: Price Coach.** The platform shows built-in pricing guidance based on live marketplace data for similar hardware, so you can see how your price compares with the market before saving it.

> **Simplified pricing.** The pricing form offers a simplified USD/day view: set a single USD number and the BTC/CLORE prices follow the exchange rate automatically. Per-day prices are capped at 0.015 BTC / 600,000 CLORE / $1,000.

<figure><img src="/files/zTJ5SuPVjG9S8epsWjye" alt=""><figcaption></figcaption></figure>

### Setting Prices for GPU Servers:

If you not only have a mining rig but also a powerful GPU server capable of handling AI tasks, you can set the Spot price based on mining profitability, while the Rent price can be based on the anticipated cost of services.

CLORE Coin Bonus:

In addition to rental payments (e.g., $2.96), renters receive a bonus in CLORE coins. You can find the exact amount using the calculator on the website.

Remember to account for the marketplace fee. The total fee (10% on-demand, 2.5% Spot) is split 50/50 between renter and host, so your host share is 5% on-demand and 1.25% on Spot, plus the [extra hoster fee](/for-hosts/host-fees) on non-CLORE payments. For a day's on-demand rental at $2.96 this gives:

```
$2.812 + CLORE rewards on top
```

> 📚 See also: [GPU Mining Profitability — Mining vs Renting on Clore.ai](https://blog.clore.ai/gpu-mining-profitability/)


# Server Rating System

Understanding how server ratings work on Clore.ai and how to improve yours.

## What is Server Rating?

Server rating is a score that indicates the reliability and quality of your server. Higher ratings lead to:

* Better visibility in marketplace
* More rentals
* Higher trust from renters

## Rating Components

Your server rating is based on several factors:

### 1. Uptime

The most important factor - how often your server is online and available.

| Uptime | Impact                       |
| ------ | ---------------------------- |
| 99%+   | Excellent - top rankings     |
| 95-99% | Good                         |
| 90-95% | Average                      |
| <90%   | Poor - may affect visibility |

### 2. Successful Rentals

Completed rentals without issues improve your rating.

* ✅ Rental completed normally
* ✅ No complaints from renter
* ✅ Server performed as advertised

### 3. Response Time

How quickly your server responds to connections and requests.

### 4. Specs Accuracy

Server performs according to listed specifications:

* GPU model matches
* RAM amount correct
* Storage available as listed
* Network speed as advertised

### 5. Renter Reviews

Feedback from users who rented your server.

## How to Improve Your Rating

### Maximize Uptime

```bash
# Monitor your server status
systemctl status clore-hosting

# Enable auto-start on boot
systemctl enable clore-hosting

# Check for issues
journalctl -u clore-hosting -f
```

**Tips:**

* Use stable power supply (UPS recommended)
* Stable internet connection
* Monitor server health regularly
* Set up alerts for downtime

### Ensure Accurate Specs

1. Verify GPU is detected: `nvidia-smi`
2. Check RAM: `free -h`
3. Check storage: `df -h`
4. Test network speed: `speedtest-cli`

### Maintain Good Performance

* Keep drivers updated
* Monitor temperatures
* Clear unused Docker images
* Maintain sufficient free disk space

### Respond to Issues Quickly

* Monitor your Telegram notifications
* Respond to renter messages promptly
* Fix reported issues fast

## Viewing Your Rating

1. Go to **My Servers**
2. Click on your server
3. View rating and statistics

### Statistics Available

* Total uptime percentage
* Number of rentals
* Average rental duration
* Earnings history

## Rating Impact on Earnings

Higher-rated servers typically:

* Get rented more often
* Can command slightly higher prices
* Appear higher in search results
* Build repeat customers

## Common Rating Issues

### Low Uptime

**Causes:**

* Unstable internet
* Power outages
* Server crashes
* Clore software not running

**Solutions:**

* Use UPS for power backup
* Get business-grade internet
* Monitor and restart services automatically
* Check logs for crash reasons

### Spec Mismatch

**Causes:**

* GPU not detected properly
* RAM reporting incorrectly
* Storage mounted wrong

**Solutions:**

* Reinstall NVIDIA drivers
* Check hardware connections
* Verify Clore hosting configuration

### Slow Response

**Causes:**

* Network issues
* High server load
* Firewall blocking

**Solutions:**

* Check network configuration
* Monitor server resources
* Verify port forwarding

## Best Practices

1. **Monitor 24/7** - Use monitoring tools or Telegram bot
2. **Automate restarts** - Set up automatic service recovery
3. **Keep updated** - Update Clore hosting software regularly
4. **Honest listings** - Only list specs you can deliver
5. **Good communication** - Respond to renters quickly
6. **Regular maintenance** - Schedule maintenance during low-demand times


# Telegram bot

The Telegram Bot Notifies You of the Start and End of Rentals, and Also Provides Information About Your Servers — a Useful Tool for Convenient Management.

To connect the bot, follow these steps:

1\. Go to your CLORE account, open the Settings section, and click Link account.

<figure><img src="/files/bBjaevkryJ89iFZzO9uB" alt=""><figcaption></figcaption></figure>

2. The bot will automatically open in Telegram. Send it the /start command.

<figure><img src="/files/GWqMvdLWcNknEsouUf67" alt=""><figcaption></figcaption></figure>

3\. Copy the provided link and paste it into the browser where you are logged in on the CLORE website.

<figure><img src="/files/o4MTBXbXulDG526qM7r4" alt=""><figcaption></figcaption></figure>

Done, the bot is connected! Now you will receive notifications about the status of your servers directly in Telegram.


# Host Fee Guide

Understanding how fees affect your earnings as a Clore.ai host.

## Order Creation Fee — 10-Minute Grace Period

Each order has a one-time **creation fee** charged to the renter. Previously, this fee was charged immediately when the order was created. Now, it is only charged **after the 10-minute mark**:

* The renter can deploy, inspect, and test the server for up to 10 minutes
* If the renter cancels within 10 minutes — **no creation fee is charged**
* After 10 minutes, the creation fee is charged **once**

This encourages renters to try servers risk-free, which means **more orders for you**.

***

## How You Get Paid

When a renter uses your server (after the 10-minute trial), the settlement is calculated in CLORE. Your earnings depend on three factors:

1. **Your listed price** (you set this)
2. **Base marketplace fee** (always applies)
3. **Extra hoster fee** (only for non-CLORE payments)

***

## Base Fee (All Currencies)

A base marketplace fee is deducted from every rental:

| Order Type | Total Base Fee | Your Share (Hoster) | Renter Pays |
| ---------- | -------------- | ------------------- | ----------- |
| On-Demand  | 10%            | 5%                  | 5%          |
| Spot       | 2.5%           | 1.25%               | 1.25%       |

The base fee is **split 50/50** between you and the renter.

### Reduce Base Fee with PoH

Hold CLORE in [Proof of Holding](/proof-of-holding/overview) to reduce your base fee by up to **50%**:

| CLORE in PoH | Your Base Fee Reduction |
| ------------ | ----------------------- |
| 200,000      | 5%                      |
| 500,000      | 13%                     |
| 1,000,000    | 25%                     |
| 2,000,000+   | 50% (maximum)           |

**Example:** On-demand with 1M CLORE in PoH → your base fee drops from 5% to 3.75%.

***

## Extra Hoster Fee (Non-CLORE Only)

When renters pay with **BTC, USDT, or USDC**, an additional **15% extra hoster fee** is applied to your earnings:

| Payment Currency | Extra Hoster Fee |
| ---------------- | ---------------- |
| **CLORE**        | **0%**           |
| Bitcoin (BTC)    | 15%              |
| USDT / USDC      | 15%              |

> **This is the biggest fee difference.** A renter paying with BTC costs you 15% more in fees compared to CLORE.

### Reduce Extra Fee with MFP Lock

Lock CLORE in [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) tiers to reduce or eliminate the extra hoster fee:

| MFP Tier     | Max Reduction | Lock Capacity | Unlock Period |
| ------------ | ------------- | ------------- | ------------- |
| Tier 1       | 30%           | MFP × 1,000   | 7 days        |
| Tier 2       | 30%           | MFP × 4,000   | 14 days       |
| Tier 3       | 40%           | MFP × 2,000   | 28 days       |
| **All Full** | **100%**      | MFP × 7,000   | —             |

When all three tiers are fully locked, the 15% extra fee is **completely eliminated**.

***

## Earnings Comparison

**100 CLORE/day listed price, On-Demand rental:**

| Scenario         | Payment  | Base Fee | Extra Fee | You Receive     |
| ---------------- | -------- | -------- | --------- | --------------- |
| No PoH, No MFP   | CLORE    | 5%       | 0%        | **95 CLORE**    |
| No PoH, No MFP   | BTC/USDT | 5%       | 15%       | **80 CLORE**    |
| 1M PoH, No MFP   | BTC/USDT | 3.75%    | 15%       | **81.25 CLORE** |
| 1M PoH, Full MFP | BTC/USDT | 3.75%    | 0%        | **96.25 CLORE** |
| 2M PoH, Full MFP | CLORE    | 2.5%     | 0%        | **97.5 CLORE**  |

***

## Quick Tips for Hosts

1. **Accept CLORE payments** — zero extra fees
2. **Stake in PoH** — reduces base fees for all currencies
3. **Lock MFP** — eliminates extra fees on BTC/USDT/USDC orders
4. **Price competitively** — check similar GPUs on the [Marketplace](https://clore.ai/marketplace)
5. **Maintain high uptime** — affects your [Server Rating](/for-hosts/server-settings/server-rating)

***

## Related

{% content-ref url="/pages/3Fy2rz08NzsQhx7vBl4g" %}
[Fee Structure](/for-renters/fee-structure)
{% endcontent-ref %}

{% content-ref url="/pages/oEJbVy1ap9ut7R2BNoXV" %}
[MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics)
{% endcontent-ref %}

{% content-ref url="/pages/Db4Ut4ssjHaKEmdt6plm" %}
[Overview](/proof-of-holding/overview)
{% endcontent-ref %}


# MFP Lock

### 1. What is MFP and how is it calculated

**MFP (Maximum Fair Price)** is an internal quality score for every server on the Clore.ai network.\
The algorithm considers **all key hardware characteristics**: GPU model and count, PCI Express version and lane width, internet-link speed, CPU model, and many other metrics.\
The more powerful and stable the configuration, the higher the MFP.\
The current upper limit is **200**.

<figure><img src="/files/OHxjKqCLCpsKGq4BJ7zi" alt=""><figcaption></figcaption></figure>

***

### 2. Why MFP Lock is needed

A score alone is not enough-the network must see that the owner **believes in the server** and is willing to back its quality with coins.\
A host therefore **locks** CLORE in MFP Lock, confirming the server's MFP and unlocking access to a share of each block reward.

***

### 3. Reward flows in a block

| Income source   | Share of emission | Explanation                                             |
| --------------- | ----------------- | ------------------------------------------------------- |
| **PoH Staking** | 20 %              | Users without servers can simply stake their coins here |
| **MFP Lock**    | 20 %              | Split equally: **Tier 1 (10 %)** and **Tier 2 (10 %)**  |

With a reward of 313 CLORE per block and 1 440 blocks per day, each basket (Tier 1 and Tier 2) receives **45 072 CLORE per day**.

***

### 4. How much must be locked

*Base rule: **1 MFP = 1 000 CLORE***.

| Tier       | Multiplier | Stake range     | Example at MFP = 20                     |
| ---------- | ---------- | --------------- | --------------------------------------- |
| **Tier 1** | 1×         | ≤ 100 % MFP     | ≤ 20 000 CLORE                          |
| **Tier 2** | 4×         | 100 - 500 % MFP | up to 80 000 CLORE more (total 100 000) |
| **Tier 3** | 2×         | 500 - 700 % MFP | up to 40 000 CLORE more (total 140 000) |

**Maximum total lock: MFP × 7 000 CLORE**

***

### 5. Daily-reward formula and limits

For each basket *k* ∈ {T1, T2}:

<figure><img src="/files/n5AVXsz3RSIuP3ZKKJIx" alt=""><figcaption></figcaption></figure>

#### Rent-price ceiling

A server cannot receive from **each** tier more than its daily rental price.\
For a rent of **300 CLORE / day** the maximums are:

| Source                  | Ceiling             |
| ----------------------- | ------------------- |
| MFP Lock Tier 1         | 300 CLORE           |
| MFP Lock Tier 2         | 300 CLORE           |
| **Total from MFP Lock** | **600 CLORE / day** |

Together with rent, a server rented out for 300 CLORE can earn **up to 900 CLORE per day**.

#### Burning surplus

Anything above these limits is automatically sent to a burn wallet.

***

### 5.1 Extra fee reduction through MFP Lock

In addition to emission rewards, MFP Lock provides a powerful benefit for hosts: **reduction of extra hoster fees** on non-CLORE currency orders.

When renters pay with Bitcoin or other non-CLORE currencies, an extra hoster fee (e.g., 15%) is deducted from the host's earnings. By locking CLORE across all three tiers, hosts can **reduce this extra fee to 0%**.

#### How the reduction works

Each tier contributes proportionally to how full it is:

| Tier       | Max Reduction | Capacity          |
| ---------- | ------------- | ----------------- |
| **Tier 1** | 30%           | MFP × 1 000 CLORE |
| **Tier 2** | 30%           | MFP × 4 000 CLORE |
| **Tier 3** | 40%           | MFP × 2 000 CLORE |
| **Total**  | **100%**      | MFP × 7 000 CLORE |

**Formula per tier:**

```
tier_reduction = round(fill_ratio × max_reduction_percent)
```

**Combined reduction:**

```
total_reduction = min(100, T1_reduction + T2_reduction + T3_reduction)
```

The reductions are **additive** - when all three tiers are fully locked, the total reduction reaches **100%**, eliminating the extra hoster fee entirely.

#### Example

Server with MFP = 20, Bitcoin order with 15% extra hoster fee:

| Lock State     | T1 Reduction | T2 Reduction | T3 Reduction | Total | Extra Fee |
| -------------- | ------------ | ------------ | ------------ | ----- | --------- |
| Empty          | 0%           | 0%           | 0%           | 0%    | 15%       |
| T1 full only   | 30%          | 0%           | 0%           | 30%   | 10.5%     |
| T1 + T2 full   | 30%          | 30%          | 0%           | 60%   | 6%        |
| All tiers full | 30%          | 30%          | 40%          | 100%  | **0%**    |
| T1 half only   | 15%          | 0%           | 0%           | 15%   | 12.75%    |

> **Key insight:** Even partially filling tiers provides proportional benefit. A host with just Tier 1 fully locked already saves 30% on extra fees.

#### What fees are NOT affected

* **Base marketplace fees** (10% on-demand, 2.5% spot) - these are reduced by [PoH](/proof-of-holding/overview), not MFP Lock
* **Renter extra fees** - only hoster extra fees are reduced
* **CLORE currency orders** - CLORE has 0% extra fee, so the reduction has no effect

***

### 6. The "Lock → Unlock" process

* **Lock** - The host locks the desired amount; 24 h later the server starts receiving rewards.
* **Unlock (per-tier).** Each tier has its own unlock period:

| Tier   | Unlock Period | What is returned  |
| ------ | ------------- | ----------------- |
| Tier 1 | 7 days        | Full Tier 1 stake |
| Tier 2 | 14 days       | Full Tier 2 stake |
| Tier 3 | 28 days       | Full Tier 3 stake |

When **"Unlock All"** is pressed, each tier's unlock timer starts simultaneously. The funds become available in stages:

* **Day 7** — Tier 1 stake returned
* **Day 14** — Tier 2 stake returned
* **Day 28** — Tier 3 stake returned

After each waiting period, the funds are automatically moved back to the system wallet.

> **Note:** You can only unlock tiers that have passed their waiting period. For example, if you locked 10 days ago, you can unlock Tier 1 (7 days) but must wait for Tier 2 (14 days) and Tier 3 (28 days).

***

### 7. Reward examples

*(all examples use a server rented out for 300 CLORE / day)*

| Scenario                                            | T1 pool share | T1 payout | T2 payout | **Total** | **Burned**               |
| --------------------------------------------------- | ------------- | --------- | --------- | --------- | ------------------------ |
| **A.** You are alone, Stake = 65 000 (T1 = 100 %)   | 100 %         | 300       | 0         | **600**   | 45 072 - 300 = 44 772    |
| **B.** 100 hosts, each 65 000                       | 1 %           | 300 (max) | 0         | **600**   | 45 072 - 30 000 = 15 072 |
| **C.** 99 × 65 000 and you 1 000                    | 0.01538 %     | 6.93      | 0         | **≈ 307** | pool burns ≈ 15 359      |
| **D.** You are alone, Stake = 325 000 (T1 + T2 max) | 100 %         | 300       | 300       | **900**   | 90 144                   |

#### Large scenario: **500 servers (MFP = 30), all rented at 300 CLORE / day**

* **150** servers - full Tier 1 (30 000 CLORE)
* **150** servers - half Tier 1 (15 000 CLORE)
* **150** servers - full Tier 1 + half Tier 2 (30 000 + 60 000 CLORE)
* **50** servers - full Tier 1 + full Tier 2 (30 000 + 120 000 CLORE)

| Group (count)        | Stake T1 / server | Stake T2 / server | T1 payout / server | T2 payout / server\* | **Total / server†** |
| -------------------- | ----------------- | ----------------- | ------------------ | -------------------- | ------------------- |
| Full T1 (150)        | 30 000            | -                 | **106.05**         | 0                    | **≈ 406**           |
| Half T1 (150)        | 15 000            | -                 | **53.03**          | 0                    | **≈ 353**           |
| Full T1 + ½ T2 (150) | 30 000            | 60 000            | **106.05**         | **180.29**           | **≈ 586**           |
| Full T1 + T2 (50)    | 30 000            | 120 000           | **106.05**         | **300.00**           | **≈ 706**           |

\* T2 payouts are capped at 300 CLORE / day.\
† Total = (T1 payout + T2 payout + rent 300 CLORE).

**Pool outcomes (45 072 CLORE / day each):**

* **Tier 1:** fully distributed - nothing burned.
* **Tier 2:** 42 043.2 CLORE paid → **3 028.8 CLORE burned**.

Even with 500 servers and strong competition, part of Tier 2 remains unclaimed and is burned daily, maintaining token scarcity and enhancing reward value.

> **Tier 3 reminder:** Tier 3 does not earn emission rewards. Its benefit is the extra hoster fee reduction described in Section 5.1.

***

### 8. Dynamic MFP changes

If the server configuration changes (GPUs added/removed, CPU swapped, etc.), the score is recalculated:

* **MFP rises** → any already-locked free coins automatically fill Tier 1 then Tier 2.
* **MFP falls** → excess stake moves to *unused lock balance*. **unused lock balance can be unlocked after 10 days.**

***

### 9. Fund security

* **Two-factor authentication.** Always enable 2-FA-it is the first and most important defense layer.
* **Cold & hot wallets.** Most assets are stored offline; hot wallets hold only the operational minimum.
* **2.5 years incident-free.** The platform has never been hacked; in any force-majeure the company covers user losses.
* **Unlock timers.** Even if an attacker initiates a withdrawal, the minimum 7-day period gives the owner time to react.

***

### 10. Why MFP Lock is profitable

* **Transparent economics.** Higher hardware quality and higher stake mean a larger share of the pool—no hidden tricks.
* **Guaranteed break-even.** A full Tier 1 covers the rental price; a filled Tier 2 doubles that bonus.
* **Extra fee elimination.** A fully locked server (all 3 tiers) pays 0% extra hoster fee on non-CLORE orders—making BTC rentals as profitable as CLORE.
* **Clear ceiling.** A server rented for 300 CLORE can earn up to **900 CLORE/day** from emissions (T1 + T2 + rent).
* **Network self-regulation.** Surplus rewards are burned, sustaining token scarcity and price stability.
* **Security.** Per-tier unlock timers (7/14/28 days), 2-FA, and cold storage protect the stake even if an account is compromised.

**MFP Lock** is more than a "lock coins" button. It rewards investment in quality hardware, stabilises the network economy, and turns hosting on Clore.ai into a predictable, secure revenue stream.

### <mark style="color:red;">**Important information:**</mark>

If you locked MFP and attempted self-renting (renting your own server from another account):\
• Both accounts are permanently banned. Unban is impossible.\
• All related GPUs (UUID) are blacklisted.\
• All funds on the account are blocked; no refunds will be issued.

> 📚 For background on how MFP evolved, see [Understanding Maximum Fair Price (MFP) on the CLORE Blockchain](https://blog.clore.ai/understanding-maximum-fair-price-mfp-on-the-clore-blockchain-the-shift-to-usd/)


# Troubleshooting

***

**Issue:** Server is online in HiveOS but appears offline on Clore.ai.\
**Symptom:** Server shows as online in HiveOS, but remains offline on the Clore platform.

***

#### 1. Check the status of the Clore service:

Run the command:

```
systemctl status clore-hosting.service
```

If the response is`Active: inactive (dead)` — the service is not running and is not enabled on startup.

2. Enable autostart and launch the service:

```
systemctl enable clore-hosting.service
systemctl start clore-hosting.service
```

Check again:

```
systemctl status clore-hosting.service
```

If the service status is`active (running)` — the server should appear online on the Clore platform.

3. To avoid issues after reboot, make sure to add all necessary components to autostart:

```
systemctl enable clore-hosting.service
systemctl enable docker.service
systemctl enable docker.socket
```

4. Check functionality after reboot:

```
reboot
```

After reboot, the server should start automatically and appear online.


# Docker Failure

**Problem**\
Clore.ai marks the rig as *Docker Failure* and keeps it offline, even though HiveOS is running.

**Symptoms**

* A “Docker Failure” icon is shown in the Clore panel.
* In the **My Servers** section, GPUs are displayed as *0x Unknown* or the GPU count keeps changing.

***

**Cause 1: Unstable GPU or Riser**

Clore cannot initialize a GPU if it's disconnected or unstable.\
Even if HiveOS sees the GPU, Clore can’t use it → **Docker Failure**.

**Solution: Restart and Check Hardware**

1. Check the GPU or riser, make sure everything is securely connected.
2. Reboot the rig:

```
reboot
```

If the error returns after reboot, the issue is likely with the GPU, motherboard, or risers.

***

#### Cause 2: Corrupted Python Environment (Miniconda)

Clore hangs on startup if the directory `/opt/clore-hosting/miniconda-env` is corrupted.

**Solution: Remove the environment and restart**

```
sudo systemctl stop clore-hosting.service
sudo rm -rf /opt/clore-hosting/miniconda-env
sudo systemctl start clore-hosting.service
```

***

#### Cause 3: Dependency installation is stuck

If Clore doesn't start, it may be due to a frozen installation of dependencies (e.g. aiofiles, docker, etc.).

**Solution: Reinstall dependencies**

```
sudo /opt/clore-hosting/clore.sh --reinstall
```

***

#### Cause 4: Unstable Docker version installed (e.g., 28.\*)

Recommended version: **27.5.1**\
Crashes are common with Docker 28+.

**Solution: Downgrade Docker**

```bash
sudo apt install \
docker-ce=5:27.5.1-1~ubuntu.22.04~jammy \
docker-ce-cli=5:27.5.1-1~ubuntu.22.04~jammy \
containerd.io -y
```

***

#### Cause 5: Required services are not enabled on startup

After reboot, the system doesn't launch Docker and Clore Hosting → server goes offline.

**Solution: Enable services on startup**

```
sudo systemctl enable clore-hosting.service
sudo systemctl enable docker.service
sudo systemctl enable docker.socket
```

***

#### Cause 6: Driver doesn't detect GPUs (`nvidia-smi → No devices found`)

If HiveOS doesn't detect the GPU, Clore can't work with it → results in Docker Failure.

**Solution: Reinstall the driver**

```
nvidia-driver-update --force
```

***

#### If the issue persists — fully remove the server from Clore, change the token, and re-add it.

This often helps if internal configs are broken.

***

<mark style="color:blue;">**Docker Failure**</mark> <mark style="color:blue;">almost always means that</mark> <mark style="color:blue;">**Clore doesn't see the GPU**</mark><mark style="color:blue;">.</mark>\ <mark style="color:blue;">In</mark> <mark style="color:blue;">**90% of cases**</mark><mark style="color:blue;">, the cause is either a disabled service or an unstable GPU/risers.</mark>\ <mark style="color:blue;">Fix the root issue, enable services on startup — and your rig will stay online.</mark>


# How to Reinstall Drivers

Sometimes, an error may occur when attempting to reinstall drivers. To resolve it, follow these steps:

<figure><img src="/files/YAGnqXAawyiEu0lfCclB" alt=""><figcaption></figcaption></figure>

1. Put the rig in maintenance mode in HiveOS by disabling driver loading.
2. Disable services:

   ```
   systemctl disable clore-hosting.service
   systemctl disable docker.service
   systemctl disable docker.socket
   ```
3. Reboot the system:

   ```
   reboot
   ```
4. Install the required drivers:
   * To install a specific driver version, use the command:

     ```
     nvidia-driver-update 580.126.18 --force
     ```
   * Or to install a driver from a link on the NVIDIA website:

     ```
     nvidia-driver-update https://us.download.nvidia.com/XFree86/Linux-x86_64/580.126.18/NVIDIA-Linux-x86_64-580.126.18.run
     ```

**Recommended Driver Versions:**

| Branch      | Version    | Support Until | Best For                         |
| ----------- | ---------- | ------------- | -------------------------------- |
| R580 (LTSB) | 580.126.18 | Aug 2028      | Recommended stable for most GPUs |
| R590        | 590.48.01  | Dec 2026      | Latest features, RTX 50 series   |

**Note:** For RTX 5090/5080 GPUs, use R590 branch (minimum). For all other GPUs, R580 LTSB is recommended for stability. 5. Update the system (optional):

````
```
apt update
apt upgrade
apt autoremove
```
````

6\. Disable maintenance mode and re-enable all services:

````
```
systemctl enable clore-hosting.service
systemctl enable docker.service
systemctl enable docker.socket
reboot
```
````


# Reinstall While Keeping Token

**Method One (Preferred)**

The token file is located at /opt/clore-hosting/client/auth. For convenient copying, you can use a program with a file manager, such as Snowflake:

<figure><img src="/files/4R3odz4Lb4jLEOxwGorK" alt=""><figcaption></figcaption></figure>

1. Ensure that your computer and the server are on the same local network.
2. Drag the auth file onto your computer or a flash drive.
3. Reinstall Hive OS with the command:

   ```
   hive-replace -y --stable
   ```
4. Follow the installation instructions, and instead of using the command:

   ```
   /opt/clore-hosting/clore.sh --init-token <token>
   ```

   simply return the auth file back to the`/opt/clore-hosting/client/`directory
5. Reboot the server:

   ```
   reboot
   ```

**Method Two**

1. Connect to the server via SSH or open Hive Shell and run the command:

   ```
   mcedit /opt/clore-hosting/client/auth
   ```

   The auth file will open in a text editor. Copy its contents and save it on your computer, for example, in a text editor.

<figure><img src="/files/EdNtCENzb8HNbXdHtwv0" alt=""><figcaption></figcaption></figure>

2. Press F10 to exit the editor.
3. Reinstall Hive OS with the command:

   ```
   hive-replace -y --stable
   ```
4. After completing the installation, run the following commands in sequence:

   ```
   cd /opt/clore-hosting/client/
   touch auth
   mcedit auth
   ```

   A blank auth file will open in the text editor. Paste the previously saved content of the old file.
5. Press F2 to save and F10 to exit.
6. Reboot the server:

   ```
   reboot
   ```


# Docker Disk Space Cleanup

## Docker Disk Space Cleanup (Hosts)

Docker data grows over time on host machines.\
Old rentals leave behind stopped containers, unused images, volumes, and build cache.\
If you don’t clean it up, the root disk fills up and deployments start failing.

{% hint style="info" %}
Run cleanup only when the server is **not rented** and you don’t need any old container data.\
If you’re unsure, stop here and do a disk usage check first.
{% endhint %}

***

## 1) Check disk usage

### OS-level disk space (`df -h`)

This shows free space on each mounted filesystem.

```bash
df -h
```

You care mainly about `/` (root) and the partition that holds `/var/lib/docker`.

### Docker-level disk usage (`docker system df`)

This shows what Docker is storing and how much is reclaimable.

```bash
docker system df
```

If you want more detail per image/container:

```bash
docker system df -v
```

***

## 2) Complete cleanup (recommended)

This is the “reset Docker leftovers from previous rentals” command.\
It removes unused containers, images, networks, and **unused volumes**.

```bash
docker system prune -a --volumes
```

{% hint style="warning" %}
`-a` removes **all unused images**, not just “dangling” ones.\
That includes cached images you might want for faster future deployments.
{% endhint %}

***

## 3) Individual cleanup commands (more control)

Use these when you want to clean a specific category.\
They’re safer for incremental maintenance.

### Containers (stopped only)

```bash
docker container prune
```

### Images (unused)

Remove dangling layers only:

```bash
docker image prune
```

Remove all unused images (same risk as `system prune -a`):

```bash
docker image prune -a
```

### Volumes (unused)

```bash
docker volume prune
```

### Networks (unused)

```bash
docker network prune
```

***

## 4) Best practices for host maintenance

1. Clean up **between rentals** or during planned downtime.
2. Keep a safety buffer.\
   Aim for **10–20 GB** free on the root disk at all times.
3. Check Docker usage regularly:
   * `df -h`
   * `docker system df`
4. Prefer incremental cleanup first:
   * `docker container prune`
   * `docker image prune`
   * `docker volume prune`
5. Use full cleanup only when needed:
   * `docker system prune -a --volumes`
6. If disk usage keeps growing fast, investigate:
   * Large volumes created by renters.
   * Logs under `/var/lib/docker/containers/*/*.log`.
   * Build cache from frequent image rebuilds.


# Advanced

Advanced configuration and features for experienced Clore.ai hosts and enterprise users.

## What's in This Section

### [Bare Metal Servers](/for-hosts/advanced/bare-metal)

Dedicated physical servers with full root access, no virtualization, and no power limits. Ideal for AI/ML training, HPC, 3D rendering, and high-performance workloads.

* Tier 3+ data center requirements
* Hardware specifications (64+ threads, 128GB+ RAM, NVMe)
* InfiniBand/NVLink support
* SLA 99.99% with 24/7 NOC

Want to rent bare metal instead of providing it? See [Bare Metal Rental](/for-renters/bare-metal-rental) for the self-serve configurator; the requirements above are for providers listing capacity.

### [HiveOS Integration](/for-hosts/advanced/hiveos-integration)

Connect your HiveOS-managed rigs to the Clore.ai marketplace. Run mining workloads alongside GPU rentals.

### [Clore Fleet — Mass Onboarding](/for-hosts/advanced/clore-fleet-mass-onboarding)

Onboard multiple servers at once using Clore Fleet. Designed for hosts with 10+ machines who need automated provisioning and management.

### [Clore Partners](/for-hosts/advanced/clore-partners-high-volume-rental-system-for-major-clients)

High-volume rental system for data centers and major infrastructure providers. Custom pricing, dedicated support, and enterprise SLAs.


# Bare Metal

## Clore Bare Metal — Requirements and Guide

**Clore Bare Metal** are physical (non-virtualized) servers with full root access, no sharing, and no power limits. Suitable for AI/ML, HPC, 3D rendering, and any heavy workloads.

**Available GPUs (examples):** B200, H100, H200, A100, L40S, RTX 5090, RTX 4090, etc.\
**Locations (start):** USA, Japan, Hong Kong, and others\
**SLA:** Tier 3 and above data centers, target uptime **99.99%**.

***

### 1) What is Bare Metal on Clore

* You get a whole physical machine (CPU, RAM, disks, network, GPU).
* Full root access/SSH and, when available, IPMI/KVM for OS reinstallation.
* No PL limits / isolating layers — performance matches the hardware.
* Differs from container-based rentals (HiveOS/Docker) in that resources are not shared.

***

### 2) Mandatory infrastructure requirements (for providers)

**2.1 Data center**

* Minimum **Tier 3** (Uptime Institute or a recognized local equivalent).
* Documents: DC letter/certificate, redundancy description (power N+1/2N, cooling, network).
* **SLA 99.99%** with a 24/7 NOC.
* Compliance with fire safety standards; availability of emergency procedures (RPO/RTO).
* **Legal entities only.** Home/office “server rooms” are not accepted.

**2.2 Hardware base (minimum)**

* **CPU:** from 64 threads.
* **RAM:** from 128 GB (256 GB+ recommended for multi-GPU/HPC).
* **Storage:** NVMe SSD ≥ 1 TB, throughput ≥ 1 GB/s (RAID1/10 recommended for system and data).
* **Network:** ≥ 1 Gbps symmetric (10 Gbps preferred, L2/L3 redundancy, static IPv4; IPv6 is a plus).
* **GPU (tier):** L40S / H200 and above or equivalents resilient to heavy workload:\
  B200, H100, H200, A100, L40S, RTX 4090/5090 (**server A-series and data-center cards preferred**).

**2.3 High-performance interconnects (preferred)**

* **InfiniBand** (EDR/HDR/NDR) for distributed training/HPC.
* **NVLink/NVSwitch** — desirable for multi-GPU within a node.

#### 2.4 Reliability and replacement

* In case of hardware failure — **one-for-one** replacement (identical or strictly equivalent configuration) with no SLA degradation.
* Mandatory stock of spare parts / “hot” spares.

#### 2.5 Security and data hygiene

* Disk sterilization between rentals: **blkdiscard/secure erase/1-pass zero/TRIM** (logging).
* IPMI isolation, closed **mgmt** perimeter, ACL/DDoS profile.
* OS images — vetted, with up-to-date microcodes/patches, support for **NVIDIA** drivers.

***

### 3) Minimum commercial terms

* **Minimum rental term:** from **2 weeks**.
* **Pricing:** price lists competitive by geolocation (accounting for traffic/electricity/VAT costs).
* **API integration** is mandatory/desired (depending on volume) for auto-provisioning, extensions, and monitoring.

***

### 4) Software and image requirements

* **OS:** Ubuntu 22.04/24.04 LTS, Rocky/RHEL 9; on request — Windows Server (with licensing).
* **GPU stack:** NVIDIA 550.xx+ (or those recommended for specific GPUs), CUDA 12.2/12.4+.
* **Management:** SSH (required), IPMI/KVM (preferred) with temporary accounts for the renter.
* **Containerization:** Docker/Podman on request; Kubernetes — allowed if a master is provisioned within the same DC.

***

### 5) How a provider can connect to Bare Metal

1. **Application & verification:**
   * Legal entity, official contract with a Tier 3+ DC, SLA 99.99%, 24/7 NOC.
   * Document package: Tier/equivalent certificate, SLA, fire safety, redundancy scheme.
   * Acceptance tests: public IPv4, screenshot/access to IPMI (KVM), iPerf3/disk performance results.
2. **SKU catalog & pricing:**
   * Standardized cards (GPU composition, CPU threads, RAM, NVMe, network, IB/NVLink, DC/location, traffic limits).
   * Prices tied to geography. Minimum term — 2 weeks.
3. **Operational policies:**
   * Incident response time: ≤ 15 min; hardware replacement: equivalent immediately.
   * Logging of disk sterilization, closure of admin access after return, audit.
   * Monthly reports on uptime/incidents.

### 6) Network and throughput requirements

* Minimum **1 Gbps** (symmetric), preferably **10 Gbps** with redundancy.
* Public IPv4, rDNS support on request; IPv6 is desirable.
* Basic ACLs, anti-DDoS profile, dedicated **mgmt-VLAN** for IPMI.
* For **InfiniBand** — direct L2 segmentation within the rack/room and OFED availability.

***

### 7) Example workloads

* **Multi-GPU LLM training:** 8×L40S/NVLink or an IB cluster of A100/H100/H200 nodes.
* **Video rendering:** 4×RTX 4090/5090 with local NVMe cache and **10 Gbps** egress.
* **HFT/trading:** low latencies, CPU **64–128** threads, RAM **256–512 GB**, NVMe **RAID1** and **10 Gbps** network.
* **Genomics/HPC:** A100/H100 with IB **HDR/NDR**, **SLURM** / MPI support.

***

## Comparison of Standard Rental and Bare Metal

| Parameter                     | Standard rental (HiveOS/Docker)                         | Bare Metal                                                         |
| ----------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------ |
| What it is                    | Container/environment inside the host OS                | Entire physical server                                             |
| Resources (CPU/RAM/bandwidth) | Shared by scheduler; cgroup quotas, possible throttling | Exclusive; predictable CPU/RAM/bandwidth                           |
| Root/privileges               | root inside container, no BIOS access                   | Full server root; BIOS/UEFI access                                 |
| GPU drivers (CUDA/NVIDIA)     | Version defined by the host                             | You install required versions (CUDA/OFED, etc.)                    |
| GPU control                   | Passthrough with restrictions (PL/OC per host policy)   | Full PL/OC control; NVLink/NVSwitch (if present)                   |
| IPMI/KVM/Virtual Media        | No                                                      | Yes (remote console, ISO mounting)                                 |
| Storage                       | Host volumes/mounts; bandwidth may fluctuate            | Direct NVMe/RAID; stable IOPS/throughput                           |
| Network                       | Ports/NAT/shared bandwidth                              | Dedicated NIC 1–10G+; rDNS, VLAN; public IPv4                      |
| Reliability / SLA             | Depends on host; no guaranteed like-for-like swap       | DC Tier 3+, target SLA 99.99%, mandatory like-for-like replacement |
| Minimum term                  | Usually hours/days                                      | From 2 weeks                                                       |
| Cost                          | Lower                                                   | Higher (exclusive + data center)                                   |
| Time to start                 | Seconds–minutes                                         | from 1h up to 48h to start                                         |
| HPC / InfiniBand              | Usually no                                              | Recommended (InfiniBand), NVLink/NVSwitch                          |
| Best for                      | Quick tasks, tests, mining, short sessions              | AI/ML/HPC, production workloads, long projects                     |
| Requirements for provider     | Basic                                                   | Legal entity, DC Tier 3+, 24/7 NOC, regional pricing, API          |
| Security / data               | Within host policies                                    | Disk sanitization between rentals, isolated mgmt (IPMI)            |

## FAQ

**How is Bare Metal different from container rental?**\
Bare Metal is **entirely your physical machine** (CPU/RAM/Disk/Net/GPU). In container rental, resources are shared and you work in an isolated environment.

**Is IPMI required?**\
Preferred. It speeds up OS reinstallation and provides KVM access, especially for network/SSH issues.

**Can nodes be interconnected over IB?**\
Yes, InfiniBand is encouraged for distributed training/HPC. Specify the IB bandwidth/type in the SKU.

**What’s the minimum for GPUs?**\
L40S / H200 level and above, or an equivalent resilient to heavy workloads (B200, H100, A100, etc.).

**What if the server “goes down”?**\
The provider must promptly deliver an **identical replacement** with no degradation (SLA 99.99%).


# HiveOS Integration

This section shows how to add a server (rig) to Clore.ai using HiveOS: system preparation, installing Clore Hosting, applying the token, starting a flightsheet

### **Step 1. Create a flightsheet and get the token**

* Open **HiveOS → Workers → your rig → Create Flight Sheet**.
* In the **Coin** field, type **clore rentals** and select it from the list.
* Click **Get token**.

<figure><img src="/files/9WZrmmPvveV1ibQoniQK" alt=""><figcaption></figcaption></figure>

### **Step 2. Authorize on Clore.ai via HiveOS**

* After clicking **Get token**, you’ll be redirected to **Clore.ai**.
* Click **Login with HiveOS** and approve the sign-in.

<figure><img src="/files/n8MbuAPm00mRJRCCKurH" alt=""><figcaption></figcaption></figure>

### **Step 3. Set price/duration and push the token to HiveOS**

* Set the **price** and **maximum rental duration**.
* Click **Apply the token to HiveOS** — the token will be auto-filled into the **Wallet** field of your HiveOS flightsheet.
* Go back to HiveOS and continue completing the flightsheet.

<figure><img src="/files/LFfdti1qjTcKCPIxqQPy" alt=""><figcaption></figcaption></figure>

### **Step 4. Create the flightsheet**

* In HiveOS, click **Create Flight Sheet**.
* Verify the **Wallet** contains your token and **Coin: clore rentals** is selected.
* Click **Create**.

<figure><img src="/files/aS5Ke5zrukko2UuESMt0" alt=""><figcaption></figcaption></figure>

### **Step 5. Apply the flightsheet to farms**

* Open **HiveOS → Farms** and go to the workers list.
* Select the required farms/workers (use **Select all** if needed).
* Click the **Rocket / Apply** action.
* From the dropdown, choose the created **Flight Sheet** and click **Apply to selected**.

<figure><img src="/files/3zvjHslGZInuJZTMgpew" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/ervpdx2LtbJnI4007Cpj" alt=""><figcaption></figcaption></figure>

Done. Good job — now go to [**Clore.ai**](https://clore.ai/my-servers) and handle everything else in **Marketplace → My Servers**.

### Idle mining

You’ve got two options:

* If your server already shows up in “My Servers” on our site and is online — just disable the clore.rental flight sheet in HiveOS and apply your own mining flight sheet for idle mining.
* Add a second miner to our clore.rental flight sheet — edit the sheet in HiveOS, click Add miner, configure your coin/pool/miner, and save. It will run when the rig isn’t rented.

<figure><img src="/files/VKnA4brOqyyMSdAb9k44" alt=""><figcaption></figcaption></figure>


# Clore Fleet (Mass onboarding)

Instruction for Using Clore Fleet

<figure><img src="/files/Nn0Rho074G7XQSeDa7iW" alt=""><figcaption></figcaption></figure>

To add any number of servers to Clore.ai, follow these steps:

## Important: Do NOT enable Clore Fleet on machines that are already listed on the Clore platform! This will onboard the machine under a new name, cancel the original customer order, and delete the order data.

1. Navigate to [Clore Marketplace - My Servers](https://clore.ai/my-servers).
2. Find and click the “Mass Onboard” button to start the process.
3. Select Your Operating System:

Choose your operating system. In this guide, we’ll be using HiveOS as an example (the steps are similar for a standard Ubuntu setup).

4. Select One of the Two Pricing Models:

<figure><img src="/files/aBNX03mbe25mHrkfJl9x" alt=""><figcaption></figcaption></figure>

### You can choose between two pricing systems for your servers:

a) Static Pricing

<figure><img src="/files/wAa2BDcAbdpbdrnwqEmK" alt="" width="563"><figcaption></figcaption></figure>

• You can set a specific value in Clore or BTC for your server rentals.

• Alternatively, you can enable automatic price recalculation in USD. For example, if you set a price of $10, the system will automatically adjust the rental price in Clore/BTC every two hours to match $10 in USD.

b) Mining Profitability-Based Pricing

• If you run another miner alongside Clore Fleet, such as mining Clore Coin, the system will monitor your server’s profitability.

• For instance, if your server earns $5 on a specific mining algorithm and you set a multiplier of x2, your rig will be listed on the marketplace for $10.

• If profitability changes, the marketplace price will be adjusted accordingly.

### **We have 2 types of rentals**

On-Demand Rental: The server is rented at a fixed price set by the host. It remains reserved for the client until the contract is canceled or their balance is insufficient to continue the service.

Spot Market: A quick transaction at auction-based rates. The minimum price is set by the host and is typically lower than the on-demand rate. The server is not reserved, and higher bids can outbid the current offer.

<figure><img src="/files/ivBxVIpBKInXqc6FIINq" alt="" width="563"><figcaption></figcaption></figure>

5. Select the Rental Period:

Choose the period for which you plan to rent out your rig. We recommend a minimum rental duration of 150 hours.

6. Choose Additional Options:

### There are three optional settings you can configure:

a) Override offer specifics on each machine reboot or change of configuration

• With this option, you won’t be able to edit server prices through the Clore.ai website after onboarding.

• Price adjustments will only be possible by changing your unique authorization token set as a wallet address on the HiveOS platform.

b) Use Overclocking and Power Limits from HiveOS for Orders by Default

• If selected, your Power Limit and core/memory clock settings will be copied from HiveOS.

• This option is crucial for miners who do not want to lease their rigs at full Power Limit.

• Note: Enabling this option will automatically list your server on the Power Efficiency Market on the Clore Marketplace, a specialized marketplace for mining servers.

c) Keep Clore.ai running even when miner is disabled

• This option allows you to remove Clore Fleet from the flight sheet an it will still be working in a background

• It’s designed for miners using custom settings that are not officially supported by HiveOS, such as mining Aleo.

• If you enable this option and later wish to remove Clore Fleet from your server, run the following command:

```
systemctl stop clore-onboarding.service ; systemctl disable clore-onboarding.service ; systemctl stop clore-hosting.service ; systemctl disable clore-hosting.service ; systemctl stop docker.service ; systemctl disable docker.service
```

7. Start the Miner

To simplify, just copy the configuration displayed in the field on the left and use the Import from Clipboard feature when creating a new flight sheet in HiveOS.

<figure><img src="/files/tdmvi00BbfYz18mWcj4C" alt="" width="375"><figcaption><p>copy this</p></figcaption></figure>

<figure><img src="/files/dPMxPvDoia5FRircdroz" alt="" width="375"><figcaption><p>paste here</p></figcaption></figure>

8. Specify Your Unique Authorization Token as the Wallet

<figure><img src="/files/J6gxLYiHbF4tiDCVot0b" alt=""><figcaption><p>Example of</p></figcaption></figure>

**IMPORTANT: Never share your unique authorization token with anyone. This token allows others to control your servers.**

9. If you chose “Keep Clore.ai running even when the miner is disabled,” you can set any FlightSheet you want; it won’t affect Clore Fleet.


# Clore Partners

Clore Partners is a premium system designed to match top-tier hardware providers with large-scale renters such as Vast.ai and other high-demand platforms. If you’re a company or service provider that frequently rents out a significant volume of compute resources, Clore Partners offers a dedicated and reliable solution.

***

## 💰 Hosters: Why Join Clore Partners?

Clore Partners provides several key benefits for server owners (hosters):

#### ✅ 1. Dollar-Pegged Payouts

Your income is always tied to the US dollar, not the fluctuating $CLORE price.

You earn a fixed dollar value, no matter what happens on the market.

• Example: If your rig is rented for $10/day:

• At $CLORE = $0.04 → you get 250 CLORE

• At $CLORE = $0.02 → you get 500 CLORE

📌 You always get the full dollar equivalent.

#### 🎁 2. Bonus Rewards: +10% to +100% in Extra CLORE

In addition to your stable base payout, you receive bonus rewards in $CLORE on top.

• Base rewards range from +10% to +30%

• If you have **MFP Tier 1** and **Tier 2** locked, your bonus can reach up to **+100%**

Example:

If your daily income is **$10**, you can automatically receive up to **$10** extra in **$CLORE**.

💡 This means your total income can potentially double without any extra effort.

#### 🧲 3. High-Volume, Long-Term Renters

Clore Partners is tailored for integration with large platforms like Vast.ai, which means your rigs are prioritized for:

• Bulk rentals

• Longer sessions

• More consistent usage

#### 🚀 4. Priority Matching = More Earnings

Partner rigs are prioritized in the rental queue.

That means:

• Faster matching

• Higher rental frequency

• More revenue over time

#### 🧠 5. No Downsides

There is zero downside to enabling Clore Partners on your server:

• It does not prevent normal rentals

• It does not reduce your visibility

• It simply adds another layer of earning potential

## Just make sure your rig meets the minimum requirements and toggle the “Apply to Clore Partners” option — and you’re in.

This results in fewer idle hours and more total uptime.

***

## 📋 Minimum Requirements to Qualify as a Clore Partner Host

To be eligible to supply your servers to Clore Partners, please make sure your rigs meet the following conditions:

🔘 Step 1: Enable Clore Partners

• In your rig settings on the Clore dashboard, make sure to click the “Apply to work with Clore Partners” button for each server.

💻 Step 2: Hardware Requirements

✅ Mandatory Specs

• GPUs: RTX 3070, 3090, 3090 Ti, 4090, 5080, 5090, or NVIDIA A4000, A5000, A6000

• CPU: Minimum 16 threads

• RAM: Minimum 32 GB

• Network Speed: Minimum 100 Mbps

• Storage: Minimum 500 GB SSD

• CUDA Version: At least 12.2<br>

⭐ Recommended for Higher Rent Probability

• CPU: 32+ threads

• Power Limits: Full PL required (rig must be listed in the Mainline Marketplace)

• RAM: 64+ GB

• Internet: 500+ Mbps

• SSD: 1.5+ TB

• CUDA: 12.4+

• Reliability 99.5%+

• PCIe Interface: At least x4 lane with PCIe version 3.0 or higher

## 💵 Average Partner Pricing + Daily Revenue

| GPU        | Daily Revenue                         |
| ---------- | ------------------------------------- |
| RTX 3070   | $0.714 + \~$0.13 (from Clore Rewards) |
| RTX 3090   | $1.836 + \~$0.33 (from Clore Rewards) |
| RTX 3090Ti | $2.448 + \~$0.44 (from Clore Rewards) |
| RTX 4090   | $3.060 + \~$0.55 (from Clore Rewards) |
| RTX 5080   | $2.448 + \~$0.44 (from Clore Rewards) |
| RTX 5090   | $4.488 + \~$0.81 (from Clore Rewards) |
| A4000      | $1.224 + \~$0.22 (from Clore Rewards) |
| A5000      | $2.040 + \~$0.37 (from Clore Rewards) |
| A6000      | $6.120 + \~$1.10 (from Clore Rewards) |

## ⚙️ What Happens After You Enable Clore Partners

Once you click the “Apply to Clore Partners” button in your server settings:

• ✅ We will automatically set the rental prices for you.

These prices may slightly differ from the ones listed in the table, as they are adjusted to match the needs and budgets of large-scale renters. Rest assured — pricing is always fair and market-aligned.

• 🕒 The minimum rental duration will be set automatically — starting from 14 days or more.

This ensures longer rental sessions, reduced downtime, and more consistent income without frequent renter switches.

<br>

📌 Everything is handled automatically — you don’t need to configure anything manually.

## 🤝 How to Join as a Client (Large-Scale Renter)

If you run a platform or business that regularly requires $20,000/month or more in rental volume, you may qualify to become a Clore Partner renter.

To get started, simply reach out to us at:

📧 <marketing@clore.ai><br>

Our team will evaluate your needs and initiate the onboarding process to provide you with access to our partner-level server supply.


# Banned GPUs

<figure><img src="/files/oGgXfjPj9qXivdhYFpaB" alt=""><figcaption></figcaption></figure>

This page lists every graphics card that has been banned from Clore.ai for breaking our platform rules - most commonly self‑renting, hardware substitution, or other prohibited behaviour.

*We update this register continuously. Before commissioning equipment - or buying a used GPU on a marketplace - check here to ensure the card is not already banned.*

***

### 1. List of banned GPUs

#### Finding your GPU ID in Hive OS

If your rig runs Hive OS, you can retrieve the GPU ID (the **GPU‑xxxxxxxx‑…** string used in the list below) as follows:

1. Log in to your Hive OS dashboard and open the worker that hosts the card.
2. Click **Hive Shell ▶ Start** to launch a remote terminal, then open the provided link.
3. In the shell window:

   ```
   nvidia-smi --query-gpu=gpu_uuid --format=csv,noheader
   ```
4. The command prints one line per GPU, e.g.

   ```
   GPU-c441ddc8-aa02-7c35-a8e7-d56d7a170c10
   ```

   Copy the line that corresponds to the card you want to check.
5. Compare this ID with the banned‑GPU list below. If it appears, the card is currently blocked.

<mark style="color:blue;">Check the ban status of a GPU by its</mark> <mark style="color:blue;">**GPU UUID**</mark><mark style="color:blue;">:</mark> [<mark style="color:blue;">https://clore.ai/check-gpu-ban</mark>](https://clore.ai/check-gpu-ban)

> ```markdown
> gpu id
> GPU-001f446c-b829-9906-de2d-00a654343ca4
> GPU-9e3e1a55-d59e-b4db-6dcf-dfdf9d7ae890
> GPU-e86d0491-7923-aebd-a433-90d1ea0bed05
> GPU-235a4f2c-020e-7dfb-dc0d-da45bf7ac374
> GPU-458eb4fb-b3e8-83ba-db72-3aad9aaa470b
> GPU-9bb2555e-701c-c14a-5552-c86ee4eb3d30
> GPU-eb33c412-40fa-591a-83ab-f271e0c47eaf
> GPU-0b7c7d6b-3501-00c2-a1dd-05bf337e73c0
> GPU-c07c3ce5-0abb-99f5-4ced-0cfa5efbd8b7
> GPU-ad67b2bf-866f-7b25-df05-126a47ba1abb
> GPU-d4745bdd-3f4a-4f48-6db6-8eb784bbf628
> GPU-42827971-4e94-ba68-b9c8-b4499bace407
> GPU-2662c707-22ea-a046-c3f0-5034779e3df7
> GPU-400e5eb2-5410-579c-ad2e-41c933ea91b9
> GPU-cd06f17d-45c5-90fb-13cb-54848b3ed99b
> GPU-97bf3fa8-f80a-fb53-6309-4c345f66ac76
> GPU-c441ddc8-aa02-7c35-a8e7-d56d7a170c10
> GPU-15d2cf9d-8255-27a1-44ac-08305313340f
> GPU-d8420e5b-a5f7-334c-470c-fb524c17a878
> GPU-737f2b11-8cf6-4d49-ce80-183d3b0d6b07
> GPU-3449ec77-8480-57f7-3af6-de9c23ad4966
> GPU-90f225de-6513-6661-5205-566b16cde85b
> GPU-715fb37b-ed08-b368-91a3-1f3bf19abb1c
> GPU-a9c894ca-cf0a-0b20-9e14-402b7b2449e5
> GPU-e2d23c8f-b1cc-9963-4092-87ca2dbb6f00
> GPU-77b96020-b655-40ee-8597-2e5f30437548
> GPU-d01ced75-3bf3-22d5-8a31-dc9591407ceb
> GPU-4465b211-f19e-8af2-8838-563d0d44aac5
> GPU-de9d48fc-a5ff-4563-9503-f5df43813d5e
> GPU-dedc615c-3ce4-f30a-220a-2b4ca1125649
> GPU-a2915c18-4ca4-edfb-043c-7357c59b3388
> GPU-0f59c74a-fbf0-a5d5-c3be-8a369a899340
> GPU-16a9ed81-e42b-26cb-da69-154982aa4b8a
> GPU-49acd816-2d66-a877-835c-4a577000462b
> GPU-e911a35f-a83a-2899-4894-fbf0e70b8a68
> GPU-3a755856-b6ac-1466-fbbc-8b0c547b7bb8
> GPU-28729e77-dfbe-a1c4-1ed5-50928ba42dd2
> GPU-debc8ee0-77d7-624d-8e9a-fc644d3f85df
> GPU-ef54915a-76da-2c88-c324-235a414eaf6b
> GPU-fb87fd2a-d4e1-4d8f-7c21-c666a11508b2
> GPU-1aaba1b7-e8a6-16a1-e8e2-8e755bf814a8 unbanned
> GPU-041e804f-0a90-c358-08af-438f4ac657bc unbanned
> GPU-2463a92d-6d75-403d-8720-fa6aac348175 unbanned
> GPU-d49a5960-33ba-86d6-90b4-7410942159a2 unbanned
> GPU-3d7df8e1-1112-414b-3f36-3d94abad64fd unbanned
> GPU-6f80e6ee-379f-35df-b2fd-46bb50865a2d unbanned
> GPU-68f34cb2-7438-8523-57af-bef5e4a8f757 unbanned
> GPU-55f4ba58-2dff-4b75-5ed8-e2f87c9063bf unbanned
> GPU-3aa6aee8-8a71-e4c6-81d3-c8294d57ff37 unbanned
> GPU-dcab6394-d6bd-546d-c23c-06afc5cf2cfa unbanned
> GPU-d7b77b1e-93e2-f625-9acf-a7939292293d unbanned 
> GPU-17dacd7b-756c-95fd-17ed-66364f728630 unbanned
> GPU-477aa585-98c7-3bae-7543-00838d3cb436 unbanned
> GPU-96918233-c889-9687-3bb5-aec69b187ed4 unbanned 
> GPU-42901281-9596-d360-b1a6-c976bbe7288a unbanned
> GPU-f1e83180-ff31-48fb-ed29-01d5769dc18e unbanned
> GPU-55286ca4-1d86-3d00-9b2f-3431529a028f unbanned
> GPU-c487e592-b332-b6d1-b2ab-9f84f8123189 unbanned
> GPU-9ad364c0-81f9-c393-4d31-6c7bf3ca3457 unbanned
> GPU-c12757c2-f349-318a-2462-699c5a680d98 unbanned
> GPU-34e41d33-1aee-c61a-ad52-006b60a1e4d9 unbanned
> GPU-f3e7e76d-3722-dda4-0f66-9ce65a7c7f69 unbanned
> GPU-8635ef91-203d-a9a1-0793-fd267145e628 unbanned
> GPU-b92daf7a-4d6f-f068-fd1d-15e44d484e58
> GPU-0698290e-d1e8-c34f-d0d5-a48ddc063240
> GPU-e92db3cc-a10f-73c0-dd83-69333702d1fa
> GPU-1fd837e8-cb9a-4e2f-561c-0f295fc5f34c
> GPU-fd259989-8cc2-6d8b-fb48-57f17642c6ad
> GPU-35312688-d4d1-8d60-17ae-4e9ac58b3ceb
> GPU-6d8bb980-939d-9d93-aefa-bd3157014e5c
> GPU-f4f64750-ece9-e05c-df8b-162c497b23ce
> GPU-569da1a5-5a5e-3c91-7482-1fbf54ca8125
> GPU-23bed51a-62d8-c139-c363-a86d57c81050
> GPU-f7492b41-dc13-06fc-3721-77269116cd87
> GPU-183abad4-0857-74cf-bb99-d772ff327228
> GPU-791d59ee-5e21-06b8-9caf-c860431c473c
> GPU-4e564a7e-3b47-7a57-e091-a7dde6f4136c
> GPU-b7b5c361-5acd-82ad-9105-ff134071298d
> GPU-bdf73977-fc70-7892-bb8e-259fb7cb270e
> GPU-ade63ada-e427-4415-9d94-cfb9c8a72582
> GPU-67be8b22-7d92-48ec-adec-bbdd1e7db670
> GPU-458cc5f5-768b-dfb1-eb12-d4a8f8ac2d37
> GPU-1379be3c-62f9-9813-0c7a-af0a440ac91c
> GPU-4792be81-c719-bcf7-2974-136ca4fe124e
> GPU-a10f2da8-053b-7867-b7fb-f547378ba7b7
> GPU-d4e9a8d0-8cf1-8557-0f7c-f8c53268d04b
> GPU-dc54ab79-956f-6be9-5d81-883cd204cee4
> GPU-3328e759-df3a-cd5d-a9ff-7c23a0a981ec
> GPU-4d1de856-26c3-1707-e4a1-a64dc3e326e1
> GPU-35e71dc9-ed54-2668-5e78-46e44d5f801a
> GPU-300f5dff-4f94-8183-2406-b4eed34a22df
> GPU-a1a4a4b6-f7ed-9973-fc59-8c2b1994bcde
> GPU-056a9981-f981-cf15-0bb3-9798c6e0fae2
> GPU-3f24a77d-1724-17ef-f0e8-7f9b3b7cf459
> GPU-97a77b08-e0bd-e1fa-12a5-edc080d1e818
> GPU-71971c8a-fe83-96d0-32e2-0a5576eb7213
> GPU-442c3174-a718-d7e7-dea5-5b79af953f9f
> GPU-62dc4fcd-0c5d-9692-065c-d529bdb474ff
> GPU-2a397b11-4e4a-5032-3e26-2664fdf3b31b
> GPU-85516a2b-4c76-ce29-8d94-5c7b4b3300d1
> GPU-d48ad9d4-1be3-1955-cafb-2ef27e8d16ff
> GPU-fc48e85c-be7a-1d43-6f1e-31bc6500f0a5
> GPU-c0351aa7-6d76-36a7-60fc-b106fcbdac2c unbanned
> GPU-32c902bf-40b7-10f7-31d9-c2286ae08dc8 unbanned
> GPU-0ed779de-2a8e-5b4d-173d-b849135a0ddf unbanned
> GPU-80ec3f00-1a7c-58e7-b30b-ba54ac17b87e unbanned
> GPU-7e780b6b-ad75-3ed3-d49c-3c22b7b7ce5f unbanned
> GPU-58e2e00f-5477-f848-b6b8-8ac480d3a859 unbanned
> GPU-0ab6973b-c48a-6e84-0a63-308042920ea3 unbanned
> GPU-b97d80e6-c9da-3206-e284-02f9c51ac00b unbanned
> GPU-a8101a2b-6f3b-08bd-2a69-8c921afe8b9c unbanned
> GPU-246e414a-99fd-272f-6d73-03821e6e9790 unbanned
> GPU-e0dd7322-1802-3a22-d059-0b4e8025d52b
> GPU-8cebfe61-8b1d-2426-0402-8a82b47a9607
> GPU-69251897-19a2-6b0e-13af-97b0fd47f8b0
> GPU-83235d4a-7c42-be3e-7de2-298b9d184827
> GPU-f2f797d9-863b-3349-a4ec-9828ffff0066
> GPU-9ead8642-1ec1-21d0-18e2-3218fd3886f4
> GPU-15839fd4-4031-31b7-aed7-5fbfc7e22511
> GPU-c0951a9d-541e-8b6a-8648-5f20b2a75f81
> GPU-448c5c5d-3f61-9c19-4d3b-d49860c6bac1
> GPU-fa4c2d7d-488f-4bb4-dfd5-e7e1dca1309d
> GPU-953ff25c-b452-0001-7b10-26575a0a7f6a
> GPU-dbd80a22-9f16-322d-9a33-38c86c634ca6
> GPU-9ce62cc3-c94a-5e9e-a518-547277520c2a
> GPU-7a36ce70-a667-42bc-56e9-43da80ecab3d
> GPU-b733d11f-656a-fd59-69e6-7ee24d06f3b1
> GPU-b1ef05a5-43d2-87fb-c757-b3e464c45c21
> GPU-1bd6bc81-9000-fb44-9735-fb3ffe98d092
> GPU-d996c3f4-7ad9-3661-0ded-3ff1a533d2bb
> GPU-724920de-2967-54a7-22ee-c5e5884ca2b0
> GPU-551e59ea-18d2-211c-7903-0b14f18314e5
> GPU-0fac6ee7-64fd-3ac4-00b2-0925ea64242a
> GPU-9ce484ef-52ce-3a8e-a718-bb90d805705e
> GPU-4b26a962-fe40-612d-0bfa-47e389560a61
> GPU-8016c45e-399e-7cfd-fd16-ccd175cbfa08
> GPU-3d6ecdff-adac-f0a0-9961-caf5ca3028a3
> GPU-e7f9a8ef-33e6-654b-b4ed-fb1d228c4854
> GPU-b9a7f57f-7be0-dad1-2444-7a341a8f5a04
> GPU-18f26066-05ec-5629-b84d-2aee889d74fc
> GPU-71c7a0a7-f85b-22eb-a85d-42c4f20fa296
> GPU-8a562903-9a6f-d461-66dc-69b0f23c6eb5
> GPU-2691b2be-ba9e-6342-23fe-da4d9619d3b1
> GPU-ccd51b6a-4d17-f471-9d6c-9f82be65f0a7
> GPU-d302774d-ae2b-f717-cd3f-b77fa923e649
> GPU-5213d6b5-47d4-5c16-e9bf-3ea56af0f08e
> GPU-c5f8d656-c929-473d-f92b-de5e3de44bb2
> GPU-c0b10487-c902-37fb-7ffd-3e8e224ba186
> GPU-cc5f01bb-35b1-41e4-36f6-7acc93c5532d
> GPU-d4c9c371-cbe9-7940-cc51-ce6749f3af42
> GPU-05553974-7880-0085-6e8a-a38a3166ab07
> GPU-33ab04bd-2274-6b3c-7799-cbe37ffe4754
> GPU-16fc8673-aae4-99cb-85ad-47fa9e812783
> GPU-54168e15-1a11-c962-b3be-fb1d48c38895
> GPU-fd445f1a-93be-24ed-bbeb-1acae3bb355f
> GPU-c9082381-37bd-8695-91c1-be94e6936ee0
> GPU-67ce2037-a909-4256-76d7-d4e954441eca
> GPU-ba25180a-00c8-ca20-bbea-9755b4a4259d
> GPU-38af1205-bea5-59a4-278f-392d6004c539
> GPU-281feaf8-d865-36ec-5d9b-026de11e1d65
> GPU-4a911e2b-b7ad-6cad-157e-4935b8450d68
> GPU-50ab043d-2111-08eb-fb59-cd77e1ef74aa
> GPU-93ed4acf-74c7-1625-c217-8d69bb825625
> GPU-e97d1f46-a938-ac9d-d249-c9d35e34549e
> GPU-944bf9a7-6dc2-b965-b5ea-7a67e7ae8dd1
> GPU-35cdffcf-a044-6f83-de0a-7fcee2b04e11
> GPU-12867de7-6d3b-8c19-d00b-0fb320595724
> GPU-9096128f-4b73-75b2-0cc0-38871f6b5b02
> GPU-5d14724e-385b-5fc1-40ce-bd9ef5a92ec9
> GPU-272b7999-433a-7903-f156-3aac646cfdc1
> GPU-b98e47e7-9c81-7030-fe7f-2566c1879a7b
> GPU-3698e903-ffca-ed96-6174-0928ff22ed44
> GPU-0a168991-0aec-f3c5-bbfc-c401f4ab720c
> GPU-630477d8-490c-3f1c-d53f-c0d13840a3e5
> GPU-f32ef427-2736-2d0f-13e5-4cba374af017
> GPU-63cf052b-dec3-0a08-34e4-4e93498af173
> GPU-fa53b7fa-b183-b0ca-9f97-76bf49970630
> GPU-8d2ea040-def5-7d9e-67d7-56c3fe2d2dab
> GPU-6203e3b8-3cf0-2214-6963-0fe4d549106d
> GPU-e73c59df-5062-d390-80b2-8a67df07b3bb
> GPU-8bfca0fe-80d7-5c34-cb4b-2690e770a2b5
> GPU-892f07e1-bad2-e30b-3802-54782c3d715e
> GPU-09c9280a-17f5-1ae6-1a26-ec021024fca4
> GPU-73ac1332-0aec-e9bd-8f46-53952b98fb60
> GPU-6331a78f-1ff3-ff9c-3c9d-c5ddfcdac1ac
> GPU-a6589932-3222-82a2-7505-9a1d9e3cf7c6
> GPU-ee5d85f3-91b1-4fca-0a63-012fdfc2faa5
> GPU-02c1b93f-8a84-ec9b-1d90-8ad5446f97eb
> GPU-03a2fd83-0ee1-eaa8-2f3d-52d89c173195
> GPU-3e9c8cb8-84bc-8344-5a08-af29b9c4d9b2
> GPU-80fc2d08-f0db-a39f-4cf8-14c2173ed1ee
> GPU-373e4272-b84d-5a71-146a-2ebd95a877cf
> GPU-34972143-afe5-5128-85f8-e5eabaafc519
> GPU-6572ebd5-c8bb-dfde-dc55-cd3ca3ae0867
> GPU-bbbfb298-c20b-416b-a188-f7a28ea033cc
> GPU-fde09741-67a3-489d-bb82-799b03b43abc
> GPU-c95eb546-bd4d-48be-afa4-034572690aa1
> GPU-fefc5d39-f591-4852-8c53-1ffcb40d8369
> GPU-d65c5845-98d8-4213-b216-5d7eb52ae6fe
> GPU-3e5c63fd-07aa-4165-bdc3-99a45c8ae9dc
> GPU-8c4c874c-add5-4431-b90c-9b9633947087
> GPU-8b9f5e8d-8e1f-4d5e-9a6b-4b2e5e8f4b63
> GPU-e7b2a3d1-1c9b-4f8f-9c6b-2f7a9f9b4c1a
> GPU-2b7e8c3c-5e5f-4f5e-bb1b-6b8a5b4b0d2e
> GPU-df8f0e1c-3d3b-4f8c-9e2d-1b8a4c9e3f7d
> GPU-5b9a1e2f-7c3c-4d5a-8f5e-9c8b7a1e4f2b
> GPU-b9da3209-0166-47c5-a0fe-c6b000f8ac42
> GPU-4ac9006c-2ff8-4ecd-90e4-990a6376d63c
> GPU-78d9f801-64db-47c9-9b4b-fd9818c82e55
> GPU-79c72fe5-4aa5-4c35-8afe-6a001b977892
> GPU-07e900f2-96ab-4040-9ee2-3343e2e9dedb
> GPU-94667130-ab9e-45e0-ac1a-f80d79930558
> GPU-494da28c-16af-4960-b07a-bd2788797651
> GPU-592727f8-b3c9-4e8c-95b1-74eb42be78ce
> GPU-26dd369c-3e8f-46cd-bf11-fd271730820f
> GPU-03a31755-ea00-4f8b-b281-12f042439db6
> GPU-5afcec8e-05ec-4291-b2cf-b6560a3994eb
> GPU-2a0e3fd3-24af-4005-a031-ed1eef29d99f
> GPU-ab86ac29-02ea-4a9f-8f6a-ae571dd23e80
> GPU-a685138e-e2d3-4542-838d-35f47135f5d7
> GPU-95ffdc7e-ebb7-423f-ac47-2fa2f8f211cd
> GPU-d2648e02-ea98-4f99-b5b7-bcee3484d47c
> GPU-981fce94-a2a8-4c8d-9dda-cf4b00dd7fcc
> GPU-93b089a1-deff-41c0-88db-b35e8172eab3
> GPU-3fcc749b-20f8-471c-9b01-713f30b68d2a
> GPU-cb604b1b-e826-49c5-8c6a-baa8093c7935
> GPU-c71b713c-a038-4033-9300-8d5d508894fd
> GPU-7c680f26-d3db-4076-8dfd-f94a70dfa4b3
> GPU-9de8763b-394d-45d5-9a8f-53d65ff8dec1
> GPU-248a28ef-4cee-4966-88d4-796bb8f46c8f
> GPU-f949028c-5fb4-41c9-883d-94f0ce735286
> GPU-ecf5577d-52f5-49de-b83b-22fdaa574a45
> GPU-b3fabd06-639b-410d-8a5f-b24f1744a305
> GPU-3ead0308-4b27-4139-9ff2-6dab84865388
> GPU-db2c069e-70e4-4914-82d3-7be7064e647a
> GPU-d8b1e59d-c650-4241-928b-72ec203ea847
> GPU-566b09df-4ce1-4f55-9e91-95fcbe625b0d
> GPU-b74a98c0-66cc-497f-b7d9-ff072b7f9586
> GPU-83500079-c80b-4035-a636-9e3eec4c954b
> GPU-93575677-5d79-4311-ad26-b76ddbddc906
> GPU-c4f2fc6f-c429-40f5-92ba-830f274f054a
> GPU-4b0c3aef-5c1e-4b08-bc5f-0c0f3a4e5c1a
> GPU-9f1a3e2b-7c4f-485f-b0e5-3d1f4c7e8b6b
> GPU-50941b82-29dc-47bf-a24c-a226cc7ea987
> GPU-12a4494a-480c-4663-adc5-028a3264f9cb
> GPU-60ae4ac7-8b9c-4f21-80c2-3600c2232517
> GPU-e375b039-78a3-466c-9adf-27f6513095b2
> GPU-a55a7be4-d0ef-4f33-bf1c-f0b654403e95
> GPU-c5a0f719-967a-46ac-a73f-3903a755bfdb
> GPU-3f9dcf9c-d7e6-4b21-803e-2b7cfd5227f1
> GPU-919dc56f-37ed-4d77-bf52-870f7a5a2c83
> GPU-b1f4c7a0-c26e-4cc5-8982-a7792ceae276
> GPU-60bebe62-86dd-4dc9-9643-1931cc2c4857
> GPU-2041f5d4-ab3e-4670-a170-8489e362fc4d
> GPU-5eff3ed7-ebd4-4703-ba02-2ce632ecbb1d
> GPU-d44cc645-88f5-4611-8d25-c9c2276d604c
> GPU-7d67a83d-e3cc-4843-97ce-a2ce86d88f67
> GPU-a1294240-7d73-498d-a8eb-a797d9516c1b
> GPU-8bdf4b76-9767-403e-b03f-64c9754afae3
> GPU-73ba7d5f-770a-4443-99f6-448bb265a9ed
> GPU-b2dcc1f4-6a22-4034-9efe-89b0b600f094
> GPU-cc4937ff-3661-457f-bb44-0843e0536915
> GPU-dd281e8e-05da-4651-a043-21e858c5baec
> GPU-d3cde7d2-38f0-4bb1-a43e-4b17dcb4ac51
> GPU-e3042b17-8361-4ea9-9a38-494569cce307
> GPU-1c2416fb-8145-4910-b473-df4163aeea09
> GPU-3fd8e822-0ce8-4ad5-be04-7b520ed88536
> GPU-0c884c25-75ea-4e69-8c34-9d4768c52d9d
> GPU-1ce4cc33-567d-4b4f-97f7-830184aeab6f
> GPU-0fdc359b-5e6f-42ef-832e-e253a70380cd
> GPU-0ae40dc6-7d79-48a7-8b67-10e76f9f723b
> GPU-dd05e091-697c-48f0-b526-fdacec53b728
> GPU-b6761af5-45d2-42f7-8f20-aa61e50b9996
> GPU-4003f699-46df-43cc-a968-9f0a3f089824
> GPU-0337b777-b9ea-4dda-bdd7-6f376fe1f436
> GPU-d990103b-58c1-47bf-8258-1339099fa9ed
> GPU-a6d107fb-0063-431e-bb80-57720c7ec688
> GPU-67388370-9285-4164-a909-df7d4161652f
> GPU-87bc88d5-d05a-4d81-8106-7a08e56f307d
> GPU-0f3a5b2c-2c8e-4e5a-8b4b-4f7e8c4a1234
> GPU-7d3b0a59-1f5d-4a4e-9f11-1f2e4c7b5678
> GPU-9b2c8f9a-4f0e-4c3b-8e2b-7a5d9f1a9876
> GPU-5bbb5988-7e65-676a-51e2-5eae0c6e3bf9
> GPU-ed53e036-742d-a73c-d2dc-bfd6ef27632b
> GPU-dd240824-83d4-ba99-bec6-1b7d957e292c
> GPU-ea71b6aa-6fb7-ebd3-1241-621093032187
> GPU-35f3b3b7-fdde-4de2-c70a-ff9f9475bd9b
> GPU-1b8d45fc-1cf4-d2cb-7368-1965477e7b7b
> GPU-41640276-daca-45fa-b64a-74b19fb9c6c3
> GPU-21d986eb-edbd-b8d1-5cdf-9dd747e5b7be
> GPU-4575facd-8049-a8ed-964e-b4e2347a95ed
> GPU-85cc6011-c5e1-dbb4-c212-3c4d85d17398
> GPU-419f1439-f3c1-d45f-c0c8-d91426b017cd
> GPU-185672fc-460e-eccc-3cc8-e3f27674f324
> GPU-1b4be2db-a1fd-1f59-52f2-aba63adfc05c
> GPU-3de7cce8-599b-24ce-7dc9-2c278bfbed0f
> GPU-4701a8f6-4ffb-51d0-100e-01c32f72280d
> GPU-7d504e13-77fb-9810-560d-693d09a43527
> GPU-a8af1643-a5b8-0d02-f636-e90a308889cf
> GPU-c6b61828-e2d4-0249-2782-5731842e6386
> GPU-c3ee6d8a-c05f-220f-4214-68625c37236a
> GPU-b44159e9-03fc-d66b-1746-d5ab8129cad2
> GPU-dc010f19-62cd-0e22-b510-aed3062498a9
> GPU-7bb9c200-18c0-07be-aec2-548f11e29ce8
> GPU-c16a3542-caeb-a63a-1375-0c470ac177fc
> GPU-761584e3-8b16-cc7f-ad98-f840ff590292
> GPU-9bafb349-b71e-f91a-e3da-590fd69ad52d
> GPU-fb647df4-45fd-81ef-cad3-c42defc41e2e
> GPU-592045af-b62a-00a5-a06e-735682273438
> GPU-7fdac777-6989-39b8-9e43-bf6f24ec469d
> GPU-dcb4fcc5-ea2a-d758-f5c1-e595c01aff91
> GPU-2e028194-da3f-0824-6960-4e5b80db6abf
> GPU-c8702c4e-75df-12e2-5a65-23b722a48062
> GPU-786527a2-9e69-95a7-8e41-1a849680394d
> GPU-eea2ca0d-6903-1231-2fa3-6e6e5e79f159
> GPU-2d840d9c-f48e-3ea6-ff19-91526b53a331
> GPU-fc67d3bc-43a8-d2d3-6dad-56cb005f1c6b
> GPU-f3e8d1d0-b147-4f92-a69b-4b3da8d84bf1
> GPU-11451540-4616-b445-1664-7f3ad1971e86
> GPU-96e6094c-e6fc-4c37-88e5-322a4868a5fd
> GPU-b3da4c6c-6477-c424-7f2c-616d029d78bb
> GPU-fd9dd891-6293-8f07-079c-b0b50cd8dd53
> GPU-9851d52d-3b63-7be4-1f15-d34e879f4d51
> GPU-711f6603-346c-fddd-a914-ec8dc078d4fc
> GPU-a800ded6-4c67-721d-18ce-bedcc79c5fcb
> GPU-fb074718-ee8f-2b26-1342-664df10ae7ff
> GPU-47caf6c9-79d3-98bc-c8e5-baef38e188b9
> GPU-9ad4b9c5-6fcd-d456-6455-91b28714cacb
> GPU-cf6f91f7-febc-74e9-bbb0-7a98a5788f87
> GPU-8fbcb79c-76d0-2357-337b-e4914ce12692
> GPU-b49d5f5a-e4ca-f367-e1c0-cbb16c09ba1c
> GPU-12e74bb0-a0b9-f29b-9407-d6f87caaf756
> GPU-5f3495e8-4883-7561-a7a3-606a0a5753df
> GPU-015eef9e-a0bd-aa44-040d-aad8cd5629eb
> GPU-b90382dc-f274-27f7-2df4-7bf07f9c9e5c
> GPU-0a6b1b04-6b88-ada9-fefd-977f5b44c3ab
> GPU-153273c8-9418-ea9b-3d0e-5738a29f51c2
> GPU-b7c24b2f-2824-164b-4688-769c3533be51
> GPU-e09d1221-a46f-7f77-d999-bd32125447fd
> GPU-b98ca4bc-656e-5f20-4da8-68ef0f621336
> GPU-edb6b69f-4fc7-4b62-8c47-79bf26976c4b
> GPU-db25fbf5-1d2e-ad9c-3de6-e357348fcb2c
> GPU-89895123-2166-32d7-261b-9364cd7d20b0
> GPU-5fc426cf-fffe-c3bc-c2fa-ac5be9b1588a
> GPU-29ca18fc-6fa7-f68e-704b-864333cd07fd
> GPU-d7f39e65-e3b3-7233-1f3f-019d38a7a945
> GPU-7e2444ec-da46-8272-fc0e-b56aafa0d890
> GPU-eba599db-aef5-67a0-f53c-ec90932f8e95
> GPU-427a8949-bd5f-1204-1775-ce9e0aa924cd
> GPU-99ffb474-3214-ff44-1d1e-b96cdde979af
> GPU-9fd05f14-e1af-f6ab-da1f-f45f65739319
> GPU-8fdf7d4c-7714-f431-d93b-d181f6e44840
> GPU-fec677c0-c341-36e5-7bc2-98928f78546c
> GPU-7427f52f-0e28-aaf3-f0c1-0c122ebd43f5
> GPU-c0a20aee-05d8-09ec-c28e-1fa64bf987ed
> GPU-3939264b-9dfa-3ee4-03de-0de900833852
> GPU-a5d7ba5b-370f-3cd2-62c5-9dea894cba12
> GPU-2ac8e81b-d2f7-c661-41ed-e7be767a2204
> GPU-8ca406f7-ea40-3ca4-8235-4bdd09075d40
> GPU-93d6ab90-3080-a12b-dcc6-e0f8b7e33e65
> GPU-f26b9602-f431-c86d-986c-a061ac5bdcf1
> GPU-44b194b1-74e0-060c-1761-fbba1670cc80
> GPU-696639ae-7174-aa15-bedc-7d63c2aaabea
> GPU-f4b5c2e8-0148-2ced-fdf0-da103a93fa2b
> GPU-e019cd6d-8817-c8c5-bcbe-95fe8a726877
> GPU-2c49c92d-13f1-eb55-2465-860142c3f22c
> GPU-4156c23d-c424-1c9a-2fea-bfe19ea34c37
> GPU-64246f41-0d4a-1c81-7153-a8877c9430b1
> GPU-9afaef7c-4f52-272d-a747-a5e67bc57f04
> GPU-56f0b96b-8f85-65fc-2db7-75b973d1f708
> GPU-9679a09c-22d3-ea2b-c7c9-ad8409933842
> GPU-fd9c1c4a-0647-bc45-b910-1be752286faf
> GPU-1fc76a36-7d0c-076a-3332-722be8d13ab4
> GPU-e5992cc6-4633-a9df-ba8e-b9903c3d02ad
> GPU-cd68e97f-4e80-d2d4-5b0b-2f72fa099019
> GPU-6fd31c30-2257-21d7-2a01-170d3b86c1d0
> GPU-e1a588c5-b6b1-d909-a81f-3781f7b4e4ed
> GPU-6f3eab84-1f22-7925-f747-ce1d3ed6236a
> GPU-e0aa7050-95fd-f368-6a2a-f828ca728be1
> GPU-71968c8a-9c56-2cf7-af62-00e91b3cdab2
> GPU-01b097c5-913e-3036-4038-1cd3b252bfe2
> GPU-c9f0ff98-3041-73f4-f451-f023f7e95ee8
> GPU-6405cdfb-b53d-06f2-d81c-352eab05f346
> GPU-fb8a549b-a550-3420-cbdf-c00a2896b64b
> GPU-826ecb18-092f-c331-df34-8222622f4f2a
> GPU-10a3ba8c-55f3-a8e4-10ea-9cb019d7e17f
> GPU-7a629807-4454-c264-80b8-0c711da1e4ba
> GPU-ee9aff4e-284f-030f-4eb9-0b38dad3badd
> GPU-8f10650c-2493-7805-fe8a-6340b81bd834
> GPU-f227e05f-d1de-bd91-109c-76140f4cb37b
> GPU-3261621b-c5f0-2c23-1c34-6104c63b2c44
> GPU-224639b2-c8cd-6c54-1b6a-cec74a3924d4 unbanned
> GPU-6c858ef2-a53e-e9f8-17af-45dc5844ea20
> GPU-0f250aac-d694-732f-c0ba-9a21d7e2c29b
> GPU-2f193d53-329a-7fad-d673-43093ad98321
> GPU-8e65927e-a227-0170-351d-0ffc2fd93938
> GPU-21d54733-5ecd-8c70-b619-dee5295fccc1
> GPU-7aaf3bf7-d09d-6113-6a6c-4729586b726a
> GPU-104e6ef6-0b82-a04d-4f1a-df5027177ae1
> GPU-9929b97c-47bf-8fac-97e3-bf91b53ffdd5
> GPU-d6707021-f5d9-c8c2-8de8-53ca4a317744
> GPU-17cf685f-b305-a41f-ef27-aa6631cabbd9
> GPU-1b766105-b94e-3354-51fc-3141edae7411
> GPU-de02ddb5-03e0-c85c-567b-7a038f6a7853
> GPU-13bca69f-b1e4-eaed-24bb-c3a849d0c4b9
> GPU-a6fee343-7842-b231-f4b5-82f6c33332db
> GPU-ba23b60a-d746-f11c-693a-2e79529bc76d
> GPU-bdf4af3b-1e7e-26d2-9577-638b9446c0b7
> GPU-1dbd4185-2c15-4e07-353f-7e583914fa4f
> GPU-6419672b-7879-bbf2-09a4-69249777ad2b
> GPU-0715f921-7249-b4c9-cd7a-76d2a0baa2d5
> GPU-b206b34b-a9ee-3600-a238-628cfd27d528 
> GPU-30af9830-d5d7-e881-f8de-8f75761be018
> GPU-67c8844f-8c39-ec98-a0d3-e139a61c5b3e
> GPU-7766a7d1-22b3-6450-b942-394698ddfcf9
> GPU-d819aa83-0258-a0e6-951e-a9b8a7b66283
> GPU-7be026ca-7dd6-ec2f-30a5-6e54304dfa83
> GPU-2be7a58e-426e-66d6-368b-b4c2b794a03c
> GPU-51cad0aa-8984-1c81-ac14-3384d1507d4a
> GPU-d7810df0-5bc2-7fbe-e3a6-4ef5c933c802
> GPU-5518978f-a831-92b0-3065-53a19942b069
> GPU-fdb5e235-f638-d7ac-6b1f-0faeb971e9ce
> GPU-f232c074-2b6e-adf1-0dae-032af86ff46c
> GPU-91c2d454-869c-a9a9-4309-c5bbb98ff126
> GPU-a36ff0ee-c605-fdd9-e740-a8fb048407f1
> GPU-97e12548-9dfa-94b1-1672-6b2a03ee2d32
> GPU-2408eace-7680-4d65-e485-1e7ae027f442
> GPU-d8d8e6f8-e834-75d9-70a5-f7b28d10c7e1
> GPU-bcfddebb-2f98-e3bd-a013-999cd91a8264
> GPU-aa98766b-2b59-5237-b9cf-112a2028b725
> GPU-c442c856-066d-b995-a9bf-3d7e841ed474
> GPU-55d7d4a6-713c-2821-75d8-3f4b2c7fa902
> GPU-77db2ee2-5d33-f2b1-3e52-d0d8141e2277
> GPU-f36c6ae6-ed4c-9f2d-acb4-cd5d7453a1ee
> GPU-83001656-b696-9874-11d9-df09e714f082
> GPU-1342b9bd-774e-7e5c-ad12-fc8f149fd572
> GPU-2e4ee815-848c-70ee-e4f9-2c3f0162ebe7
> GPU-3ed4bb6f-b79b-22e3-c358-3bcf6f9c72df
> GPU-d4437df5-08ce-e18c-db82-1209035b68cc
> GPU-47b6f4e7-0c26-c632-6ad1-c7472291b376
> GPU-1869ea32-d437-39c3-8731-69bb62d390fd
> GPU-41296b2f-7b31-3f3b-d75e-6d7c3a6ed63e
> GPU-af5a8227-ed6c-83c6-e9b2-0ffec222b1e6
> GPU-af2beaca-0e6c-1451-21e0-4dc3ceb93fbf
> GPU-8c512dc6-4c9f-5818-a715-69bb433a4735
> GPU-14ce04d2-ff87-de3c-3b20-992660faa369
> GPU-1a3aa5be-cb8c-e0bf-daa7-e2e03bce7f67
> GPU-73317204-bd25-9abb-950e-0438ec665195
> GPU-bc1df9c9-9af5-791a-37af-16a72b7f5f3c
> GPU-1c8ac443-23c5-7ddb-692b-8dc4a5fa92c0
> GPU-55d64197-7c8e-a31e-de8c-5fd3d06fb037
> GPU-c3865dda-889e-9527-b615-7ef63e32f478
> GPU-030f14a8-2538-488d-2046-63e87e59a91e
> GPU-658a913f-386d-5837-ceb0-a9062433b10b
> GPU-07b86aa2-4933-9214-e4a4-fc747753ff1f
> GPU-c97cd48c-bacb-016b-2584-c0df2bff3533
> GPU-97d6ed26-6fd7-c81c-7433-f5882dbe076b
> GPU-80fd2832-2eb8-e9d1-da6e-81344ad22218
> GPU-0d2b393b-cba4-bbbf-5e42-8b28ae3e0ec1
> GPU-959f39ba-a441-0210-deae-3199a509155b
> GPU-4837fb1c-1c48-8614-54bf-c7ea4860501d
> GPU-452303e1-8926-4ce9-3c45-8ca91ae77f58
> GPU-54bf6d9c-978d-7f2f-f5ee-e13f286fdc32
> GPU-74f0d7cf-d753-4bae-aefb-4a452e9e2bae
> GPU-38b84263-d1f3-ae46-d9bc-8514632ded7a
> GPU-747f707c-3b9a-d9ab-817a-591395e800f6
> GPU-57fbd020-4644-95b3-614d-fc565f9f31d2
> GPU-627b2f6d-d67d-6006-cbf4-0ed2f71544cf
> GPU-14ada791-a8ce-209c-31a2-596cb8689bab
> GPU-d89b5514-763d-4612-b337-0f714b296a29
> GPU-23f678ad-b071-e63f-b132-5adf5834915c
> GPU-e4cd1aa2-d401-6cc2-73ea-145c095753d5
> GPU-3885658b-85dd-c342-057c-c3730fe2aee7
> GPU-6fd67a0d-a31b-eb6b-abe5-6e4ad960e6e6
> GPU-31a5816e-540b-e34a-f440-26323ae13bec
> GPU-4f0759a8-7c6a-2353-7de4-668a9cc55870
> GPU-d0a2d0d4-5460-c5c2-aaf3-3dbb2328b39d
> GPU-33354f5f-3489-2d3e-37d9-6e041aa73a90
> GPU-470cc5a7-9b83-878c-a548-77d6229c636a
> GPU-e144e66e-4872-17b3-8636-d1ef060f50a1
> GPU-1a1c231a-d772-be23-9226-dedbc0cd4a01
> GPU-b102a56a-4554-d4a1-0d94-b8498e564301
> GPU-7f516a28-7080-b98c-c19c-237eb69cc671
> GPU-78289e5c-f4d6-60e1-7e52-bc23c2484142
> GPU-663ce6eb-812b-cca8-365d-a81a2aefe448
> GPU-7a35343d-941d-af77-41a4-db515a07aa96
> GPU-461a6b58-728e-20de-08cf-23582bc24fce
> GPU-f425d0ad-1389-9034-8fd0-7aadb3f384e9
> GPU-923bfe7d-fedc-b7f1-8dca-f8c337e0023d
> GPU-4424f43c-6b51-bf27-e358-1f6128e8ec04
> GPU-0a17cf1c-9059-1f00-0c91-739cc03ae531
> GPU-f98649ef-f487-44b4-7c7e-d5de7c71b9f9
> GPU-7fd35aaa-a588-cf19-1999-2bc107720bc1
> GPU-2cf6db25-1e8f-5a62-ef7e-d9e7792f2bcd
> GPU-d962705f-f497-8f5b-d268-e34cadb2c915
> GPU-bd9b8d01-44a4-f2f1-e3bd-7654069b7910
> GPU-72902da5-426f-b715-19f5-bcf3919f3163 unbanned
> GPU-ebdeb319-adf3-e628-2b94-e6b6dd4d5ede unbanned
> GPU-733157cc-f367-80ad-f5f8-fdf7d14e4bf3 unbanned 
> GPU-90a2a57c-84f1-1707-c327-aebd6f87c203 unbanned
> GPU-2ea5d1c8-9433-61ea-5060-e95772fbda2f unbanned
> GPU-f3f23490-ecbe-7601-05de-91d1eef33adc unbanned
> GPU-a74092ef-efcf-2d45-4ef4-a0283e2758ce unbanned
> GPU-509b6c7e-7d38-522e-004e-17154f6c866e unbanned
> GPU-22a93fb4-22a6-66e8-200d-8438ba3903df unbanned
> GPU-1a7a797d-8c4c-524d-92ff-40ea478d8c81 unbanned
> GPU-f92ea265-23ac-0276-5c5d-d3cae2e0ec99 unbanned
> GPU-e1e6246e-a645-6ee9-d55a-d57398c93d26 unbanned
> GPU-addf1b74-e4d6-5923-80a4-afe68787f8ff
> GPU-a72b9a6f-dc07-338b-9199-790403e46c08
> GPU-2222530e-607c-e56c-c3f2-9fe1e382ab0d
> GPU-18e63610-6811-5963-80ca-9c830080b170
> GPU-f03ba71e-96b5-f0d5-d48a-e94782b516c5
> GPU-d6fa4811-e1b9-880a-c58d-e3a04919fc50
> GPU-520756e2-2dbe-1e0e-3e4e-65be93cc6959
> GPU-793a8339-2b35-d749-4833-ab9aa2174d11
> GPU-2d74c45c-56dd-12f1-ff8e-d2a63f018eaf
> GPU-c2f8f2ef-691d-7275-4912-7c5af4e0c188
> GPU-a17ec95f-e3c5-949c-b7ed-1a728a40b856
> GPU-103c43b5-511d-52bf-0a44-db81ad54acd5
> GPU-bc7ba73d-5385-51cd-8dfc-7b79aed7801b
> GPU-15541bce-8e55-d8d0-2bf7-cff655f7ff34
> GPU-821b9c91-1077-1ae9-dba4-05c161bdc0f0
> GPU-caf4bd5b-db40-c7ff-bb17-d1519e121e65
> GPU-c18b2794-6db8-c2a4-1562-0f67538b6fb7
> GPU-f3def060-4a42-4efe-1dea-5661f5385286
> GPU-cb384309-8487-fe11-ae00-fd64332b3146
> GPU-5d96561f-47ec-05f1-d6b8-7d13f071d57c
> GPU-cbf4942b-5316-b50a-7c24-4d3c5c3f803d
> GPU-e25a0050-66f8-f104-7b48-eb172bad2ecf
> GPU-285472c0-0ab7-1c4a-de5e-2ec1e4b98bcf
> GPU-af044b44-dabb-0874-b574-c4ba8fe96d8c
> GPU-bd7164ad-da95-b22c-0bf8-47de012b085d
> GPU-1920753d-6597-5002-f915-27df1828d67e
> GPU-99afa20a-9d18-6a55-d663-44c7d346a5f5
> GPU-39d8122d-da32-3822-a0ff-deef739e291a
> GPU-334920ba-259d-5262-f311-e7984ffe4358
> GPU-84148b9e-6448-d7e4-b0d3-c50fe397c3e2
> GPU-31538ada-4b59-ae5f-8493-96a4b3559613
> GPU-12addaa7-7c93-76e7-dc5a-8646edd9792c
> GPU-7bacbac4-94f7-9f2f-04f2-bac941ccb56c
> GPU-37775393-3867-2e12-78c8-33d7490c131c
> GPU-c53baf8a-e9d0-7a46-6df4-9102a852c8aa
> GPU-8277843e-6eb5-04c7-e73b-c1ab8f18a47d
> GPU-3c666ff3-e48d-d431-4906-2567d2a36fe4
> GPU-45227c7a-bacc-374f-5adc-c29074866b45
> GPU-eb759823-9d07-0270-3336-f82875965f71
> GPU-bc8bdffa-e297-527f-7efa-36b23117e8d0
> GPU-9acd23fb-4523-b1e8-f633-b9ce6d2c11ae
> GPU-8ea4a09f-3800-2607-83e8-307a0125a83f
> GPU-889ac7eb-bb86-b2c7-3c4f-37001c40fecd
> GPU-03fa551e-3650-d9aa-ac44-a382f0f28755
> GPU-ef76ec9d-d383-0f5d-fbb7-4fa10f30c563
> GPU-6b7544b7-86bc-0621-b61f-7c7abcb6e23d
> GPU-859b0b5e-3817-09af-c7ae-a3fefe4eabdc
> GPU-88d15866-f649-fa48-2aad-e65a084d22f2
> GPU-19c4ff7a-98c5-bb48-4493-6e8606d24aeb
> GPU-caa8e94d-5b86-8ffa-0b84-37e8f5274589
> GPU-50520a87-3b43-1728-275b-2cd5d45eaabe
> GPU-8483095e-0d41-39e8-058a-d12ddf95345f
> GPU-a162e603-6a25-1d44-3c25-685adfc7bab2
> GPU-58220d8d-3584-6433-28da-1d4338272ed8
> GPU-ef25309b-191f-2e74-b2f9-82726d48cde8
> GPU-433a3acb-7de9-0a5b-e7d0-efb883b81b6e
> GPU-39376b6f-e6ed-5085-de0c-c3ed4ed0bb87
> GPU-6f8b2db4-a628-0870-21ef-18ab0c2bd6a7
> GPU-c37b7587-c59d-20bb-2813-d2e6cf682ac2
> GPU-05ac3851-9f00-191b-c11b-9cbe1852ea01
> GPU-5c2680fc-e752-86e7-cbae-3bd0272f57ef
> GPU-6af39274-7193-5703-5ba3-b74b9dcbbe3f
> GPU-b07c8a0c-c7c6-5dda-7378-e19b330000f8
> GPU-9d3ca437-adc7-11bb-de1d-d12187dc8d16
> GPU-afa53902-9d59-bcac-777b-91b742b4be96
> GPU-95a6bc48-a872-b603-0eae-b1b4eeb4caf8
> GPU-5e82be97-465a-4942-ec6e-549d66cdf821
> GPU-9d6f1a02-3045-1385-7d97-4f22cdc91202
> GPU-b119b504-7456-4a9a-aeae-3914cda85a00
> GPU-74fb234b-80d4-736c-fd5b-0018dd013cf8
> GPU-e6415350-3d5c-a33f-b0c8-c2365e24562e
> GPU-1572c765-85f5-8b0c-f264-cc06dfad2bf9
> GPU-c4c35be4-cd9c-e06c-7fd7-74d3ac770fa5
> GPU-e12af7f5-8c8f-092d-c7c3-7319632b5c1f
> GPU-9cd15e26-fe3d-e605-0fbd-6e1e5b5607c8
> GPU-9325f3d5-e436-a94b-467d-9d2306b202b5
> GPU-ce9126cd-ca2f-806e-a801-dbd49c7b83d1
> GPU-4bdb7342-304e-d8dc-ea11-396e12070e1c
> GPU-9a59aeee-62d9-2fbb-49ab-f4c123b3ab4d
> GPU-21256502-b6fb-9878-5502-2fc096bac6cf
> GPU-e8b91161-4444-f87c-9ac6-9a5b24ae6cb6
> GPU-b48e2ef7-8481-b646-c099-c5ce6c87ce4e
> GPU-96ed2016-304e-018d-9e5e-97a1c254e5fd
> GPU-bb4ddfe6-456e-4ced-dd79-2dedd95e6964
> GPU-f39bd8bf-1ccb-1b66-b298-b51a8966e11c
> GPU-657e5856-7c71-3c82-afa9-24d5bb28bc8a
> GPU-c62b6bc0-c55d-0b56-fff9-f7811ab5d36b
> GPU-e3a5efcc-27e6-0d1e-6e3c-45ae938747d6
> GPU-a041e7e5-1335-6027-251a-992d038f28ba
> GPU-6e604e89-2363-ac9a-ea86-e546fbf1de4d
> GPU-8bde9b68-1a94-692a-5e17-41a9b98fbf37
> GPU-5bd66683-ffb5-ac26-df7e-ee096628b139
> GPU-652bff72-61e9-0123-0a85-bcd3102b660e
> GPU-f8201177-6a8c-c2c9-b3cd-1c7f3304715a
> GPU-97e69965-050f-9c2f-5956-421d60430e2a
> GPU-3547c442-dbb1-c59c-3ecc-7f0790787ccd
> GPU-9fb61f4d-b4fd-4037-ae04-2af32effa527
> GPU-82602eb8-38b8-cf00-66cf-62ac8c2bb8a5
> GPU-a83f67d2-cae2-c20e-56dd-f9e8eda7fd88
> GPU-9285c38c-fc1b-3b7b-be65-ad5e3a3fdeac
> GPU-172ef0f4-b9f9-6d87-f213-a73a0e695090
> GPU-42e92f47-bb15-20dc-aa54-37f0a4552497
> ```

***

### 2. Why a GPU can be banned

* **Self‑renting** (renting your own hardware to yourself to farm rewards).
* **Any tampering with the Clore Docker container** or the software environment provided to renters.
* **Hardware substitution** — e.g. listing a farm with *5 × RTX 5090* but swapping them for *1 × GTX 1060* during an active rental.
* **Artificially inflating reported core count, VRAM, or similar specs.**
* **Exploiting platform vulnerabilities** to boost metrics such as MFP or otherwise gain unfair advantage.
* Bypassing service fees, abusing referral or reward programmes, or any other fraud.

***


# API

> 💡 **Python Developers:** Use the official [clore-ai Python SDK](/developers/python-sdk) instead of raw API calls. It handles authentication, rate limiting, retries, and error handling automatically.
>
> **Quick start:** `pip install clore-ai` — See the [Python SDK docs](/developers/python-sdk) and [CLI Reference](/developers/cli-guide).

> 🧭 Prefer a browsable reference? Open the interactive API docs at [clore.ai/api-docs](https://clore.ai/api-docs). Base URL for all endpoints on this page: `https://api.clore.ai/v1/`.

<div data-full-width="true"><figure><img src="/files/cG3egNCJEuF7U7op74Ub" alt=""><figcaption></figcaption></figure></div>

### Introduction <a href="#introduction" id="introduction"></a>

[CLORE.AI](http://clore.ai/) api can be used to automate deployments of your workloads onto [CLORE.AI](http://clore.ai/)

Firstly you need to get an API key\
![alt text](https://clore.ai/assets/html/api-export1.png)\
![alt text](https://clore.ai/assets/html/api-export2.png)

***

### API responses <a href="#api-responses" id="api-responses"></a>

Responses are returned in JSON format, may have different fields

Always returned field is code, indicating status

**code field**

| code | Description                      |
| ---- | -------------------------------- |
| `0`  | NORMAL                           |
| `1`  | DATABASE ERROR                   |
| `2`  | INVALID INPUT DATA               |
| `3`  | INVALID API TOKEN                |
| `4`  | INVALID ENDPOINT                 |
| `5`  | EXCEEDED 1 request/second limit  |
| `6`  | Error specified in `error` field |

***

### Endpoints <a href="#endpoints" id="endpoints"></a>

#### 1. `wallets` <a href="#id-1-wallets" id="id-1-wallets"></a>

**About**

Return wallets and balances

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Output**

| Field     | Type       | Description      |
| --------- | ---------- | ---------------- |
| `code`    | `int`      | Status code      |
| `wallets` | `[]string` | Array of wallets |

**Example**

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/wallets'
```

**Output:**

```
{
  "wallets": [
    {
      "name": "bitcoin",
      "deposit": "tb1q6erw7v02t7hakgmlcl4wfnlykzqj05alndruwr",
      "balance": 0.00153176,
      "withdrawal_fee": 0.0001
    }
  ],
  "code": 0
}
```

#### 2. `my_servers` <a href="#id-2-my_servers" id="id-2-my_servers"></a>

**About**

Returns your servers that you are providing to [clore.ai](http://clore.ai/) marketplace

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Output**

| Field                                 | Type       | Description                                                              |
| ------------------------------------- | ---------- | ------------------------------------------------------------------------ |
| `code`                                | `int`      | Status code                                                              |
| `limit`                               | `int`      | Maximum number of servers you can own                                    |
| `servers`                             | `[]string` | Array of servers                                                         |
| `servers[x].name`                     | `string`   | User selected server name                                                |
| `servers[x].connected`                | `string`   | Was the server ever connected to [clore.ai](http://clore.ai/)            |
| `servers[x].visibility`               | `string`   | Visibility on marketplace                                                |
| `servers[x].pricing`                  | `[]string` | Price/day on-demand                                                      |
| `servers[x].online`                   | `bool`     | Is server online                                                         |
| `servers[x].min_spot_pricing`         | `[]string` | Minimal price/day to rent for on spot market                             |
| `servers[x].init_token`               | `string`   | Initialization token                                                     |
| `servers[x].specs`                    | `[]string` | Server specifications                                                    |
| `servers[x].remaining_time`           | `int`      | Rental end time of the latest active order (ms epoch, 0 when not rented) |
| `servers[x].rental_state`             | `string`   | Aggregate GPU rental state: `none`, `partial`, or `full`                 |
| `servers[x].rented_gpus`              | `int`      | GPUs currently rented across active orders                               |
| `servers[x].total_gpus`               | `int`      | Total GPUs on the server                                                 |
| `servers[x].rented_price_by_currency` | `object`   | Sum of active order prices keyed by currency                             |
| `servers[x].update_channel`           | `string`   | Hosting-agent update channel: `stable` or `beta`                         |
| `servers[x].requested_channel`        | `string`   | Channel switch requested from the web UI (present only when requested)   |
| `servers[x].target_backend_version`   | `int`      | Backend version the current channel targets                              |

**Example**

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/my_servers'
```

**Output:**

```
{
  "servers": [
    {
      "name": "Michael",
      "connected": false,
      "visibility": "hidden",
      "pricing": { "bitcoin": 0, "usd": 0 },
      "online": false,
      "min_spot_pricing": { "bitcoin": 0, "usd": 0 },
      "init_token": "qnwVIMsZPjUWS7jw6gAbTOeoGQNgTH9XVxJaiCEbG0OlmfjF"
    },
    {
      "name": "Jan Vykydal",
      "connected": true,
      "visibility": "public",
      "pricing": { "bitcoin": 0.00010337, "usd": 0 },
      "online": false,
      "min_spot_pricing": { "bitcoin": 0.00005168, "usd": 0 },
      "specs": {
        "mb": "Z590 GAMING X",
        "cpu": "Intel Core i9-11900",
        "cpus": "8/16",
        "ram": 64,
        "disk": "NVMe 512GB",
        "disk_speed": 2000,
        "gpu": "1x GeForce GTX 1080 Ti",
        "gpuram": 11,
        "net": {
          "down":119.61,
          "up":25.24
        }
      }
    }
  ],
  "limit": 16,
  "code": 0
}
```

USD values in `pricing` objects reflect the USD-based pricing mode where the host enabled it.

#### 3. `server_config` <a href="#id-3-server_config" id="id-3-server_config"></a>

**About**

Get configuration of specific server

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field         | Type     | Mandatory | Description |
| ------------- | -------- | --------- | ----------- |
| `server_name` | `string` | Yes       | Server name |

**Output**

| Field                   | Type       | Description                                                                                    |
| ----------------------- | ---------- | ---------------------------------------------------------------------------------------------- |
| `code`                  | `int`      | Status code                                                                                    |
| `creation_completed`    | `bool`     | Is server creation complete                                                                    |
| `config`                | `[]string` | Config of server                                                                               |
| `config.name`           | `string`   | User selected server name                                                                      |
| `config.connected`      | `bool`     | Was the server ever connected to [clore.ai](http://clore.ai/)                                  |
| `config.visibility`     | `string`   | Visibility on marketplace                                                                      |
| `config.pricing`        | `[]string` | Price/day on-demand                                                                            |
| `config.spot_pricing`   | `[]string` | Minimal price/day to rent for on spot market                                                   |
| `config.mrl`            | `int`      | Maximum rental length in hours                                                                 |
| `config.online`         | `bool`     | Is server online                                                                               |
| `config.initialized`    | `bool`     | Was the server ever connected to [clore.ai](http://clore.ai/)                                  |
| `config.id`             | `int`      | Unique server ID                                                                               |
| `config.rental_status`  | `int`      | 0 - not rented \| 1 - Rented on spot market \| 2 - Rented On Demand                            |
| `config.specs`          | `[]string` | Server specifications                                                                          |
| `config.background_job` | `[]string` | Background job when not rented                                                                 |
| `config.tenants`        | `[]object` | Active orders on the server: `rental_type`, `currency`, `price`, `remaining_time`, `gpu_count` |

> `server_config` also returns the same rental-state and update-channel fields as `my_servers`: `remaining_time`, `rental_state`, `rented_gpus`, `total_gpus`, `rented_price_by_currency`, `update_channel`, `requested_channel`, `target_backend_version`.

**Example**

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "server_name":"Jan Vykydal"
}' 'https://api.clore.ai/v1/server_config'
```

**Output:**

```
{
  "config": {
    "name": "Jan Vykydal",
    "connected": true,
    "visibility": "public",
    "pricing": { "bitcoin": 0.00010337, "usd": 0 },
    "spot_pricing": { "bitcoin": 0.00005168, "usd": 0 },
    "mrl": 72,
    "online": false,
    "initialized": true,
    "id": 4,
    "rental_status": 0,
    "specs": {
      "mb": "Z590 GAMING X",
      "cpu": "Intel Core i9-11900",
      "cpus": "8/16",
      "ram": 64,
      "disk": "NVMe 512GB",
      "disk_speed": 2000,
      "gpu": "1x GeForce GTX 1080 Ti",
      "gpuram": 11,
      "net": {
        "down":119.61,
        "up":25.24
      }
    },
    "background_job": {
      "times_updated": 1,
      "image": "cloreai/ubuntu20.04-jupyter",
      "command": "",
      "env": []
    }
  },
  "creation_completed": true,
  "code": 0
}
```

USD values in `pricing` objects reflect the USD-based pricing mode where the host enabled it.

#### 4. `marketplace` <a href="#id-4-marketplace" id="id-4-marketplace"></a>

**About**

Get marketplace

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Output**

| Field                           | Type               | Description                                                                                          |
| ------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------- |
| `code`                          | `int`              | Status code                                                                                          |
| `my_servers`                    | `[]string`         | Array of server ids you are providing to [clore.ai](http://clore.ai/) (can't be rented)              |
| `servers`                       | `[]string`         | Array of public servers on marketplace                                                               |
| `servers[x].id`                 | `int`              | Unique server ID                                                                                     |
| `servers[x].owner`              | `int`              | Unique owner ID                                                                                      |
| `servers[x].mrl`                | `int`              | Maximum rental length in hours                                                                       |
| `servers[x].price.on_demand`    | `[]string`         | On demand price per day                                                                              |
| `servers[x].spot`               | `[]string`         | Minimal spot market price per day                                                                    |
| `servers[x].rented`             | `bool`             | Is server rented on demand                                                                           |
| `servers[x].specs`              | `[]string`         | Server specifications                                                                                |
| `servers[x].partial_gpu_rental` | `object` \| `bool` | `false` when partial rental is unavailable, otherwise `{ total_gpus, available_gpus, free_indices }` |

**Example**

Get marketplace

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/marketplace'
```

**Output:**

```
{
  "servers": [
    {
      "id": 6,
      "owner": 4,
      "mrl": 73,
      "price": { "on_demand": { "bitcoin": 0.00001 },
      "spot": { "bitcoin": 0.000001 }},
      "rented": false,
      "specs": {
        "mb": "Z590 GAMING X",
        "cpu": "11th Gen Intel(R) Core(TM) i9-11900 @ 2.50GHz",
        "cpus": "8/16",
        "ram": 62.67348861694336,
        "disk": "NVMe disk 247.3623046875GB",
        "disk_speed": 0,
        "gpu": "1x NVIDIA GeForce GTX 1080 Ti",
        "gpuram": 11,
        "net": {
          "up": 26.38,
          "down": 118.42,
          "cc": "CZ"
        }
      }
    }
  ],
  "my_servers": [1, 2, 4],
  "code": 0
}
```

#### 5. `my_orders` <a href="#id-5-my_orders" id="id-5-my_orders"></a>

**About**

Get your orders

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Query string**

| Field              | Type   | Mandatory | Description                       |
| ------------------ | ------ | --------- | --------------------------------- |
| `return_completed` | `bool` | No        | Return completed (expired) orders |

**Output**

| Field                    | Type       | Description                                       |
| ------------------------ | ---------- | ------------------------------------------------- |
| `code`                   | `int`      | Status code                                       |
| `limit`                  | `int`      | Maximum count of active orders                    |
| `orders`                 | `[]string` | Array of orders                                   |
| `orders[x].id`           | `int`      | Unique order ID                                   |
| `orders[x].fee`          | `float`    | Fee (%) paid to [clore.ai](http://clore.ai/)      |
| `orders[x].creation_fee` | `float`    | Creation fee paid to [clore.ai](http://clore.ai/) |
| `orders[x].price`        | `float`    | Order price (cost) per day                        |
| `orders[x].mrl`          | `int`      | Maximum order rental length in seconds            |
| `orders[x].image`        | `string`   | Used docker image                                 |
| `orders[x].currency`     | `string`   | Currency used for billing                         |
| `orders[x].spend`        | `float`    | Money spend on the order                          |
| `orders[x].ct`           | `int`      | Creation time (UNIX time)                         |
| `orders[x].p`            | `int`      | Currently used proxy cluster                      |
| `orders[x].specs`        | `[]string` | Server specifications                             |
| `orders[x].si`           | `int`      | Unique server ID                                  |
| `orders[x].pub_cluster`  | `[]string` | Public endpoints with forwarded ports             |
| `orders[x].tcp_ports`    | `[]string` | TCP port forwarding                               |
| `orders[x].http_port`    | `string`   | Container port forwarded through HTTPS proxy      |
| `orders[x].spot`         | `bool`     | Indication that it is spot order                  |
| `orders[x].expired`      | `bool`     | Indication that the order has expired             |

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/my_orders?return_completed=true'
```

**Output:**

```
{
  "orders": [
    {
      "id": 38,
      "fee": 5,
      "creation_fee": 3e-7,
      "price": 0.00001,
      "mrl": 262800,
      "image": "cloreai/ubuntu20.04-jupyter",
      "currency": "bitcoin",
      "spend": 6.944444444444445e-9,
      "ct": 1667401396,
      "p": 1,
      "specs": {
        "mb": "Z590 GAMING X",
        "cpu": "11th Gen Intel(R) Core(TM) i9-11900 @ 2.50GHz",
        "cpus": "8/16",
        "ram": 62.67348861694336,
        "disk": "NVMe disk 247.3623046875GB",
        "disk_speed": 2000,
        "gpu": "1x NVIDIA GeForce GTX 1080 Ti",
        "gpuram": 11,
        "net": {
          "up": 26.38,
          "down": 118.42,
        }
      },
      "si": 6,
      "pub_cluster": [ "n1.c1.clorecloud.net", "n2.c1.clorecloud.net" ],
      "tcp_ports": [ "22:10000" ],
      "http_port": "8888"
    },{
      "id": 36,
      "fee": 2.5,
      "creation_fee": 1e-7,
      "price": 0.00001,
      "mrl": 262800,
      "image": "cloreai/ubuntu20.04-jupyter",
      "currency": "bitcoin",
      "spend": 1.3888888888888888e-7,
      "ct": 1667248597,
      "specs": {
        "mb": "Z590 GAMING X",
        "cpu": "11th Gen Intel(R) Core(TM) i9-11900 @ 2.50GHz",
        "cpus": "8/16",
        "ram": 62.67348861694336,
        "disk": "NVMe disk 247.3623046875GB",
        "disk_speed": 2000,
        "gpu": "1x NVIDIA GeForce GTX 1080 Ti",
        "gpuram": 11,
        "net": {
          "up": 26.38,
          "down": 118.42,
        }
      },
      "si": 6,
      "spot": true,
      "expired": true,
      "tcp_ports": []
    }
  ],
  "limit": 13,
  "code": 0
}
```

#### 6. `spot_marketplace` <a href="#id-6-spot_marketplace" id="id-6-spot_marketplace"></a>

**About**

Get spot marketplace for specific server

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Query string**

| Field    | Type  | Mandatory | Description      |
| -------- | ----- | --------- | ---------------- |
| `market` | `int` | Yes       | Unique server ID |

**Output**

| Field                       | Type     | Description                                          |
| --------------------------- | -------- | ---------------------------------------------------- |
| `code`                      | `int`    | Status code                                          |
| `exists`                    | `bool`   | Verification that the market exists                  |
| `market`                    | `object` | Marketplace                                          |
| `market.offers`             | `array`  | Rental offers for the server                         |
| `market.offers[x].offer_id` | `int`    | Unique offer ID                                      |
| `market.offers[x].bid`      | `float`  | Offered price per day                                |
| `market.offers[x].active`   | `bool`   | This offer is currently used                         |
| `market.offers[x].my`       | `bool`   | This offer is owned by me                            |
| `market.server`             | `object` | Server information                                   |
| `market.server.min_pricing` | `object` | Minimal offer price per day                          |
| `market.server.mrl`         | `int`    | Maximum rental length in seconds                     |
| `market.server.visibility`  | `string` | You can create offers only when visibility is public |
| `market.server.online`      | `bool`   | Server is online                                     |

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/spot_marketplace?market=6'
```

**Output:**

```
{
  "market": {
    "offers": [
      {
        "offer_id": 39,
        "bid": 0.0000042,
        "active": true,
        "my": true
      }
    ],
    "server": {
      "min_pricing": {
        "bitcoin": 0.000001
      },
      "mrl": 262800,
      "visibility": "public",
      "online": true
    }
  },
  "exists": true,
  "code": 0
}
```

#### 7. `set_server_settings` <a href="#id-7-set_server_settings" id="id-7-set_server_settings"></a>

**About**

Configure settings of server you are providing on [clore.ai](http://clore.ai/) marketplace

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field          | Type     | Mandatory | Description                             |
| -------------- | -------- | --------- | --------------------------------------- |
| `name`         | `string` | Yes       | User selected server name               |
| `availability` | `bool`   | Yes       | Can server be rented                    |
| `mrl`          | `int`    | Yes       | Maximum server rental length            |
| `on_demand`    | `float`  | Yes       | Price per day for your server on demand |
| `spot`         | `float`  | Yes       | Minimum price per day for SPOT offer    |

> Per-day prices are capped at 0.015 BTC / 600,000 CLORE / 1,000 USD.

**Output**

| Field  | Type  | Description |
| ------ | ----- | ----------- |
| `code` | `int` | Status code |

**Example**

Let's create a send proof for a transaction sent from the current wallet.

**Input:**

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "name": "Jan Vykydal",
    "availability":true,
    "mrl":96,
    "on_demand":0.0001,
    "spot":0.00000113
}' 'https://api.clore.ai/v1/set_server_settings'
```

**Output:**

```
{
  "code": 0
}
```

#### 8. `set_spot_price` <a href="#id-8-set_spot_price" id="id-8-set_spot_price"></a>

**About**

Set price per day on your SPOT market offer

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field           | Type    | Mandatory | Description                |
| --------------- | ------- | --------- | -------------------------- |
| `order_id`      | `int`   | Yes       | Unique offer ID            |
| `desired_price` | `float` | Yes       | Your offered price per day |

**Example**

Let's try to update spot market price

**Input 1 (Step down was too big):**

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "order_id":39,
    "desired_price":0.00000200
}' 'https://api.clore.ai/v1/set_spot_price'
```

**Possible output 1 (Step down was too big):**\
You can lower spot market offer price by max of 0.00000100 ₿

| Field      | Type     | Description                                                 |
| ---------- | -------- | ----------------------------------------------------------- |
| `code`     | `int`    | Status code                                                 |
| `error`    | `string` | Error description field                                     |
| `max_step` | `float`  | Lowest possible value to what you can currently lower price |

```
{
    "error":"exceeded_max_step",
    "max_step":0.0000032,
    "code":6
}
```

**Input 2 (Valid price step down):**

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "order_id":39,
    "desired_price":0.00000320
}' 'https://api.clore.ai/v1/set_spot_price'
```

**Possible output 2 (Valid price step down):**

```
{
    "code": 0
}
```

**Input 3 (Lower price even more after sending Input 2):**

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "order_id":39,
    "desired_price":0.00000220
}' 'https://api.clore.ai/v1/set_spot_price
```

**Possible output 3 (Lower price even more after sending Input 2):**\
You can lower spot price once in 600 seconds

| Field              | Type     | Description                                             |
| ------------------ | -------- | ------------------------------------------------------- |
| `code`             | `int`    | Status code                                             |
| `error`            | `string` | Error description field                                 |
| `time_to_lowering` | `float`  | Remaining time (sec) to next possibility to lower price |

```
{
    "error":"can_lower_every_600_seconds",
    "time_to_lowering":513,
    "code":6
}
```

#### 9. `cancel_order` <a href="#id-9-cancel_order" id="id-9-cancel_order"></a>

**About**

Cancel an active order or spot offer

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field   | Type     | Mandatory | Description                                                                                                                          |
| ------- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `id`    | `int`    | Yes       | Unique order/offer ID                                                                                                                |
| `issue` | `string` | No        | If you have encountered any issues with the server you can report them to [clore.ai](http://clore.ai/) team, maximum 2048 characters |

**Output**

| Field  | Type  | Description |
| ------ | ----- | ----------- |
| `code` | `int` | Status code |

**Example**

Cancel order/offer

**Input:**\
In this example we are reporting issues with GPU #1, if you don't have issues, don't include issue field.\
You can write any message to text field and we will investigate it

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "id":39,
    "issue":"GPU #1 Was overheating and throttling"
}' 'https://api.clore.ai/v1/cancel_order'
```

**Output:**

```
{
  "code": 0
}
```

#### 10. `create_order` <a href="#id-10-create_order" id="id-10-create_order"></a>

**About**

You can create spot offer or on demand order with this endpoint\
This endpoint also allows only 1 request in 5 seconds

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field                | Type     | Mandatory | Description                                                                                                                                                                   |
| -------------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `currency`           | `string` | Yes       | Currency name                                                                                                                                                                 |
| `image`              | `string` | Yes       | Valid image from dockerhub                                                                                                                                                    |
| `renting_server`     | `int`    | Yes       | ID of server you want to rent                                                                                                                                                 |
| `type`               | `string` | Yes       | `on-demand` OR `spot`                                                                                                                                                         |
| `spotprice`          | `float`  | Depends   | Offering price per day on spot market, required when making spot order                                                                                                        |
| `ports`              | `object` | No        | Port forwarding configuration, max 5 records                                                                                                                                  |
| `env`                | `object` | No        | <p>Environment variables, limited to 12000 characters in total when stringified.<br>Variable name - 128 symbols max<br>Variable value - 1536 symbols max</p>                  |
| `jupyter_token`      | `string` | No        | Jupyter token for images that has jupyter notebooks, maximum 32 characters \*                                                                                                 |
| `ssh_key`            | `string` | No        | SSH key for images with SSH, maximum 3072 characters \*                                                                                                                       |
| `ssh_password`       | `string` | No        | SSH password for images with SSH, maximum 32 characters \*                                                                                                                    |
| `command`            | `string` | No        | Command will be run on server after order creation                                                                                                                            |
| `required_price`     | `float`  | No        | Specify price for what you want to start the order, if machine owner changes the price, then order will not start (on demand only)                                            |
| `gpu_count`          | `int`    | No        | Number of GPUs to rent on servers that support [partial rental](/for-renters/partial-gpu-rental). Omit to rent the whole rig. On-demand only, homogeneous rigs only           |
| `gpu_indices`        | `[]int`  | No        | 0-based GPU slot indices to rent; length must equal `gpu_count`. Take free slots from `partial_gpu_rental.free_indices`. Omit to let the backend pick free GPUs automatically |
| `autossh_entrypoint` | `bool`   | No        | Use clore.ai entrypoint, that autometically deploy SSH server and custom `/root/onstart.sh` script                                                                            |

\* To fields marked with star you can only input characters from this regexp group `/^[a-zA-Z0-9\s-=.@+/]+$/`

**Output**

| Field  | Type  | Description |
| ------ | ----- | ----------- |
| `code` | `int` | Status code |

**Example**

**Input 1 (Create spot offer):**

```
curl -XPOST -H 'auth: 6FcuR7ibcwKR1Z32lEFoSotzUUtzKO2H' -H "Content-type: application/json" -d '
{
    "currency":"bitcoin",
    "image":"cloreai/ubuntu20.04-jupyter",
    "renting_server":6,
    "type":"spot",
    "spotprice":0.000001,
    "ports":{
        "22":"tcp",
        "8888":"http"
    },
    "env":{
        "VARIABLE_NAME":"VARIABLE_VALUE",
    },
    "jupyter_token":"hoZluOjbCOQ5D5yH7R",
    "ssh_password":"Hpcj08ZaOpCbTmn1Eu",
    "command":"#!/bin/sh\napt update -y && apt install htop"
}' 'https://api.clore.ai/v1/create_order'
```

**Output 1 (Create spot offer):**

```
{
  "code":0
}
```

**Input 2 (Create on demand order):**

```
curl -XPOST -H 'auth: 6FcuR7ibcwKR1Z32lEFoSotzUUtzKO2H' -H "Content-type: application/json" -d '
{
    "currency":"bitcoin",
    "image":"cloreai/ubuntu20.04-jupyter",
    "renting_server":6,
    "type":"on-demand",
    "ports":{
        "22":"tcp",
        "8888":"http"
    },
    "env":{
        "VARIABLE_NAME":"VARIABLE_VALUE",
    },
    "jupyter_token":"hoZluOjbCOQ5D5yH7R",
    "ssh_password":"Hpcj08ZaOpCbTmn1Eu",
    "command":"#!/bin/sh\napt update -y && apt install htop"
}' 'https://api.clore.ai/v1/create_order'
```

**Output 2 (Create on demand order):**

```
{
  "code":0
}
```

#### 11. `renter_fees` <a href="#id-11-renter_fees" id="id-11-renter_fees"></a>

**About**

Get the current fee configuration: base renter fees, PoH fee reduction settings, order creation fees, and your PoH balance

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Output**

| Field                              | Type     | Description                                           |
| ---------------------------------- | -------- | ----------------------------------------------------- |
| `code`                             | `int`    | Status code                                           |
| `renter_fees.on_demand`            | `object` | Base fee (%) per currency for on-demand orders        |
| `renter_fees.spot`                 | `object` | Base fee (%) per currency for spot orders             |
| `poh_config.max_reduction_percent` | `int`    | Maximum base-fee reduction through PoH (%)            |
| `poh_config.max_poh_amount`        | `int`    | PoH balance at which the maximum reduction is reached |
| `creation_fees.on_demand`          | `object` | Order creation fee per currency                       |
| `creation_fees.spot`               | `object` | Spot order creation fee per currency                  |
| `poh_balance`                      | `float`  | Your PoH CLORE balance                                |

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/renter_fees'
```

**Output:**

```
{
  "renter_fees": {
    "on_demand": {
      "base": { "bitcoin": 5, "CLORE-Blockchain": 5, "USD-Blockchain": 5 }
    },
    "spot": {
      "base": { "bitcoin": 1.25, "CLORE-Blockchain": 1.25, "USD-Blockchain": 1.25 }
    }
  },
  "poh_config": { "max_reduction_percent": 50, "max_poh_amount": 2000000 },
  "creation_fees": {
    "on_demand": { "bitcoin": 3e-7, "CLORE-Blockchain": 0.5, "USD-Blockchain": 0.1 },
    "spot": { "bitcoin": 1e-7, "CLORE-Blockchain": 0.05, "USD-Blockchain": 0.005 }
  },
  "poh_balance": 5000,
  "code": 0
}
```

#### 12. `cancel_orders` <a href="#id-12-cancel_orders" id="id-12-cancel_orders"></a>

**About**

Cancel multiple orders at once (up to 256 per request)

**Headers**

| Field          | Type     | Mandatory | Description                |
| -------------- | -------- | --------- | -------------------------- |
| `auth`         | `string` | Yes       | API token                  |
| `Content-type` | `string` | Yes       | Must be `application/json` |

**Body**

| Field       | Type    | Mandatory | Description                      |
| ----------- | ------- | --------- | -------------------------------- |
| `order_ids` | `[]int` | Yes       | Order IDs to cancel, maximum 256 |

**Output**

| Field              | Type    | Description                           |
| ------------------ | ------- | ------------------------------------- |
| `code`             | `int`   | Status code                           |
| `failed_to_cancel` | `[]int` | Order IDs that could not be cancelled |

**Input:**

```
curl -XPOST -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' -H "Content-type: application/json" -d '{
    "order_ids": [38, 39]
}' 'https://api.clore.ai/v1/cancel_orders'
```

**Output:**

```
{
  "failed_to_cancel": [],
  "code": 0
}
```

#### 13. `poh_balance` <a href="#id-13-poh_balance" id="id-13-poh_balance"></a>

**About**

Get your Proof of Holding balance: items, totals, and PoH Marketplace offer amounts

**Headers**

| Field  | Type     | Mandatory | Description |
| ------ | -------- | --------- | ----------- |
| `auth` | `string` | Yes       | API token   |

**Output**

| Field                                         | Type       | Description                             |
| --------------------------------------------- | ---------- | --------------------------------------- |
| `code`                                        | `int`      | Status code                             |
| `items`                                       | `[]object` | Your PoH items with individual balances |
| `balance.total`                               | `float`    | Total PoH balance across all items      |
| `balance.free_amount`                         | `float`    | PoH balance not committed to offers     |
| `balance.reward_amount`                       | `float`    | Balance eligible for rewards            |
| `balance.offers.rented_amount`                | `float`    | Amount rented out to other users        |
| `balance.offers.rented_on_marketplace_amount` | `float`    | Amount listed on the PoH Marketplace    |
| `balance.offers.leased_amount`                | `float`    | Amount leased from other users          |

**Input:**

```
curl -XGET -H 'auth: b8qwqRAL5W7YDyDJeB4XANVvKndbrrPk' 'https://api.clore.ai/v1/poh_balance'
```

**Output:**

```
{
  "items": [ { "balance": 5000 } ],
  "balance": {
    "total": 5000,
    "free_amount": 4000,
    "reward_amount": 5000,
    "offers": {
      "rented_amount": 0,
      "rented_on_marketplace_amount": 1000,
      "rented_finished_amount": 0,
      "leased_amount": 0
    }
  },
  "code": 0
}
```

#### 14. `server-config.json` <a href="#id-14-server-config-json" id="id-14-server-config-json"></a>

**About**

Get static API configuration data. No authentication required

**Input:**

```
curl -XGET 'https://api.clore.ai/v1/server-config.json'
```

#### 15. `get_relay` <a href="#id-15-get_relay" id="id-15-get_relay"></a>

**About**

Get TCP forwarding relay nodes for your detected country. No authentication required

**Output**

| Field     | Type     | Description           |
| --------- | -------- | --------------------- |
| `code`    | `int`    | Status code           |
| `country` | `string` | Detected country code |
| `nodes`   | `object` | Available relay nodes |

**Input:**

```
curl -XGET 'https://api.clore.ai/v1/get_relay'
```


# Spot Market

Auction-style interruptible rentals: bid your own price per day and rent GPUs at the lowest cost on the marketplace.

SPOT is the auction side of the Clore.ai marketplace. Instead of paying the listed on-demand price, you place a **bid**: your own price per day for a server. While your offer is the winning one, the server runs your workload and you are billed per minute at your bid. It is the cheapest way to rent, in exchange for the risk of being interrupted.

|                  |                                                                                          |
| ---------------- | ---------------------------------------------------------------------------------------- |
| **Pricing**      | Your bid, at or above the host's minimum spot price                                      |
| **Billing**      | Per minute, at your bid price, only while your offer holds the server                    |
| **Interruption** | Another renter can outbid you at any time; an on-demand rental also takes the server     |
| **Base fee**     | 2.5% (vs 10% on-demand), reducible by up to 50% with [PoH](/proof-of-holding/overview)   |
| **Scope**        | Whole servers only ([partial rental](/for-renters/partial-gpu-rental) is on-demand only) |

## How Bidding Works

1. Pick a server on the marketplace and press **Rent**. Go through the usual steps (image, configuration), and at the **Confirm & pay** step switch to **Rent on spot**.
2. Enter your bid: any price per day at or above the host's minimum spot price. If it beats the current winning offer, your order starts.
3. While your offer holds the server, you are billed per minute at your bid. If someone outbids you, your order ends with the **Outbid** status and billing stops immediately.
4. Your offer competes with others on the same server. You can check the queue and manage bids from **My Orders**.

### Managing Your Bid

* **Raising** your bid is unrestricted: do it any time to defend the server.
* **Lowering** is rate-limited to protect hosts: at most once every 10 minutes, by a limited step per move, and never below the host's minimum spot price.

{% hint style="info" %}
Spot orders cannot be extended and can end at any moment. Checkpoint your work: save intermediate results to persistent storage or externally, so an interruption costs you minutes, not hours.
{% endhint %}

## What Runs Well on Spot

* Cryptocurrency mining and other restartable compute
* Batch processing and render queues with checkpoints
* Hyperparameter sweeps and experiments
* Anything fault-tolerant where price matters more than uptime

Not sure whether you need Spot or On-Demand? See the side-by-side guide:

{% content-ref url="/pages/E7V0uUw9wjm5QB2hYGx2" %}
[On-Demand vs Spot](/for-renters/on-demand-vs-spot)
{% endcontent-ref %}

## For Hosts

You set the **minimum spot price** per server (or per GPU, with optional USD-based auto-recalculation) in your server settings. Spot keeps your hardware earning between on-demand rentals: an on-demand order always takes priority and replaces the active spot order.

{% content-ref url="/pages/BP9ytpSVPwfichcYVOR0" %}
[Server Management](/for-hosts/server-settings)
{% endcontent-ref %}

## Automation

Spot is fully scriptable through the public API: `create_order` with `"type": "spot"` and `spotprice`, `spot_marketplace` to read the offer queue for a server, and `set_spot_price` to manage your bid. The [Python SDK and CLI](/developers/python-sdk) cover the same operations.

{% content-ref url="/pages/VsUJqQ2pg4ZUv2TbJaMy" %}
[API](/for-hosts/api)
{% endcontent-ref %}


# Overview

<div data-full-width="true"><figure><img src="/files/I6weUOnqVTHp5L52qmWs" alt=""><figcaption></figcaption></figure></div>

## What is Proof of Holding (PoH)?

Proof of Holding is Clore.ai's staking system that rewards users for holding CLORE tokens. Simply stake your tokens and earn passive income — no lockups, no penalties, full control of your funds.

***

## PoH Staking Rewards

| Parameter              | Value                                          |
| ---------------------- | ---------------------------------------------- |
| **Daily distribution** | 82,000 CLORE                                   |
| **APY**                | Up to 15%                                      |
| **Emission decay**     | -30% per year                                  |
| **Minimum stake**      | No minimum                                     |
| **Claim period**       | 7 days (rewards from day 1 claimable on day 8) |
| **Expiration**         | Unclaimed rewards expire after 30 days         |

### How Rewards Work

1. Stake CLORE tokens via the [PoH section](https://clore.ai/proof-of-holding)
2. Rewards accumulate daily based on your staked amount
3. Claim rewards after 7-day maturation period
4. Rewards expire if not claimed within 30 days

{% content-ref url="/pages/Ni1bVthc7jH6E6wnugoc" %}
[PoH Staking](/proof-of-holding/poh-staking)
{% endcontent-ref %}

***

## Fee Reduction

Staking CLORE reduces marketplace **base fees** for **both renters and hosters**:

| CLORE in PoH | Fee Reduction |
| ------------ | ------------- |
| 0            | 0%            |
| 200,000      | 5%            |
| 500,000      | 13%           |
| 1,000,000    | 25%           |
| 1,500,000    | 38%           |
| 2,000,000+   | 50% (maximum) |

The reduction is calculated using a simple proportional formula:

```
reduction = round(min(your_poh / 2,000,000, 1) × 50)%
```

**Key changes:**

* Applies to **both renter and hoster** fees (previously renter-only)
* Works for **all currencies** — same formula regardless of payment currency
* Reduces **base fees only** (does not affect extra currency fees)

> **Example:** With 1,000,000 CLORE in PoH, both you and the host get a 25% reduction on your respective halves of the base fee.

***

## How to Stake

1. Go to [Marketplace → Proof of Holding](https://clore.ai/proof-of-holding)
2. Connect your MetaMask wallet
3. Sign the verification message
4. Done — your CLORE balance is now staked

Rewards accumulate automatically. Claim after 7-day maturation period.

***

## POH Marketplace

Don't have CLORE tokens? Rent them on the POH Marketplace!

The POH Marketplace lets you:

* **Rent CLORE** to reduce your rental fees
* **Lend CLORE** to earn passive income (5-250% APY)

Rental terms: 1 hour to 90 days, minimum 500 CLORE.

{% content-ref url="/pages/ZXxfwYY21Yz9PCv9X7ja" %}
[PoH Marketplace](/proof-of-holding/poh-marketplace)
{% endcontent-ref %}

> 📚 See also: [Proof of Holding 2.0 — GPU-Based Rewards and Optimized Earnings](https://blog.clore.ai/proof-of-holding-20-gpu-based-rewards-and-optimized-earnings/)


# PoH Staking

**PoH Staking** allows CLORE holders to earn passive income by staking their tokens. Connect your wallet, sign a message, and start earning — it's that simple.

## Benefits

* **Earn up to 15% APY** on your CLORE holdings
* **82,000 CLORE distributed daily** to stakers
* **Reduce marketplace fees** by up to 50%
* **No lockup** — withdraw anytime without penalties
* **Full control** — tokens stay in your wallet

***

## How to Stake

### Step 1: Go to PoH Section

Navigate to [clore.ai/proof-of-holding](https://clore.ai/proof-of-holding)

### Step 2: Connect MetaMask

Click **Connect Wallet** and select MetaMask. Approve the connection request in MetaMask.

### Step 3: Sign Verification

Sign the message in MetaMask to verify wallet ownership. This does not transfer any tokens.

### Step 4: Done

Your CLORE balance is now staked. Rewards accumulate automatically.

***

## Claiming Rewards

| Parameter              | Value                    |
| ---------------------- | ------------------------ |
| Maturation period      | 7 days                   |
| Day 1 reward claimable | Day 8                    |
| Expiration             | 30 days after maturation |

Go to [clore.ai/proof-of-holding](https://clore.ai/proof-of-holding) and click **Claim** to collect your rewards.

> **Important:** Unclaimed rewards expire after 30 days. Claim regularly!

***

## Fee Reduction

Staking CLORE reduces marketplace base fees for **both renters and hosters**:

| CLORE in PoH | Fee Reduction |
| ------------ | ------------- |
| 0            | 0%            |
| 200,000      | 5%            |
| 500,000      | 13%           |
| 1,000,000    | 25%           |
| 2,000,000+   | 50% (maximum) |

Reduction scales linearly: `round(min(poh / 2,000,000, 1) × 50)%`

See [Fee Structure](/for-renters/fee-structure) for full details on fee reduction mechanisms.


# PoH Marketplace

<div data-full-width="true"><figure><img src="/files/6TOJ1CgqmlNILAN3QAx4" alt=""><figcaption></figcaption></figure></div>

## Introduction to the PoH Marketplace

The Proof of Holding (POH) Marketplace is an innovative platform within Clore.ai that allows users to maximize the utility of their CLORE coins. The marketplace serves a dual purpose: it enables coin holders to earn passive income by renting out their CLORE coins, and it allows others to rent these coins to increase their earnings from the network’s rewards system. By participating in the PoH Marketplace, users can either profit from holding CLORE coins or enhance their staking power to receive a larger share of network rewards.

## Purpose and Benefits of the PoH Marketplace

**PoH Marketplace** is designed as a convenient, transparent, and secure platform with the primary goal of reducing server rental fees and encouraging user participation in PoH staking.

For renters, it’s a way to **reduce rental fees** by up to **50%** by having coins in PoH.\
For coin holders, it’s an opportunity to **stake their coins in PoH** and earn rewards without selling their assets.

This mechanism highlights **Clore.ai’s** commitment to building a user-focused ecosystem where everyone can effectively use **CLORE** both as a payment method and a source of income.

The PoH Marketplace is built into the platform: open [clore.ai/poh-marketplace](https://clore.ai/poh-marketplace) to browse offers, lend, or rent CLORE directly from your account.

## **Key Functions of the PoH Marketplace:**

#### 1. Listing and Filtering Offers:

• Users can view all offers on the marketplace, including those seeking to rent coins and those offering coins for rent.

• Filters allow users to narrow down listings by type (rental or lending), amount of coins, rental duration, and cost.

#### 2. Creating Offers:

• Users can create their own offers, specifying the amount of CLORE coins (minimum 500 coins), rental duration (from 1 hour to 90 days), and the annual rental rate (from 5% to 250%).

• An online calculator dynamically calculates the exact amount of coins users will pay or receive based on the entered details. For instance, if an owner wants to rent out 10,000 CLORE coins for 3 days at an annual rate of 250%, the calculator will show the precise earnings before the offer is published.

#### 3. Managing Offers:

• Owners and Renters can manage their active offers, including editing, deleting, or accepting counteroffers from other users. Users can also track the status of their offers and receive notifications about important updates, such as when someone has accepted their offer or made a counteroffer.

#### 4. Negotiations and Order Completion:

• The marketplace supports negotiations, where users can propose counteroffers by adjusting the amount of coins, rental period, or rate. These negotiations allow for flexibility and mutual agreement between parties.

• Both Owners and Renters can close active orders, either upon completion or early, with a small fee applied for early closures to compensate the counterparty.

#### 5. Order History and Repeat Offers:

• Users can view the history of their completed orders, including detailed information such as the rental period, amount of CLORE coins, annual percentage rate, and total profit earned.

• The system allows users to repeat previously completed orders, with pre-filled parameters to streamline the process. Users can adjust any details before reposting the offer to the marketplace.

***

## Related Pages

{% content-ref url="/pages/Ni1bVthc7jH6E6wnugoc" %}
[PoH Staking](/proof-of-holding/poh-staking)
{% endcontent-ref %}

{% content-ref url="/pages/y32ehuugYQLXgvnOYh9m" %}
[Use cases](/proof-of-holding/use-cases)
{% endcontent-ref %}


# Use cases

## **Case 1**

**Listing Coins for Rent:**

**Role:** Coin Owner/ Coin renter

**Action:** A user who holds Clore coins and has verified them in the Proof of Holding (PoH) system can list their coins for rent on the PoH marketplace. And verified user which wants to apply for a coins rent can add his offer to the marketplace.

<figure><img src="/files/qucoJy3aAwUjVnhMq1mH" alt="" width="353"><figcaption></figcaption></figure>

**Process:**

The user navigates to the PoH marketplace and selects the option to list coins for rent.

Where specify the amount of coins (minimum 500), the rental duration (minimum 1 hour, maximum 90 days), and the desired annual rental rate (between 5% and 250%).

The system calculates the potential earnings using an online calculator.

Upon review, the user publishes the offer, which then appears in the marketplace listings.

## **Case 2**

**Searching for Coins to Rent:**

**Role:** Renter

**Action:** A user looking to rent Clore coins can search the PoH marketplace for suitable rental offers.

<div data-full-width="false"><figure><img src="/files/RZN5EhSOzGU8hKbmE0Zu" alt=""><figcaption><p>applying for the offer from the marketplace page</p></figcaption></figure></div>

<div><figure><img src="/files/AieVzfd0qg2Nmu1BBBLP" alt=""><figcaption><p>proceeding selected offer</p></figcaption></figure> <figure><img src="/files/OR1qouhDMEMbPtlsJ0Wr" alt=""><figcaption><p>filtering</p></figcaption></figure></div>

**Process:**

The user browses the marketplace listings and uses filters to narrow down search results based on the amount of coins available, rental duration, and rental rate.

Upon finding a suitable offer, the renter can agree to the terms and initiate the rental agreement.

The rented coins are transferred to the renter's account for the agreed-upon duration, and the associated rental fee is deducted or credited based on the terms.

## **Case 3**

**Creating a Custom Rental Offer:**

**Role:** Renter/owner

**Action:** If a user cannot find an existing rental offer that meets their requirements, they can create a custom offer seeking to rent coins.

<figure><img src="/files/DiR8q7Q7bv4sdRRmxELK" alt=""><figcaption><p>selecting offer in the marketplace page</p></figcaption></figure>

<figure><img src="/files/DXaGvQFnIVyUJcd5aTUV" alt="" width="370"><figcaption><p>managing the offer parameters</p></figcaption></figure>

**Process:**

**The user specifies the number of coins they wish to rent, the rental period, and the annual rental rate they are willing to pay.**

**This custom offer is posted to the marketplace, where it can be seen by coin owners.**

**Interested coin owners can then respond by agreeing to the rental terms, effectively creating a new rental transaction.**

## **Case 4**

**Managing Active Rentals:**

**Role:** Coin Owner or Renter

**Action:** Both parties can manage active rentals through the PoH marketplace interface.

<figure><img src="/files/FxeIVYFdHZTERZ0LFmlD" alt=""><figcaption></figcaption></figure>

**Process:**

**Coin owners can monitor which of their coins are rented out, track the rental duration, and view their earnings in real time.**

**Renters can see the duration remaining on their rented coins and manage how they utilize these coins during the rental period.**

**Both parties receive notifications regarding the status of their rentals, including the start and end of the rental period and any changes in rental conditions.**

## **Case 5**

**Ending a Rental Early:**

**Role:** Coin Owner or Renter

**Action:** Either party can end the rental agreement early, subject to conditions.

<div><figure><img src="/files/mZnLd2bE5wC3cd8Qq2oJ" alt=""><figcaption></figcaption></figure> <figure><img src="/files/8hXGtcwHPdlVleNIjZ6A" alt=""><figcaption></figcaption></figure></div>

**Process:**

**If a coin owner needs to regain control of their rented coins before the agreed period ends, they can terminate the rental. A small penalty fee (3-5% of the daily profit) is charged for early termination.**

**Renters can also decide to end the rental early if their needs change. They will stop paying the rental fee from the moment the coins are returned.**

**Early termination must be confirmed by both parties, ensuring clear communication and agreement.**

## **Case 6**

**Receiving and Reviewing Notifications:**

**Role:** Coin Owner or Renter

**Action:** Users receive notifications about the status of their offers and rentals. And both users can communicate with each other via messenger

<div><figure><img src="/files/pClR1NroLk17w30MgSq9" alt=""><figcaption><p>notification to owner his cois got leased</p></figcaption></figure> <figure><img src="/files/AjoVwzQDi2vGNxEGAPuK" alt="" width="563"><figcaption><p>notification to offer owner with modified parameters</p></figcaption></figure></div>

**Process:**

**Notifications include confirmations of successful listings, rentals, and terminations.**

**Users are informed when someone takes up their rental offer, or if a counter-offer is made.**

**System-generated alerts help maintain transparency and ensure all parties are informed about the state of their rental transactions.**

## **Case 7**

**Reviewing Rental History:**

**Role:** Coin Owner or Renter

**Action:** Users can view the history of their rental activities.

<figure><img src="/files/S9smhwA3k3VfBmRhuZrN" alt="" width="336"><figcaption><p>finished order with details and status</p></figcaption></figure>

**Process:**

**The marketplace provides a detailed history section where users can review all completed rental transactions, including past offers, amounts, rental rates, and durations.**

**Users can access detailed reports on each transaction, helping them track performance and optimize future rental strategies. Users can repeat the offer to publish it again to the marketplace.**


# Team

<div data-full-width="true"><figure><img src="/files/QtDoss9AxHswNsQixIpw" alt=""><figcaption></figcaption></figure></div>

Approved by [Certik](https://skynet.certik.com/projects/clore-ai)

![card-img](https://clore.ai/images/Jakub.webp)

## Jakub

Founder

The driving force behind Clore, Jakub boasts an extensive background in developing mining pools for Ethereum and other cryptocurrencies. His past experiences with IT infrastructure security in Czech companies have equipped him with the expertise needed to spearhead Clore. Serving not just as the CEO, Jakub is also the chief developer, visionary, and architect of the Clore project.

![card-img](https://clore.ai/images/John.webp)

## John

CEO

With a rich history as a miner, trader, and SEO specialist, John is a seasoned player in the crypto world, having been involved for over six years. His vast experience with blockchain startups is invaluable to Clore. As the CEO, he manages all operational facets, proving himself to be a versatile asset or, as some might say, a "one-man orchestra.

![card-img](https://clore.ai/images/Loony.webp)

## Loony

Head of product

An expert in Machine Learning, Loony has previously served as a Software & ML Engineer for various startups. At Clore, he takes the helm of product development. Despite his busy schedule, Loony is renowned for making himself available to the community, addressing queries across crypto chats. Some say he manages to squeeze 25 hours into his day!

![card-img](https://clore.ai/images/Val.webp)

## Val

CTO

As the CTO of Clore, Val is the mastermind behind the platform’s technical strategy. With an extensive background in software engineering and leadership roles in tech startups, Val is responsible for overseeing the development team and guiding the technical vision of Clore. His ability to foresee industry trends and adapt to emerging technologies ensures that Clore remains a leader in the crypto space, driving innovation and maintaining the platform’s competitive edge.

![card-img](https://clore.ai/images/Corey.webp)

## Corey

Desktop / Mobile Apps Lead

Corey leads the development of Clore’s desktop and mobile applications, ensuring a seamless and intuitive user experience. Previously, he created the Vidulum Wallet, showcasing his expertise in secure and efficient app development. His commitment to innovation and quality helps Clore.ai stay at the forefront of the GPU leasing marketplace.

![card-img](https://clore.ai/images/Nik.webp)

## Nik

Developer

Nik is a versatile developer with a knack for problem-solving and optimization. His previous work on high-performance applications and system integrations has equipped him with the skills needed to enhance Clore’s platform. At Clore, Nik focuses on refining the platform’s performance, ensuring that every operation runs smoothly and efficiently, contributing to an overall superior user experience.

![card-img](https://clore.ai/images/Oksana.webp)

## Oksana

Developer

Oksana is a highly skilled developer with a passion for coding that drives her to constantly push the boundaries of what’s possible. With experience in both frontend and backend development, she brings versatility and expertise to the Clore team. Her innovative solutions and deep understanding of blockchain technologies play a crucial role in advancing the platform’s capabilities, ensuring that Clore remains at the cutting edge of the industry.

![card-img](https://clore.ai/images/Egor.webp)

## Egor

Developer

Egor is a backend wizard at Clore, bringing years of experience in developing scalable systems and robust architectures. His deep knowledge of distributed systems and blockchain technology enables him to craft solutions that are both efficient and secure. Egor’s work behind the scenes ensures that Clore can handle the demands of a growing user base, making him a key contributor to the platform’s success.

![card-img](https://clore.ai/images/Daniel.webp)

## Daniel

Developer

Daniel is a frontend specialist who excels at transforming complex functionalities into user-friendly interfaces. With a background in UI/UX design and development, he ensures that every feature of Clore is not only functional but also intuitive and accessible. Daniel’s ability to blend design with functionality helps create a seamless experience for Clore’s users, making him a vital part of the team.

![card-img](https://clore.ai/images/Max.webp)

## Max

QA Engineer

Max is the guardian of quality at Clore, meticulously testing every aspect of the platform to ensure it meets the highest standards. With a background in software testing across fintech and blockchain industries, Max is adept at identifying potential issues before they reach the users. His attention to detail and commitment to excellence make him an indispensable part of the Clore team, ensuring a smooth and reliable user experience.

![card-img](https://clore.ai/images/Nikita.webp)

## Nikita

Web / Front

With a knack for creating intuitive and responsive web interfaces, Nikita has previously worked on diverse projects, from e-commerce sites to complex web applications. At Clore, he ensures that the platform is not just functional but also user-friendly, making every user's journey seamless and enjoyable.

![card-img](https://clore.ai/images/Vlad.webp)

## Vlad

Designer

An artisan in the realm of design, Vlad's previous endeavors include crafting visual experiences for startups and established brands alike. At Clore, he's responsible for the platform's aesthetic appeal, ensuring every pixel aligns with the company's vision and resonates with its users.

![card-img](https://clore.ai/images/anon.svg)

## Anon

Developer

A coding prodigy, he has previously worked on developing intricate algorithms and backend systems for tech startups. At Clore, he's instrumental in refining the platform's backend mechanics, ensuring scalability and efficiency as the user base grows.

![card-img](https://clore.ai/images/anon.svg)

## Anon

Developer

Known for his problem-solving prowess, he has been pivotal in several tech projects, focusing on optimization and system integrations. In Clore, he collaborates closely with the team, fine-tuning the platform's features and ensuring a smooth user experience.


# Links

**Website**: <https://clore.ai/>

**Marketplace**: <https://clore.ai/marketplace>

**POH Marketplace**: [https://clore.ai/poh-marketplace](https://clore.ai/poh-marketplace?amount_min=40000)

**Roadmap**: <https://clore.ai/roadmap>

**Proof of Holding**: <https://clore.ai/pdf/Proof-Of-Holding-Whitepaper>

**BitcoinTalk**: <https://bitcointalk.org/index.php?topic=5428175>

**Twitter**: <https://twitter.com/clore_ai>

**Telegram**: <https://t.me/clorechat>

**Discord**: <https://discord.gg/clore-ai>


# FAQ

<div data-full-width="true"><figure><img src="/files/MtAciAjG3T9RNxye8Lol" alt=""><figcaption></figcaption></figure></div>

## General

**What is Clore.ai?**

Clore.ai is a decentralized peer-to-peer GPU marketplace that connects GPU owners (hosts) with users who need computing power (renters). It's built for AI development, machine learning, rendering, and other GPU-intensive workloads.

**How does Clore.ai work?**

Clore.ai connects GPU owners with those needing computing power through a peer-to-peer marketplace. Hosts list their servers with pricing, renters browse and rent them by deploying Docker containers. Transactions are facilitated with Clore Coin and other supported currencies.

**What is Clore Coin (CLORE)?**

Clore Coin is the native cryptocurrency of the Clore.ai platform, used for transactions, fee payments, and rewarding users through the Proof of Holding (PoH) system. Paying with CLORE avoids extra currency conversion fees.

**What type of GPUs does Clore.ai offer?**

Clore.ai provides a wide range of GPUs, including NVIDIA GeForce, RTX, and professional-grade cards suitable for AI/ML, 3D rendering, and mining. Available hardware depends on what hosts list on the marketplace at any given time.

**Is Clore.ai decentralized?**

Yes. Clore.ai is a P2P marketplace — servers are owned and operated by individual hosts around the world. There is no single centralized provider. This means hardware availability, uptime, and performance vary by host.

**What can I use Clore.ai for?**

Common use cases include AI/ML model training and inference, LLM fine-tuning, image/video rendering, scientific computing, and cryptocurrency mining. Any GPU-intensive Docker-based workload is supported.

**What is the minimum and maximum duration for GPU leasing?**

Clore.ai allows GPU leasing for a minimum of 6 hours. The maximum duration is 3000 hours (approximately 125 days), offering flexibility for both short-term projects and long-term needs.

***

## Account & Security

**How do I create an account on Clore.ai?**

Visit [clore.ai](https://clore.ai) and register with your email address. After email verification, you can deposit funds and start renting or listing servers.

**How do I enable two-factor authentication (2FA)?**

Go to your account security settings and enable 2FA using an authenticator app (e.g., Google Authenticator or Authy). 2FA is strongly recommended to protect your account and funds.

**What do I do if I lose access to my 2FA device?**

Contact Clore.ai support with your account details and proof of ownership. Recovery without 2FA requires manual verification by the team.

**How do I reset my password?**

Use the "Forgot Password" link on the login page. A reset link will be sent to your registered email address.

**How do I delete my account?**

Contact Clore.ai support to request account deletion. Ensure you withdraw all funds before submitting the request, as balances on a deleted account may not be recoverable.

**Does Clore.ai require KYC (Know Your Customer) verification?**

Basic usage does not require KYC. However, for large withdrawals or compliance-related cases, the platform may request identity verification. Check the current requirements in your account settings.

**How do I generate an API key?**

Go to your account settings and navigate to the API section. Generate a new token there. Keep your API key private — it grants full access to your account via the API.

***

## Billing & Payments

**What payment methods does Clore.ai accept?**

You can deposit and pay using **Clore Coin (CLORE)**, **Bitcoin (BTC)**, **USDT**, and **USDC** (on Ethereum ERC-20 or BNB Chain BEP-20). Paying with CLORE is recommended as it carries no extra currency conversion fees.

**Can I deposit USDT or USDC?**

Yes. Clore.ai supports USDT and USDC deposits on **two networks: Ethereum (ERC-20) and BNB Chain (BEP-20)**. Both stablecoins credit to a unified USD balance. Your USDT/USDC deposit address is the same as your CLORE deposit address. Send only on Ethereum or BNB Chain; tokens sent on other networks (Polygon, Tron, etc.) cannot be recovered.

**Can I pay for GPU rentals directly with USDT or USDC?**

Yes. USDT/USDC deposits credit a unified USD balance, and you can pay for GPU rentals directly from it, alongside CLORE and Bitcoin. Note that non-CLORE payments carry an extra hoster fee (see [Fee Structure](/for-renters/fee-structure)).

**Are there any fees on Clore.ai?**

Yes. On-demand orders carry a 10% base fee. Spot orders carry a 2.5% base fee. Both fees can be reduced by up to 50% through the Proof of Holding (PoH) system. Additionally, paying with non-CLORE currencies (e.g., Bitcoin) incurs an extra hoster fee that can be reduced to 0% via the MFP Lock system.

**What are extra hoster fees?**

When paying with non-CLORE currencies, an additional fee is charged. This fee can be reduced to 0% by locking CLORE in the [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) system. CLORE payments have no extra fees.

**What is the minimum deposit amount?**

Minimum deposit amounts vary by currency and network conditions. Check the deposit page in your account for current minimums. Deposits below the minimum may be lost or not credited.

**How do I withdraw funds?**

Go to your account wallet section, select the currency, enter your external wallet address, and submit a withdrawal request. Withdrawal fees vary by currency (e.g., for Bitcoin, the fee is 0.00005 BTC).

**Are there refunds if a server goes offline during my rental?**

Clore.ai is a P2P marketplace — hosts are independent operators. If a server goes offline mid-rental, the platform may provide partial compensation depending on circumstances. When cancelling an order and reporting issues via the `cancel_order` API or UI, the team investigates reported problems. For reliable workloads, use on-demand leasing which provides non-interruptible sessions.

**How long do deposits take to confirm?**

Confirmation times depend on the blockchain network. Bitcoin typically requires a few confirmations (10–60 minutes). CLORE, USDT, and USDC typically confirm within a few minutes depending on network congestion.

***

## For Renters

**How do I rent a GPU on Clore.ai?**

Browse the marketplace, filter by GPU type, price, and specs. Select a server, choose on-demand or spot leasing, configure your Docker image and port forwarding, and place your order. The container starts within minutes.

**What is the difference between Spot and On-Demand leasing?**

**On-Demand** leasing guarantees your rental for the full duration — no other user can outbid you. It carries a 10% base fee. **Spot** leasing is cheaper (2.5% base fee) but interruptible — another user can outbid your offer and take over the server. Use spot for fault-tolerant workloads like mining; use on-demand for training runs and critical jobs.

**How do I connect via SSH?**

When creating an order, specify your SSH public key in the `ssh_key` field or set an `ssh_password`. After the container starts, check your order details for the forwarded port and cluster endpoint (e.g., `n1.c1.clorecloud.net:10000`). Connect with: `ssh root@n1.c1.clorecloud.net -p 10000`.

**What Docker images can I use?**

Any public Docker Hub image is supported. Clore.ai also provides pre-built images (e.g., `cloreai/ubuntu20.04-jupyter`) with common ML frameworks, Jupyter, and SSH pre-configured. You can use your own custom images from Docker Hub.

**Can I use a Jupyter notebook on my rented server?**

Yes. Use a Jupyter-enabled image (e.g., `cloreai/ubuntu20.04-jupyter`) and forward port `8888` via `http`. Set a `jupyter_token` in your order configuration. The notebook will be accessible via an HTTPS proxy URL.

**Can I run a custom startup script?**

Yes. Pass a shell script in the `command` field of `create_order`. It runs after container startup. Alternatively, enable `autossh_entrypoint` to have Clore.ai automatically deploy an SSH server and execute `/root/onstart.sh`.

**How do I forward ports?**

In the order configuration, specify up to 5 port forwarding rules. Use `"tcp"` for raw TCP ports (e.g., SSH on port 22) and `"http"` for a port you want exposed via HTTPS proxy (e.g., Jupyter on 8888).

**What happens to my data when the rental ends?**

Data stored inside the container is **not persistent** by default — it is lost when the order expires or is cancelled. Save important data to external storage (S3, Backblaze, etc.) before your rental ends. Hosts may optionally support persistent storage — check server details.

**Can I cancel an order early?**

Yes. Cancel via the UI or the `cancel_order` API endpoint. If you experienced problems with the server, include an `issue` description — the Clore.ai team will investigate it.

**How do I use order templates?**

The Clore.ai web UI allows saving order configurations as templates, so you can quickly redeploy the same Docker image, environment variables, and port settings on different servers without re-entering everything.

**What should I do if my container won't start?**

Check that your Docker image name is valid and publicly accessible on Docker Hub. Verify that your port configuration is correct. If the server appears online but the container fails, cancel the order and report the issue. Try a different server.

**How do I set environment variables for my container?**

Pass environment variables in the `env` field of `create_order` as a key-value object. Variable names are limited to 128 characters and values to 1536 characters. The total stringified env object is limited to 12,000 characters.

***

## For Hosts

**How do I list my server on Clore.ai?**

Download and install the Clore.ai host software on your server, connect it to your account using the initialization token, configure pricing and visibility, and set it to public. Your server will appear on the marketplace when online.

**What are the minimum hardware requirements to become a host?**

There are no hard-coded minimum specs, but servers with at least one NVIDIA GPU and a stable internet connection are recommended. Servers with better GPUs, more RAM, fast NVMe storage, and high-speed networking will attract more renters and earn more.

**How is host pricing set?**

You set your own on-demand price per day and the minimum spot price per day. Prices are set in BTC and CLORE, or pegged to USD with automatic recalculation at the current rate. You can update pricing anytime via `set_server_settings`.

**How do hosts earn money?**

Hosts earn the rental price minus the platform fee. The fee hosts pay depends on the renter's payment currency: CLORE payments incur no extra fee; non-CLORE payments carry an additional fee that can be reduced to 0% via MFP Lock.

**What is MFP Lock?**

MFP (Maximum Fair Price) Lock allows hosts to lock CLORE tokens to reduce or eliminate the extra fee charged when renters pay with non-CLORE currencies. Locking enough CLORE reduces this fee to 0%, maximizing host earnings.

**What is server rating and how does it affect me?**

Server rating reflects reliability based on uptime and renter feedback. Higher-rated servers appear higher in marketplace listings and attract more renters. Frequent offline periods and reported issues will lower your rating.

**What happens if my server goes offline while rented?**

If your server goes offline during an active rental, this negatively impacts your server rating. On-demand renters may receive compensation. Maintaining stable, high-uptime servers is essential for maintaining a good rating and steady earnings.

**Are there penalties for poor server performance?**

Yes. Frequent downtime, unresolved renter issues, and low ratings can result in reduced marketplace visibility or other restrictions. Operating a reliable server is important for long-term earnings.

**What is Maximum Rental Length (MRL)?**

MRL is the maximum number of hours a single rental can last on your server. You configure this via `set_server_settings` (`mrl` field, in hours). Setting a longer MRL attracts renters with longer-running workloads.

**Can I set a background job for when my server isn't rented?**

Yes. You can configure a background job (Docker image + command) that runs on your server when it's not being rented, for example to mine cryptocurrency during idle time.

**How do I update my server pricing?**

Use the `set_server_settings` API endpoint or the web UI. Changes take effect for new rentals — active orders are not affected mid-rental.

***

## Proof of Holding

**What is Proof of Holding (PoH)?**

Proof of Holding is a system where users earn rewards for holding CLORE tokens in their Clore.ai wallets. It incentivizes long-term participation in the ecosystem and rewards holders proportionally based on their holdings.

**How do I participate in Proof of Holding?**

Simply hold CLORE tokens in your Clore.ai account wallet. The longer and more you hold, the more rewards you accumulate. No staking contract or locking is required for basic PoH participation.

**What rewards do I get from PoH?**

PoH participants receive a share of platform fee revenue distributed in CLORE. Rewards are calculated based on your holding relative to the total CLORE in PoH.

**How does PoH reduce marketplace fees?**

Holding CLORE proportionally reduces the on-demand fee (base 10%) and spot fee (base 2.5%) by up to 50%. Maximum fee reduction (50%) is achieved at 2,000,000 CLORE held. The reduction is linear and applies to both renters and hosts.

**Is PoH the same as MFP Lock?**

No. PoH reduces platform fees for everyone (renters and hosts) based on CLORE holdings. MFP Lock is specifically for hosts to eliminate the extra fee charged when renters pay with non-CLORE currencies. Both mechanisms use CLORE but serve different purposes.

**Where can I track my PoH rewards?**

PoH rewards and your current balance are visible in your Clore.ai account dashboard under the PoH or wallet section.

***

## API

**How do I get an API key?**

Log in to your Clore.ai account, navigate to the API settings section, and generate a new API token. Keep it secret — it provides full access to your account.

**How do I authenticate API requests?**

Pass your API token in the `auth` header of every request: `auth: <your_token>`. Do not use `Authorization: Bearer` format — Clore.ai uses the `auth` header specifically.

**What is the base URL for the API?**

All API endpoints are at `https://api.clore.ai/v1/`. Example: `https://api.clore.ai/v1/marketplace`.

**What are the API rate limits?**

The general rate limit is **1 request per second**. Exceeding this returns code `5`. The `create_order` endpoint has an additional stricter limit of **1 request per 5 seconds**.

**What do API response codes mean?**

| Code | Meaning                                   |
| ---- | ----------------------------------------- |
| `0`  | Success (Normal)                          |
| `1`  | Database error                            |
| `2`  | Invalid input data                        |
| `3`  | Invalid API token                         |
| `4`  | Invalid endpoint                          |
| `5`  | Rate limit exceeded (1 req/sec)           |
| `6`  | Error — see the `error` field for details |

**What are the main API endpoints?**

| Endpoint                  | Method | Description                            |
| ------------------------- | ------ | -------------------------------------- |
| `/v1/wallets`             | GET    | Get wallets and balances               |
| `/v1/marketplace`         | GET    | List all public servers                |
| `/v1/my_servers`          | GET    | List your hosted servers               |
| `/v1/my_orders`           | GET    | List your active/completed orders      |
| `/v1/create_order`        | POST   | Create a spot or on-demand order       |
| `/v1/cancel_order`        | POST   | Cancel an active order                 |
| `/v1/spot_marketplace`    | GET    | Get spot offers for a specific server  |
| `/v1/set_server_settings` | POST   | Configure your server settings         |
| `/v1/set_spot_price`      | POST   | Update your spot market offer price    |
| `/v1/server_config`       | GET    | Get configuration of a specific server |

**Does Clore.ai have a WebSocket API?**

The primary API uses REST (HTTP). For real-time data needs (e.g., monitoring marketplace changes), poll the relevant endpoints within rate limits. Check the official documentation for any WebSocket support updates.

**Can I automate order deployment with the API?**

Yes. The `create_order` endpoint lets you fully automate deployment: specify the server ID, Docker image, port forwarding, environment variables, SSH keys, and startup commands programmatically.

***

## Technical

**What GPU brands and models are supported?**

Clore.ai primarily supports **NVIDIA GPUs** (GeForce, RTX, Quadro, Tesla, A-series). AMD GPU support depends on host configuration and Docker image compatibility. Available models range from GTX 1080 Ti to high-end A100/H100-class cards, depending on what hosts list.

**What CUDA versions are available?**

CUDA version depends on the Docker image you deploy and the host's driver version. Official Clore.ai images (e.g., `cloreai/ubuntu20.04-jupyter`) include CUDA support. Use NVIDIA's official CUDA images or check the host's GPU driver version in the server specs before deploying.

**What are the Docker limitations on Clore.ai?**

Each order runs a single Docker container. Multi-container setups (Docker Compose) are not natively supported — you must include all services in one image or orchestrate them via a startup script. Privileged mode availability depends on the host's configuration.

**What network speeds can I expect?**

Network speeds vary by host and are listed in the server specs as upload/download in Mbps. Check the `net.up` and `net.down` values in marketplace listings before renting for bandwidth-sensitive workloads.

**What storage types are available?**

Storage depends on the host's hardware. Server specs list disk type (NVMe, SSD, HDD) and size. Disk speed (`disk_speed` in MB/s) is also shown. For fast I/O workloads, filter for NVMe servers.

**Is persistent storage between rentals supported?**

By default, container storage is ephemeral — data is lost when the order ends. Always back up important data externally (S3, IPFS, etc.) before your rental expires. Some hosts may offer optional persistent storage — check their server description.

**What are the minimum network requirements for hosts?**

Hosts should have a stable internet connection with adequate upload bandwidth for the number of GPUs they're serving. Low-latency, high-bandwidth connections (100 Mbps+ upload) provide better renter experience and higher server ratings.

**Can I run Windows containers on Clore.ai?**

Clore.ai is Linux-based. Docker containers run on Linux hosts. Windows containers are not supported. You can run Windows applications inside a Linux container using tools like Wine if needed.

**What happens if a server runs out of disk space during my rental?**

If the container fills available disk space, applications may crash or become unresponsive. Monitor disk usage during your session and clean up temporary files. Check the available disk size in server specs before deploying disk-heavy workloads.


# Glossary

<div data-full-width="true"><figure><img src="/files/qphLXp24kX0E9FSgAzNg" alt=""><figcaption></figcaption></figure></div>

**1. Blockchain**

A decentralized digital ledger that records transactions across many computers. In the Clore.ai ecosystem, blockchain technology ensures the transparency, security, and immutability of all transactions, fostering trust among users.

**2. Clore Coin (CLORE)**

The native cryptocurrency of the Clore.ai platform. CLORE is used to facilitate transactions, reward participants, and support the platform's overall economy.

**3. Decentralization**

A system where control and decision-making are distributed among multiple participants rather than centralized in a single entity. Clore.ai leverages decentralization to create a more equitable and resilient platform for GPU leasing and cryptocurrency transactions.

**4. GPU (Graphics Processing Unit)**

A specialized electronic circuit designed to accelerate the processing of images and videos. In Clore.ai, GPUs are leased through a decentralized marketplace, allowing users to access high-performance computing power for various tasks like AI development and cryptocurrency mining.

**5. GPU Leasing**

The process of renting out GPU resources to others who need computing power. Clore.ai provides a platform where GPU owners can lease their idle hardware to users needing temporary access to powerful GPUs.

**6. Proof of Holding (PoH)**

A system within Clore.ai that rewards users for holding CLORE coins in their wallets. Unlike traditional staking, PoH offers rewards immediately, providing incentives for long-term holding and contributing to the platform's stability.

**7. Marketplace**

In the context of Clore.ai, the marketplace is a decentralized platform where users can lease or rent GPUs and CLORE coins. It operates on a peer-to-peer model, ensuring transparent and direct transactions between users.

**8. Emission Schedule**

The planned release of CLORE token rewards into circulation over time. Clore.ai's emission schedule follows a decay model (-30% per year) for PoH and marketplace rewards, ensuring controlled distribution and long-term value.

<div data-full-width="true"><figure><img src="/files/ZdL5yRqFUzFWZXWjGyhP" alt=""><figcaption></figcaption></figure></div>

**9. Staking**

In Clore.ai, staking refers to holding CLORE tokens in the Proof of Holding (PoH) system to earn rewards and reduce marketplace fees. Unlike traditional blockchain staking, PoH does not require transaction validation - simply holding tokens qualifies you for benefits.

**10. Spot Rental**

An interruptible rental option on the Clore.ai platform where GPU resources are allocated at a lower cost. Other users can overbid your offer at any time, making this option best suited for fault-tolerant workloads like cryptocurrency mining or batch processing.

Base fee: 2.5%, reducible by up to 50% with Proof of Holding.

**11. On-Demand Rental**

A flexible rental option where users can lease GPUs from the Clore.ai marketplace on an as-needed basis. This option is ideal for users with varying computational needs.

Base fee: 10%, reducible by up to 50% with Proof of Holding.

**12. MFP Lock (Maximum Fair Price Lock)**

A system allowing GPU hosts to lock CLORE tokens against their servers across three tiers. Tier 1 and Tier 2 earn emission rewards, while all three tiers contribute to reducing extra hoster fees on non-CLORE currency orders. Maximum lock: MFP × 7,000 CLORE.

**13. Extra Fees**

Additional platform fees applied to non-CLORE currency payments (e.g., Bitcoin). Extra hoster fees can be reduced to 0% through MFP Lock tiers. CLORE payments have no extra fees.

**14. USD Stablecoins (USDT/USDC)**

Stablecoins pegged to the US Dollar, supported on Clore.ai for deposits, withdrawals, and rental payments on two networks: Ethereum (ERC-20) and BNB Chain (BEP-20). Both tokens credit to a unified USD balance on the platform.

**15. Airdrop**

A distribution of cryptocurrency tokens to a specific group of users, often used as a marketing strategy or as a reward for platform participation. Clore.ai periodically conducts airdrops to reward users, particularly those participating in the Proof of Holding system.

**16. Compliance**

Adherence to legal and regulatory requirements in the jurisdictions where Clore.ai operates. The platform is committed to maintaining full compliance with relevant laws, ensuring a secure and trustworthy environment for all users.

**17. Tokenomics**

The economic structure of a cryptocurrency, including its supply, distribution, and reward mechanisms. Clore.ai’s tokenomics are designed to ensure long-term sustainability and value growth for the CLORE coin.

**18. Scalability**

The ability of the Clore.ai platform to handle increasing demand without compromising performance. Scalability is a key focus for Clore.ai, ensuring that the platform can grow and serve a larger user base effectively.

**19. Bare Metal**

Dedicated physical GPU servers and clusters rented for fixed terms (from two weeks up to five years) instead of per-minute container rentals. Configurations of up to 1,000 GPUs across multiple regions, priced in USD per GPU per hour.

**20. Partial GPU Rental**

Renting only a subset of the GPUs on a multi-GPU server and paying a proportional price. Available on-demand on homogeneous rigs whose hosts run up-to-date hosting software; each order gets an isolated container with dedicated GPUs.


# Support & Contact

How to get help and contact the Clore.ai team.

## Getting Help

### Community Support

The fastest way to get help is through our community channels:

| Channel      | Link                                               | Best For                        |
| ------------ | -------------------------------------------------- | ------------------------------- |
| **Telegram** | [t.me/clorechat](https://t.me/clorechat)           | Quick questions, community help |
| **Discord**  | [discord.gg/clore-ai](https://discord.gg/clore-ai) | Technical discussions, support  |

### Self-Service Resources

Before reaching out, check these resources:

1. **Documentation** - You're here! Browse the guides
2. **FAQ** - [Frequently Asked Questions](/help/faq)
3. **Glossary** - [Terms and definitions](/help/glossary)

## Common Issues by Category

### For Renters

| Issue                   | Where to Look                                                 |
| ----------------------- | ------------------------------------------------------------- |
| Can't connect to server | [How to Connect](/for-renters/how-to-connect)                 |
| Server not working      | [Renter Troubleshooting](/for-renters/renter-troubleshooting) |
| Billing questions       | [Fee Structure](/for-renters/fee-structure)                   |
| Deposit/withdrawal      | [Deposit & Withdrawal](/getting-started/deposit-withdrawal)   |

### For Hosts

| Issue          | Where to Look                                                                           |
| -------------- | --------------------------------------------------------------------------------------- |
| Server offline | [Server Offline](/for-hosts/server-offline-on-clore.ai)                                 |
| Docker errors  | [Docker Failure](/for-hosts/server-offline-on-clore.ai/docker-failure)                  |
| Pricing help   | [How to Price Servers](/for-hosts/server-settings/how-to-price-your-servers-for-rental) |
| Network issues | [Network Requirements](/for-hosts/installing-clore-hosting/network-requirements)        |

## Contacting Support

For direct support, use the official support page or email:

| Channel            | Contact                                      |
| ------------------ | -------------------------------------------- |
| **Support Portal** | [clore.ai/support](https://clore.ai/support) |
| **Email**          | <support@clore.ai>                           |

### What to Include

When asking for help, provide:

1. **Your username** (not password!)
2. **Server ID** (if applicable)
3. **Order ID** (if about a rental)
4. **Screenshot** of the error
5. **Steps to reproduce** the issue

### Response Times

Community channels are monitored by both team members and experienced users. Response times vary but are typically:

* **Telegram/Discord**: Minutes to hours
* **Complex issues**: May take longer for investigation

## Official Links

| Resource        | URL                                                          |
| --------------- | ------------------------------------------------------------ |
| Website         | [clore.ai](https://clore.ai)                                 |
| Support Portal  | [clore.ai/support](https://clore.ai/support)                 |
| Marketplace     | [clore.ai/marketplace](https://clore.ai/marketplace)         |
| POH Marketplace | [clore.ai/poh-marketplace](https://clore.ai/poh-marketplace) |
| Roadmap         | [clore.ai/roadmap](https://clore.ai/roadmap)                 |

## Social Media

| Platform    | Link                                                            |
| ----------- | --------------------------------------------------------------- |
| Twitter/X   | [@clore\_ai](https://twitter.com/clore_ai)                      |
| Telegram    | [t.me/clorechat](https://t.me/clorechat)                        |
| Discord     | [discord.gg/clore-ai](https://discord.gg/clore-ai)              |
| BitcoinTalk | [Forum Thread](https://bitcointalk.org/index.php?topic=5428175) |

## Telegram Bot

Hosts can use the Clore Telegram bot for:

* Server status notifications
* Rental alerts
* Quick server management

See: [Telegram Bot Setup](/for-hosts/server-settings/telegram-bot)

## Reporting Issues

### Security Issues

For security vulnerabilities, contact the team directly through Discord private message to a team member. Do not post security issues publicly.

### Bug Reports

For bugs or issues with the platform:

1. Check if already reported in community channels
2. Provide detailed reproduction steps
3. Include screenshots/logs if possible

## Tips for Getting Help Faster

1. **Search first** - Your question may already be answered
2. **Be specific** - "Server not working" vs "Server shows CUDA error after 10 minutes"
3. **Include context** - OS version, GPU model, error messages
4. **Be patient** - Community members are volunteers
5. **Share solutions** - If you solve it yourself, share how!


# Python SDK (clore-ai)

The **clore-ai** package is the official Python SDK for the [Clore.ai](https://clore.ai) GPU marketplace. It wraps the entire REST API into a clean, type-safe interface with built-in rate limiting, automatic retries, and structured error handling — so you can focus on renting GPUs, not on HTTP plumbing.

***

## Installation

```bash
pip install clore-ai
```

**Requirements:** Python 3.9+

The package installs both the Python SDK and the [`clore` CLI](/developers/cli-guide).

***

## Authentication

Get your API key from the [Clore.ai dashboard](https://clore.ai) → **API** section.

### Option 1: Environment variable (recommended)

```bash
export CLORE_API_KEY=your_api_key_here
```

The SDK reads `CLORE_API_KEY` automatically — no code changes needed.

### Option 2: CLI config file

```bash
clore config set api_key YOUR_API_KEY
```

This stores the key in `~/.clore/config.json`.

### Option 3: Pass directly in code

```python
from clore_ai import CloreAI

client = CloreAI(api_key="your_api_key_here")
```

> ⚠️ **Important:** The Clore.ai API uses the `auth` header for authentication, **not** `Authorization: Bearer`. The SDK handles this automatically.

***

## Quick Start

```python
from clore_ai import CloreAI

client = CloreAI()
servers = client.marketplace(gpu="RTX 4090", max_price_usd=5.0)
for s in servers:
    print(f"Server {s.id}: {s.gpu_model} — ${s.price_usd:.4f}/h")
```

***

## Sync Client (`CloreAI`)

### Constructor

```python
CloreAI(
    api_key: str | None = None,       # Falls back to CLORE_API_KEY env / config
    base_url: str | None = None,       # Default: https://api.clore.ai/v1
    timeout: float = 30.0,             # Request timeout in seconds
    max_retries: int = 3               # Retry attempts on rate-limit / network errors
)
```

The client supports context managers for automatic cleanup:

```python
with CloreAI() as client:
    wallets = client.wallets()
    # client.close() called automatically
```

***

### `wallets()`

Get your wallet balances and deposit addresses.

```python
wallets = client.wallets()

for wallet in wallets:
    print(f"{wallet.name}: {wallet.balance:.8f}")
    if wallet.deposit:
        print(f"  Deposit: {wallet.deposit}")
```

**Returns:** `List[Wallet]`

| Field            | Type            | Description                                                                |
| ---------------- | --------------- | -------------------------------------------------------------------------- |
| `name`           | `str`           | Currency name (e.g. `"bitcoin"`, `"CLORE-Blockchain"`, `"USD-Blockchain"`) |
| `balance`        | `float \| None` | Current balance                                                            |
| `deposit`        | `str \| None`   | Deposit address                                                            |
| `withdrawal_fee` | `float \| None` | Withdrawal fee                                                             |

***

### `marketplace()`

Search the GPU marketplace with optional client-side filters.

```python
# All available servers
servers = client.marketplace()

# Filter by GPU model and max price
servers = client.marketplace(
    gpu="RTX 4090",
    max_price_usd=5.0
)

# Multi-GPU rigs with plenty of RAM
servers = client.marketplace(
    min_gpu_count=4,
    min_ram_gb=128.0
)
```

**Parameters:**

| Parameter        | Type            | Default | Description                                            |
| ---------------- | --------------- | ------- | ------------------------------------------------------ |
| `gpu`            | `str \| None`   | `None`  | Filter by GPU model (case-insensitive substring match) |
| `min_gpu_count`  | `int \| None`   | `None`  | Minimum number of GPUs                                 |
| `min_ram_gb`     | `float \| None` | `None`  | Minimum RAM in GB                                      |
| `max_price_usd`  | `float \| None` | `None`  | Maximum price per hour in USD                          |
| `available_only` | `bool`          | `True`  | Only return servers that are available to rent         |

**Returns:** `List[MarketplaceServer]`

Each `MarketplaceServer` provides convenient properties for the most common fields, plus access to the full nested data:

| Property         | Type            | Description                                                   |
| ---------------- | --------------- | ------------------------------------------------------------- |
| `id`             | `int`           | Unique server ID                                              |
| `gpu_model`      | `str \| None`   | Primary GPU description (e.g. `"1x NVIDIA GeForce RTX 4090"`) |
| `gpu_count`      | `int`           | Number of GPUs (from `gpu_array`)                             |
| `ram_gb`         | `float \| None` | RAM in GB                                                     |
| `price_usd`      | `float \| None` | On-demand price in USD                                        |
| `spot_price_usd` | `float \| None` | Spot price in USD                                             |
| `available`      | `bool`          | Whether the server is available (not rented)                  |
| `location`       | `str \| None`   | Country code from network specs                               |

For advanced use cases, you can access the full nested structure:

| Field         | Type                   | Description                                                                                  |
| ------------- | ---------------------- | -------------------------------------------------------------------------------------------- |
| `specs`       | `ServerSpecs \| None`  | Full hardware specs (`specs.gpu`, `specs.ram`, `specs.cpu`, `specs.disk`, `specs.net`, etc.) |
| `price`       | `ServerPrice \| None`  | Full price object (`price.usd.on_demand_usd`, `price.usd.spot`, `price.on_demand`, etc.)     |
| `rented`      | `bool \| None`         | Whether the server is currently rented                                                       |
| `reliability` | `float \| None`        | Server reliability score                                                                     |
| `rating`      | `ServerRating \| None` | Server rating (`rating.avg`, `rating.cnt`)                                                   |

> **Note:** The `marketplace()` endpoint is public — it works without an API key.

***

### `my_servers()`

List servers you are providing to the Clore.ai marketplace.

```python
my_servers = client.my_servers()

for server in my_servers:
    print(f"{server.name}: {server.gpu_model} [{server.status}]")
```

**Returns:** `List[MyServer]`

| Property     | Type            | Description                                                                          |
| ------------ | --------------- | ------------------------------------------------------------------------------------ |
| `id`         | `int`           | Server ID                                                                            |
| `name`       | `str \| None`   | Server name                                                                          |
| `gpu_model`  | `str \| None`   | Primary GPU description                                                              |
| `ram_gb`     | `float \| None` | RAM in GB                                                                            |
| `status`     | `str`           | Human-readable status: `"Online"`, `"Offline"`, `"Disconnected"`, or `"Not Working"` |
| `connected`  | `bool \| None`  | Whether the server is connected                                                      |
| `online`     | `bool \| None`  | Whether the server is online                                                         |
| `visibility` | `str \| None`   | `"public"` or `"private"`                                                            |

***

### `server_config(server_name)`

Get the configuration of a specific server you host.

```python
config = client.server_config("MyGPU")

print(f"Server: {config.name}")
print(f"GPU: {config.gpu_model}")
print(f"Min rental: {config.mrl}h")
print(f"On-demand: ${config.on_demand_price}")
print(f"Spot: ${config.spot_price}")
```

**Parameters:**

| Parameter     | Type  | Description        |
| ------------- | ----- | ------------------ |
| `server_name` | `str` | Name of the server |

**Returns:** `ServerConfig`

| Property          | Type                  | Description                         |
| ----------------- | --------------------- | ----------------------------------- |
| `name`            | `str \| None`         | Server name                         |
| `gpu_model`       | `str \| None`         | Primary GPU description             |
| `mrl`             | `int \| None`         | Maximum rental length in hours      |
| `on_demand_price` | `float \| None`       | First available on-demand USD price |
| `spot_price`      | `float \| None`       | First available spot USD price      |
| `specs`           | `ServerSpecs \| None` | Full hardware specifications        |
| `connected`       | `bool \| None`        | Whether the server is connected     |
| `visibility`      | `str \| None`         | `"public"` or `"private"`           |

***

### `my_orders(include_completed)`

Get your current orders, optionally including completed/expired ones.

```python
# Active orders only
orders = client.my_orders()

# Include completed orders
all_orders = client.my_orders(include_completed=True)

for order in orders:
    print(f"Order {order.id}: {order.type} — {order.status}")
    if order.pub_cluster:
        print(f"  IP: {order.pub_cluster}")
    if order.tcp_ports:
        print(f"  Ports: {order.tcp_ports}")
```

**Parameters:**

| Parameter           | Type   | Default | Description                      |
| ------------------- | ------ | ------- | -------------------------------- |
| `include_completed` | `bool` | `False` | Include completed/expired orders |

**Returns:** `List[Order]`

| Field         | Type            | Description                   |
| ------------- | --------------- | ----------------------------- |
| `id`          | `int`           | Unique order ID               |
| `server_id`   | `int \| None`   | Server ID                     |
| `type`        | `str`           | `"on-demand"` or `"spot"`     |
| `status`      | `str \| None`   | Order status                  |
| `image`       | `str \| None`   | Docker image                  |
| `currency`    | `str \| None`   | Payment currency              |
| `price`       | `float \| None` | Order price per day           |
| `pub_cluster` | `str \| None`   | Public hostname/IP for access |
| `tcp_ports`   | `dict \| None`  | TCP port mappings             |

***

### `spot_marketplace(server_id)`

View spot market offers for a specific server.

```python
spot = client.spot_marketplace(server_id=6)

if spot.offers:
    for offer in spot.offers:
        print(f"Order {offer.order_id}: ${offer.price}/day (server {offer.server_id})")

if spot.currency_rates_in_usd:
    for coin, rate in spot.currency_rates_in_usd.items():
        print(f"  {coin}: ${rate}")
```

**Parameters:**

| Parameter   | Type  | Description        |
| ----------- | ----- | ------------------ |
| `server_id` | `int` | Server ID to check |

**Returns:** `SpotMarket`

| Field                   | Type                       | Description                                            |
| ----------------------- | -------------------------- | ------------------------------------------------------ |
| `offers`                | `List[SpotOffer] \| None`  | List of spot offers (`order_id`, `price`, `server_id`) |
| `server`                | `SpotServerInfo \| None`   | Server info (min pricing, visibility, online status)   |
| `currency_rates_in_usd` | `Dict[str, float] \| None` | Currency exchange rates in USD                         |

***

### `create_order(...)`

Create a new on-demand or spot order. This is how you rent a GPU.

#### On-demand order

```python
order = client.create_order(
    server_id=123,
    image="cloreai/ubuntu22.04-cuda12",
    type="on-demand",
    currency="bitcoin",
    ssh_password="MySecurePass123",
    ports={"22": "tcp", "8888": "http"}
)

print(f"Order created: {order.id}")
print(f"Connect: {order.pub_cluster}")
```

#### Spot order

```python
order = client.create_order(
    server_id=123,
    image="cloreai/pytorch",
    type="spot",
    currency="bitcoin",
    spot_price=0.000005,
    ssh_password="MySecurePass123",
    ports={"22": "tcp"}
)
```

**Parameters:**

| Parameter            | Type        | Required  | Description                                                                                                                     |
| -------------------- | ----------- | --------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `server_id`          | `int`       | Yes       | Server ID to rent                                                                                                               |
| `image`              | `str`       | Yes       | Docker image (e.g. `"cloreai/ubuntu22.04-cuda12"`)                                                                              |
| `type`               | `str`       | Yes       | `"on-demand"` or `"spot"`                                                                                                       |
| `currency`           | `str`       | Yes       | Payment currency (e.g. `"bitcoin"`)                                                                                             |
| `ssh_password`       | `str`       | No        | SSH password (alphanumeric, max 32 chars)                                                                                       |
| `ssh_key`            | `str`       | No        | SSH public key (max 3072 chars)                                                                                                 |
| `ports`              | `dict`      | No        | Port mappings, e.g. `{"22": "tcp", "8888": "http"}`                                                                             |
| `env`                | `dict`      | No        | Environment variables                                                                                                           |
| `jupyter_token`      | `str`       | No        | Jupyter notebook token (max 32 chars)                                                                                           |
| `command`            | `str`       | No        | Shell command to run after container start                                                                                      |
| `spot_price`         | `float`     | Spot only | Price per day for spot orders                                                                                                   |
| `required_price`     | `float`     | No        | Lock in a specific price (on-demand only)                                                                                       |
| `autossh_entrypoint` | `str`       | No        | Use Clore.ai SSH entrypoint                                                                                                     |
| `gpu_count`          | `int`       | No        | Rent only N GPUs on servers with [partial rental](/for-renters/partial-gpu-rental) (on-demand only); omit to rent the whole rig |
| `gpu_indices`        | `list[int]` | No        | Exact GPU slots from `partial_gpu_rental.free_indices`; length must equal `gpu_count`; omit for auto-pick                       |

**Returns:** raw API response (`{"code": 0}` on success); fetch the created order via `my_orders()`

> **Rate limit:** `create_order` has a special 5-second cooldown between calls. The SDK enforces this automatically.

***

### `cancel_order(order_id, issue)`

Cancel an active order or spot offer. Optionally report an issue with the server.

```python
# Simple cancel
client.cancel_order(order_id=38)

# Cancel with issue report
client.cancel_order(
    order_id=38,
    issue="GPU #1 was overheating and throttling"
)
```

**Parameters:**

| Parameter  | Type  | Required | Description                                         |
| ---------- | ----- | -------- | --------------------------------------------------- |
| `order_id` | `int` | Yes      | Order ID to cancel                                  |
| `issue`    | `str` | No       | Cancellation reason / issue report (max 2048 chars) |

**Returns:** `Dict[str, Any]`

***

### `set_server_settings(...)`

Update settings for a server you host on the marketplace.

```python
client.set_server_settings(
    name="MyGPU",
    availability=True,
    mrl=96,
    on_demand=0.0001,
    spot=0.00000113
)
```

**Parameters:**

| Parameter      | Type    | Required | Description                      |
| -------------- | ------- | -------- | -------------------------------- |
| `name`         | `str`   | Yes      | Server name                      |
| `availability` | `bool`  | No       | Whether the server can be rented |
| `mrl`          | `int`   | No       | Maximum rental length in hours   |
| `on_demand`    | `float` | No       | On-demand price per day          |
| `spot`         | `float` | No       | Minimum spot price per day       |

**Returns:** `Dict[str, Any]`

***

### `set_spot_price(order_id, price)`

Update the price on your spot market offer.

```python
client.set_spot_price(order_id=39, price=0.000003)
```

**Parameters:**

| Parameter  | Type    | Description         |
| ---------- | ------- | ------------------- |
| `order_id` | `int`   | Spot order/offer ID |
| `price`    | `float` | New price per day   |

**Returns:** `Dict[str, Any]`

> **Note:** You can only lower spot prices once every 600 seconds, and by a limited step size. The API returns `code: 6` with details if you exceed these limits.

***

## Async Client (`AsyncCloreAI`)

The `AsyncCloreAI` client provides the same methods as `CloreAI`, but all return coroutines. Use it when you need concurrent API calls or are working within an async application.

### Basic usage

```python
import asyncio
from clore_ai import AsyncCloreAI

async def main():
    async with AsyncCloreAI(api_key="your_key") as client:
        wallets = await client.wallets()
        for w in wallets:
            print(f"{w.name}: {w.balance:.8f}")

asyncio.run(main())
```

### Concurrent operations

Run multiple API calls in parallel with `asyncio.gather`:

```python
import asyncio
from clore_ai import AsyncCloreAI

async def compare_gpus():
    async with AsyncCloreAI() as client:
        # Search for multiple GPU models concurrently
        rtx4090, rtx3090, a100 = await asyncio.gather(
            client.marketplace(gpu="RTX 4090"),
            client.marketplace(gpu="RTX 3090"),
            client.marketplace(gpu="A100"),
        )

        for name, servers in [("RTX 4090", rtx4090), ("RTX 3090", rtx3090), ("A100", a100)]:
            if servers:
                cheapest = min(s.price_usd or float('inf') for s in servers)
                print(f"{name}: {len(servers)} available, cheapest ${cheapest:.4f}/h")
            else:
                print(f"{name}: none available")

asyncio.run(compare_gpus())
```

### Available methods

`AsyncCloreAI` supports all the same methods as `CloreAI`:

| Method                              | Description              |
| ----------------------------------- | ------------------------ |
| `await wallets()`                   | Get wallet balances      |
| `await marketplace(...)`            | Search marketplace       |
| `await my_servers()`                | List your hosted servers |
| `await server_config(name)`         | Get server configuration |
| `await my_orders(...)`              | List your orders         |
| `await spot_marketplace(server_id)` | Get spot market offers   |
| `await create_order(...)`           | Create a new order       |
| `await cancel_order(...)`           | Cancel an order          |
| `await set_server_settings(...)`    | Update server settings   |
| `await set_spot_price(...)`         | Update spot price        |

***

## Error Handling

The SDK provides structured exception classes for every API error code.

```python
from clore_ai import CloreAI
from clore_ai.exceptions import (
    CloreAPIError,      # Base class for all API errors
    AuthError,          # Code 3 — invalid API key
    RateLimitError,     # Code 5 — rate limit exceeded
    InvalidInputError,  # Code 2 — bad request data
    DBError,            # Code 1 — database error
    InvalidEndpointError,  # Code 4 — invalid endpoint
    FieldError,         # Code 6 — field-specific error
)

client = CloreAI()

try:
    order = client.create_order(
        server_id=123,
        image="cloreai/ubuntu22.04-cuda12",
        type="on-demand",
        currency="bitcoin",
    )
except AuthError:
    print("Invalid API key. Check your CLORE_API_KEY.")
except RateLimitError:
    print("Rate limited. The SDK retries automatically, but you hit max retries.")
except InvalidInputError as e:
    print(f"Bad request: {e}")
except FieldError as e:
    # Code 6 errors include details in the response
    print(f"Field error: {e} (details: {e.response})")
except CloreAPIError as e:
    print(f"API error: {e} (code: {e.code})")
```

### Error codes

| Code | Exception              | Description                                             |
| ---- | ---------------------- | ------------------------------------------------------- |
| 0    | —                      | Success                                                 |
| 1    | `DBError`              | Database error                                          |
| 2    | `InvalidInputError`    | Invalid input data                                      |
| 3    | `AuthError`            | Invalid API token                                       |
| 4    | `InvalidEndpointError` | Invalid endpoint                                        |
| 5    | `RateLimitError`       | Rate limit exceeded                                     |
| 6    | `FieldError`           | Error in specific field (see `error` field in response) |

All exception classes inherit from `CloreAPIError` and include:

* `e.code` — numeric error code
* `e.response` — full API response dict (when available)

***

## Rate Limiting

The SDK includes a built-in rate limiter that enforces Clore.ai's limits automatically:

| Endpoint       | Limit                   |
| -------------- | ----------------------- |
| Most endpoints | **1 request/second**    |
| `create_order` | **1 request/5 seconds** |

When the API returns a rate-limit error (code 5), the SDK applies **exponential backoff** and retries up to `max_retries` times (default: 3). You don't need to add `time.sleep()` between calls.

### How it works

1. Before each request, the rate limiter waits until the minimum interval has elapsed.
2. `create_order` calls enforce an additional 5-second cooldown.
3. On rate-limit errors, the SDK backs off exponentially: 1s → 2s → 4s → ...
4. After `max_retries` failed attempts, a `RateLimitError` is raised.

### Customize retry behavior

```python
client = CloreAI(
    max_retries=5,    # More retries for long-running scripts
    timeout=60.0      # Longer timeout for slow connections
)
```

***

## Configuration

### Config file

The CLI stores configuration in `~/.clore/config.json`:

```json
{
  "api_key": "your_api_key_here"
}
```

### Resolution order

The SDK resolves the API key in this order:

1. `api_key` argument passed to the constructor
2. `CLORE_API_KEY` environment variable
3. `api_key` field in `~/.clore/config.json`

### Environment variables

| Variable        | Description                |
| --------------- | -------------------------- |
| `CLORE_API_KEY` | API key for authentication |

***

## Next Steps

* [**CLI Reference**](/developers/cli-guide) — Use Clore.ai from your terminal
* [**REST API**](/for-hosts/api) — Raw API documentation for custom integrations
* [**On-Demand vs Spot**](/for-renters/on-demand-vs-spot) — Understand pricing models
* [**Available Docker Images**](/for-renters/docker-images) — Pre-built images for GPU workloads


# CLI Reference

The `clore` CLI lets you manage the Clore.ai GPU marketplace directly from your terminal — search for GPUs, deploy instances, SSH in, and manage orders without writing any code.

***

## Installation

```bash
pip install clore-ai
```

This installs both the Python SDK and the `clore` CLI command.

**Requirements:** Python 3.9+

***

## Configuration

### Set your API key

Get your key from the [Clore.ai dashboard](https://clore.ai) → **API** section, then configure:

```bash
clore config set api_key YOUR_API_KEY
```

This saves the key to `~/.clore/config.json`.

### Or use an environment variable

```bash
export CLORE_API_KEY=your_api_key_here
```

### View current config

```bash
# Show all config
clore config show

# Get a specific value
clore config get api_key
```

***

## Commands

| Command                               | Description                        |
| ------------------------------------- | ---------------------------------- |
| `clore search`                        | Search the GPU marketplace         |
| `clore deploy <server_id>`            | Create a new order (rent a server) |
| `clore orders`                        | List your orders                   |
| `clore cancel <order_id>`             | Cancel an order                    |
| `clore ssh <order_id>`                | SSH into an active order           |
| `clore wallets`                       | Show wallet balances               |
| `clore servers`                       | List your hosted servers           |
| `clore server-config <name>`          | Get server configuration           |
| `clore spot <server_id>`              | View spot market for a server      |
| `clore spot-price <order_id> <price>` | Set spot price                     |
| `clore config set <key> <value>`      | Set a config value                 |
| `clore config get <key>`              | Get a config value                 |
| `clore config show`                   | Show all configuration             |
| `clore --version`                     | Show version                       |

***

## Detailed Usage

### `clore search`

Search the GPU marketplace with filters and sorting.

```bash
# List all available servers (sorted by price, top 20)
clore search

# Filter by GPU model
clore search --gpu "RTX 4090"

# Filter by GPU and max price
clore search --gpu "RTX 4090" --max-price 5.0

# Multi-GPU rigs
clore search --min-gpu 4

# Sort by GPU count, show top 10
clore search --sort gpu --limit 10

# Combine filters
clore search --gpu "A100" --min-ram 128 --max-price 10.0 --sort price --limit 5
```

**Options:**

| Option        | Type   | Description                              |
| ------------- | ------ | ---------------------------------------- |
| `--gpu`       | text   | Filter by GPU model (e.g. `"RTX 4090"`)  |
| `--min-gpu`   | int    | Minimum GPU count                        |
| `--min-ram`   | float  | Minimum RAM in GB                        |
| `--max-price` | float  | Maximum price in USD/hour                |
| `--sort`      | choice | Sort by: `price` (default), `gpu`, `ram` |
| `--limit`     | int    | Max results to show (default: 20)        |

**Example output:**

```
🔍 GPU Marketplace
┏━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ ID  ┃ GPU                             ┃ Count ┃ RAM (GB) ┃ CPU                          ┃ Price/h (USD) ┃ Location ┃
┡━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ 123 │ 1x NVIDIA GeForce RTX 4090      │     1 │    64.0  │ AMD Ryzen 9 5900X            │       $0.3200 │ US       │
│ 456 │ 2x NVIDIA GeForce RTX 4090      │     2 │   128.0  │ Intel Core i9-13900K         │       $0.5800 │ DE       │
└─────┴─────────────────────────────────┴───────┴──────────┴──────────────────────────────┴───────────────┴──────────┘
Showing 2 servers
```

***

### `clore deploy <server_id>`

Create a new order to rent a GPU server.

```bash
# On-demand order with SSH
clore deploy 123 \
  --image cloreai/ubuntu22.04-cuda12 \
  --type on-demand \
  --currency bitcoin \
  --ssh-password MySecurePass123 \
  --port 22:tcp \
  --port 8888:http

# Spot order
clore deploy 123 \
  --image cloreai/pytorch \
  --type spot \
  --currency bitcoin \
  --spot-price 0.000005 \
  --port 22:tcp

# With environment variables
clore deploy 123 \
  --image cloreai/ubuntu22.04-cuda12 \
  --type on-demand \
  --currency bitcoin \
  --ssh-password MyPass123 \
  --port 22:tcp \
  --env WANDB_API_KEY=your_key \
  --env HF_TOKEN=your_token
```

**Arguments:**

| Argument    | Description                                 |
| ----------- | ------------------------------------------- |
| `server_id` | Server ID to rent (get from `clore search`) |

**Options:**

| Option           | Type   | Required | Description                                                                               |
| ---------------- | ------ | -------- | ----------------------------------------------------------------------------------------- |
| `--image`        | text   | Yes      | Docker image (e.g. `cloreai/ubuntu22.04-cuda12`)                                          |
| `--type`         | choice | Yes      | `on-demand` or `spot`                                                                     |
| `--currency`     | text   | Yes      | Payment currency (e.g. `bitcoin`)                                                         |
| `--ssh-password` | text   | No       | SSH password (alphanumeric, max 32 chars)                                                 |
| `--ssh-key`      | text   | No       | SSH public key                                                                            |
| `--port`         | text   | No       | Port mapping (repeatable), format: `PORT:PROTOCOL`                                        |
| `--env`          | text   | No       | Environment variable (repeatable), format: `KEY=VALUE`                                    |
| `--spot-price`   | float  | No       | Spot price per day (required for spot orders)                                             |
| `--gpu-count`    | int    | No       | Rent only N GPUs (partial rental, on-demand only)                                         |
| `--gpu-indices`  | text   | No       | Comma-separated GPU slots from `partial_gpu_rental.free_indices` (requires `--gpu-count`) |

***

### `clore orders`

List your active orders.

```bash
# Active orders
clore orders

# Include completed/expired orders
clore orders --completed
```

**Options:**

| Option        | Description                      |
| ------------- | -------------------------------- |
| `--completed` | Include completed/expired orders |

**Example output:**

```
📦 My Orders
┏━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ ID ┃ Server ID ┃ Type      ┃ Status ┃ Image                             ┃ IP                      ┃
┡━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ 38 │ 6         │ on-demand │ Active │ cloreai/ubuntu22.04-cuda12        │ n1.c1.clorecloud.net    │
└────┴───────────┴───────────┴────────┴───────────────────────────────────┴─────────────────────────┘
Total: 1 orders
```

***

### `clore cancel <order_id>`

Cancel an active order.

```bash
# Cancel order
clore cancel 38

# Cancel with issue report
clore cancel 38 --issue "GPU was overheating"
```

**Arguments:**

| Argument   | Description        |
| ---------- | ------------------ |
| `order_id` | Order ID to cancel |

**Options:**

| Option    | Description                        |
| --------- | ---------------------------------- |
| `--issue` | Cancellation reason / issue report |

***

### `clore ssh <order_id>`

Auto-connect via SSH to a running order. The CLI resolves the hostname and port from your order details.

```bash
# Connect as root (default)
clore ssh 38

# Connect as a different user
clore ssh 38 --user ubuntu
```

**Arguments:**

| Argument   | Description            |
| ---------- | ---------------------- |
| `order_id` | Order ID to connect to |

**Options:**

| Option   | Default | Description  |
| -------- | ------- | ------------ |
| `--user` | `root`  | SSH username |

***

### `clore wallets`

Show your wallet balances and deposit addresses.

```bash
clore wallets
```

**Example output:**

```
💰 Wallet Balances
┏━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Currency         ┃ Balance      ┃ Deposit Address                           ┃
┡━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ bitcoin          │ 0.00153176   │ tb1q6erw7v02t7hakgmlcl4wfnlykzqj05alndruwr │
│ CLORE-Blockchain │ 150.00000000 │ cLr1q8x...                                │
└──────────────────┴──────────────┴───────────────────────────────────────────┘
```

***

### `clore servers`

List servers you are hosting on the marketplace.

```bash
clore servers
```

***

### `clore server-config <name>`

Get the configuration of a specific server you host.

```bash
clore server-config "MyGPU"
```

**Example output:**

```
Server: MyGPU
ID: 42
Connected: True
Online: True
Visibility: public
Max Rental Length: 72
GPUs: NVIDIA GeForce RTX 4090
On-Demand Price (USD): $0.35
Spot Price (USD): $0.18
CPU: AMD Ryzen 9 5900X 12-Core Processor
RAM: 64.0 GB
GPU: 1x NVIDIA GeForce RTX 4090
```

***

### `clore spot <server_id>`

View current spot market offers for a server.

```bash
clore spot 6
```

**Example output:**

```
📊 Spot Market - Server 6
┏━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Order ID ┃ Price     ┃
┡━━━━━━━━━━╇━━━━━━━━━━━┩
│ 39       │ 4.2e-06   │
└──────────┴───────────┘
```

***

### `clore spot-price <order_id> <price>`

Update the price on your spot market offer.

```bash
clore spot-price 39 0.000003
```

> **Note:** You can only lower spot prices once every 600 seconds, and by a limited step size.

***

### `clore config`

Manage CLI configuration.

```bash
# Set API key
clore config set api_key YOUR_API_KEY

# Get a value
clore config get api_key

# Show all config
clore config show
```

Configuration is stored in `~/.clore/config.json`.

***

## Workflows

### Search → Rent → SSH → Cancel

A typical end-to-end workflow:

```bash
# 1. Find an RTX 4090 under $5/hour
clore search --gpu "RTX 4090" --max-price 5.0 --sort price --limit 5

# 2. Deploy on the cheapest server (e.g. ID 123)
clore deploy 123 \
  --image cloreai/ubuntu22.04-cuda12 \
  --type on-demand \
  --currency bitcoin \
  --ssh-password MyPass123 \
  --port 22:tcp \
  --port 8888:http

# 3. Check your order status
clore orders

# 4. SSH into the instance (e.g. order ID 456)
clore ssh 456

# 5. When done, cancel the order
clore cancel 456
```

### Monitor balances before renting

```bash
# Check you have enough funds
clore wallets

# Search and rent
clore search --gpu "A100" --sort price --limit 3
clore deploy 789 \
  --image cloreai/ubuntu22.04-cuda12 \
  --type on-demand \
  --currency bitcoin \
  --ssh-password SecurePass \
  --port 22:tcp
```

### Spot market workflow

```bash
# 1. Check spot offers for a server
clore spot 6

# 2. Create a spot order
clore deploy 6 \
  --image cloreai/pytorch \
  --type spot \
  --currency bitcoin \
  --spot-price 0.000005 \
  --ssh-password MyPass123 \
  --port 22:tcp

# 3. Adjust your spot price
clore spot-price 39 0.000003
```

### Server hosting management

```bash
# View your servers
clore servers

# Check a server's config
clore server-config "MyGPU"
```

***

## Next Steps

* [**Python SDK**](/developers/python-sdk) — Automate workflows with Python
* [**REST API**](/for-hosts/api) — Raw API documentation
* [**On-Demand vs Spot**](/for-renters/on-demand-vs-spot) — Understand pricing models


# CLI & SDK Guide

## Overview

Clore.ai provides a **REST API** that enables full programmatic access to the GPU marketplace — listing servers, creating orders, monitoring deployments, and canceling rentals.

> **Note:** There is no official CLI binary at this time. All automation is done directly via the REST API using tools like `curl`, Python, or Node.js.

**Base URL:** `https://api.clore.ai/v1`

**Response format:** JSON. Every response includes a `code` field indicating the status.

***

## Authentication

Generate your API key from the [Clore.ai Dashboard](https://clore.ai):

1. Log in to your account
2. Navigate to **API** section in settings
3. Generate and copy your API key

**Header format:**

```
auth: YOUR_API_KEY
```

> ⚠️ **Important:** The auth header is `auth`, **not** `Authorization: Bearer`. Using the wrong format will return code `3` (Invalid API Token).

**Example:**

```bash
curl -H 'auth: YOUR_API_KEY' 'https://api.clore.ai/v1/marketplace'
```

***

## Quick Start

### List Marketplace Servers

Browse all available GPU servers:

```bash
curl -XGET \
  -H 'auth: YOUR_API_KEY' \
  'https://api.clore.ai/v1/marketplace'
```

**Response:**

```json
{
  "servers": [
    {
      "id": 6,
      "owner": 4,
      "mrl": 73,
      "price": {
        "on_demand": { "bitcoin": 0.00001 },
        "spot": { "bitcoin": 0.000001 }
      },
      "rented": false,
      "specs": {
        "cpu": "Intel Core i9-11900",
        "ram": 62.67,
        "gpu": "1x NVIDIA GeForce GTX 1080 Ti",
        "gpuram": 11,
        "net": { "up": 26.38, "down": 118.42, "cc": "CZ" }
      }
    }
  ],
  "my_servers": [1, 2, 4],
  "code": 0
}
```

***

### Get Your Orders

```bash
curl -XGET \
  -H 'auth: YOUR_API_KEY' \
  'https://api.clore.ai/v1/my_orders'
```

Include completed/expired orders:

```bash
curl -XGET \
  -H 'auth: YOUR_API_KEY' \
  'https://api.clore.ai/v1/my_orders?return_completed=true'
```

***

### Create an Order (On-Demand)

```bash
curl -XPOST \
  -H 'auth: YOUR_API_KEY' \
  -H 'Content-type: application/json' \
  -d '{
    "currency": "bitcoin",
    "image": "cloreai/ubuntu20.04-jupyter",
    "renting_server": 6,
    "type": "on-demand",
    "ports": {
      "22": "tcp",
      "8888": "http"
    },
    "ssh_password": "YourSSHPassword123",
    "jupyter_token": "YourJupyterToken123"
  }' \
  'https://api.clore.ai/v1/create_order'
```

**Response:**

```json
{ "code": 0 }
```

***

### Create a Spot Order

Spot orders are cheaper but can be outbid. You set your price per day:

```bash
curl -XPOST \
  -H 'auth: YOUR_API_KEY' \
  -H 'Content-type: application/json' \
  -d '{
    "currency": "bitcoin",
    "image": "cloreai/ubuntu20.04-jupyter",
    "renting_server": 6,
    "type": "spot",
    "spotprice": 0.000005,
    "ports": {
      "22": "tcp",
      "8888": "http"
    },
    "ssh_password": "YourSSHPassword123"
  }' \
  'https://api.clore.ai/v1/create_order'
```

***

### Check Order Status

```bash
curl -XGET \
  -H 'auth: YOUR_API_KEY' \
  'https://api.clore.ai/v1/my_orders'
```

Active orders include `pub_cluster` (hostnames) and `tcp_ports` for SSH access:

```json
{
  "id": 38,
  "pub_cluster": ["n1.c1.clorecloud.net", "n2.c1.clorecloud.net"],
  "tcp_ports": ["22:10000"],
  "http_port": "8888",
  "expired": false
}
```

SSH into your rented server:

```bash
ssh root@n1.c1.clorecloud.net -p 10000
```

***

### Cancel an Order

```bash
curl -XPOST \
  -H 'auth: YOUR_API_KEY' \
  -H 'Content-type: application/json' \
  -d '{
    "id": 38
  }' \
  'https://api.clore.ai/v1/cancel_order'
```

Optionally report an issue with the server:

```bash
curl -XPOST \
  -H 'auth: YOUR_API_KEY' \
  -H 'Content-type: application/json' \
  -d '{
    "id": 38,
    "issue": "GPU was overheating and throttling performance"
  }' \
  'https://api.clore.ai/v1/cancel_order'
```

***

## Python SDK

A lightweight wrapper using the `requests` library. Install it with:

```bash
pip install requests
```

### CloreClient Class

```python
import requests
import time

class CloreClient:
    """Simple Python SDK for Clore.ai REST API."""
    
    BASE_URL = "https://api.clore.ai/v1"
    
    def __init__(self, api_key: str):
        self.api_key = api_key
        self.session = requests.Session()
        self.session.headers.update({"auth": api_key})
    
    def _get(self, endpoint: str, params: dict = None) -> dict:
        url = f"{self.BASE_URL}/{endpoint}"
        response = self.session.get(url, params=params)
        response.raise_for_status()
        data = response.json()
        if data.get("code") != 0:
            raise CloreAPIError(data)
        return data
    
    def _post(self, endpoint: str, body: dict) -> dict:
        url = f"{self.BASE_URL}/{endpoint}"
        self.session.headers.update({"Content-type": "application/json"})
        response = self.session.post(url, json=body)
        response.raise_for_status()
        data = response.json()
        if data.get("code") != 0:
            raise CloreAPIError(data)
        return data
    
    def list_servers(self) -> list:
        """Get all available servers on the marketplace."""
        data = self._get("marketplace")
        return data["servers"]
    
    def get_server_details(self, server_id: int) -> dict:
        """Get spot marketplace info for a specific server."""
        data = self._get("spot_marketplace", params={"market": server_id})
        return data["market"]
    
    def create_order(
        self,
        server_id: int,
        image: str,
        order_type: str = "on-demand",
        currency: str = "bitcoin",
        spotprice: float = None,
        ports: dict = None,
        ssh_password: str = None,
        ssh_key: str = None,
        jupyter_token: str = None,
        env: dict = None,
        command: str = None,
    ) -> dict:
        """
        Create an on-demand or spot order.
        
        Args:
            server_id: ID of server to rent
            image: Docker image (e.g. 'cloreai/ubuntu20.04-jupyter')
            order_type: 'on-demand' or 'spot'
            currency: 'bitcoin' (default)
            spotprice: Required for spot orders — price per day in BTC
            ports: Port forwarding, e.g. {"22": "tcp", "8888": "http"}
            ssh_password: SSH password (alphanumeric chars only)
            ssh_key: SSH public key
            jupyter_token: Jupyter notebook token
            env: Environment variables dict
            command: Shell command to run after container start
        """
        if order_type == "spot" and spotprice is None:
            raise ValueError("spotprice is required for spot orders")
        
        body = {
            "currency": currency,
            "image": image,
            "renting_server": server_id,
            "type": order_type,
        }
        
        if spotprice is not None:
            body["spotprice"] = spotprice
        if ports:
            body["ports"] = ports
        if ssh_password:
            body["ssh_password"] = ssh_password
        if ssh_key:
            body["ssh_key"] = ssh_key
        if jupyter_token:
            body["jupyter_token"] = jupyter_token
        if env:
            body["env"] = env
        if command:
            body["command"] = command
        
        return self._post("create_order", body)
    
    def get_orders(self, include_completed: bool = False) -> list:
        """Get your active (and optionally completed) orders."""
        params = {"return_completed": "true"} if include_completed else {}
        data = self._get("my_orders", params=params)
        return data["orders"]
    
    def cancel_order(self, order_id: int, issue: str = None) -> dict:
        """Cancel an order. Optionally report an issue."""
        body = {"id": order_id}
        if issue:
            body["issue"] = issue
        return self._post("cancel_order", body)
    
    def get_wallets(self) -> list:
        """Get your wallets and balances."""
        data = self._get("wallets")
        return data["wallets"]
    
    def get_my_servers(self) -> list:
        """Get servers you are providing to the marketplace."""
        data = self._get("my_servers")
        return data["servers"]


class CloreAPIError(Exception):
    """Raised when the Clore API returns a non-zero code."""
    
    ERROR_CODES = {
        0: "Normal",
        1: "Database error",
        2: "Invalid input data",
        3: "Invalid API token",
        4: "Invalid endpoint",
        5: "Rate limit exceeded (1 req/sec)",
        6: "Error (see error field)",
    }
    
    def __init__(self, response: dict):
        self.code = response.get("code")
        self.error = response.get("error", "")
        message = self.ERROR_CODES.get(self.code, f"Unknown code {self.code}")
        if self.error:
            message = f"{message}: {self.error}"
        super().__init__(f"Clore API error {self.code}: {message}")
```

### Full Working Example

```python
import os
from clore_client import CloreClient, CloreAPIError

# Initialize client
client = CloreClient(api_key=os.environ["CLORE_API_KEY"])

# 1. Browse marketplace
servers = client.list_servers()
print(f"Found {len(servers)} servers on marketplace")

# 2. Filter available RTX 4090 servers
rtx4090_servers = [
    s for s in servers
    if "4090" in s["specs"].get("gpu", "") and not s["rented"]
]

if not rtx4090_servers:
    print("No available RTX 4090 servers found")
    exit(1)

# 3. Pick the cheapest one
cheapest = min(rtx4090_servers, key=lambda s: s["price"]["on_demand"]["bitcoin"])
print(f"Cheapest RTX 4090: server ID {cheapest['id']}, "
      f"price {cheapest['price']['on_demand']['bitcoin']:.8f} BTC/day")

# 4. Create an order
try:
    client.create_order(
        server_id=cheapest["id"],
        image="cloreai/ubuntu20.04-jupyter",
        order_type="on-demand",
        currency="bitcoin",
        ports={"22": "tcp", "8888": "http"},
        ssh_password="SecurePass123",
        jupyter_token="MyToken123",
    )
    print("Order created successfully!")
except CloreAPIError as e:
    print(f"Failed to create order: {e}")
    exit(1)

# 5. Check your orders
orders = client.get_orders()
for order in orders:
    if not order.get("expired"):
        cluster = order.get("pub_cluster", [])
        tcp = order.get("tcp_ports", [])
        print(f"Order {order['id']}: server {order['si']}")
        if cluster and tcp:
            ssh_port = tcp[0].split(":")[1]
            print(f"  SSH: ssh root@{cluster[0]} -p {ssh_port}")
```

***

## Node.js Example

Using the native `fetch` API (Node.js 18+):

```javascript
const BASE_URL = 'https://api.clore.ai/v1';

class CloreClient {
  constructor(apiKey) {
    this.apiKey = apiKey;
    this.defaultHeaders = {
      'auth': apiKey,
    };
  }

  async get(endpoint, params = {}) {
    const url = new URL(`${BASE_URL}/${endpoint}`);
    Object.entries(params).forEach(([k, v]) => url.searchParams.set(k, v));
    
    const res = await fetch(url.toString(), {
      headers: this.defaultHeaders,
    });
    
    const data = await res.json();
    if (data.code !== 0) {
      throw new Error(`Clore API error ${data.code}: ${data.error || 'unknown'}`);
    }
    return data;
  }

  async post(endpoint, body) {
    const res = await fetch(`${BASE_URL}/${endpoint}`, {
      method: 'POST',
      headers: {
        ...this.defaultHeaders,
        'Content-type': 'application/json',
      },
      body: JSON.stringify(body),
    });
    
    const data = await res.json();
    if (data.code !== 0) {
      throw new Error(`Clore API error ${data.code}: ${data.error || 'unknown'}`);
    }
    return data;
  }

  async listServers() {
    const data = await this.get('marketplace');
    return data.servers;
  }

  async getOrders(includeCompleted = false) {
    const params = includeCompleted ? { return_completed: 'true' } : {};
    const data = await this.get('my_orders', params);
    return data.orders;
  }

  async createOrder({ serverId, image, type = 'on-demand', currency = 'bitcoin', spotprice, ports, sshPassword, jupyterToken, env, command }) {
    const body = {
      currency,
      image,
      renting_server: serverId,
      type,
      ...(spotprice && { spotprice }),
      ...(ports && { ports }),
      ...(sshPassword && { ssh_password: sshPassword }),
      ...(jupyterToken && { jupyter_token: jupyterToken }),
      ...(env && { env }),
      ...(command && { command }),
    };
    return this.post('create_order', body);
  }

  async cancelOrder(orderId, issue = null) {
    const body = { id: orderId };
    if (issue) body.issue = issue;
    return this.post('cancel_order', body);
  }
}

// Usage example
const client = new CloreClient(process.env.CLORE_API_KEY);

async function main() {
  // List marketplace
  const servers = await client.listServers();
  console.log(`Marketplace has ${servers.length} servers`);

  // Find available RTX 4090 servers
  const available = servers.filter(
    s => s.specs.gpu.includes('4090') && !s.rented
  );
  console.log(`Available RTX 4090 servers: ${available.length}`);

  // Get current orders
  const orders = await client.getOrders();
  const activeOrders = orders.filter(o => !o.expired);
  console.log(`Active orders: ${activeOrders.length}`);

  for (const order of activeOrders) {
    const host = order.pub_cluster?.[0];
    const portMapping = order.tcp_ports?.[0];
    if (host && portMapping) {
      const sshPort = portMapping.split(':')[1];
      console.log(`Order ${order.id} → ssh root@${host} -p ${sshPort}`);
    }
  }
}

main().catch(console.error);
```

***

## Common Workflows

### Find Cheapest RTX 4090 and Rent It

```python
from clore_client import CloreClient

client = CloreClient(api_key="YOUR_API_KEY")

def rent_cheapest_rtx4090(image="cloreai/ubuntu20.04-jupyter", ssh_password="SecurePass123"):
    servers = client.list_servers()
    
    # Filter: available RTX 4090
    candidates = [
        s for s in servers
        if "4090" in s["specs"].get("gpu", "")
        and not s["rented"]
    ]
    
    if not candidates:
        raise RuntimeError("No available RTX 4090 servers found")
    
    # Sort by on-demand BTC price
    candidates.sort(key=lambda s: s["price"]["on_demand"]["bitcoin"])
    best = candidates[0]
    
    price_btc = best["price"]["on_demand"]["bitcoin"]
    print(f"Renting server {best['id']}: {best['specs']['gpu']} @ {price_btc:.8f} BTC/day")
    
    client.create_order(
        server_id=best["id"],
        image=image,
        order_type="on-demand",
        currency="bitcoin",
        ports={"22": "tcp"},
        ssh_password=ssh_password,
    )
    
    print("Done! Check your orders for SSH connection details.")
    return best["id"]

rent_cheapest_rtx4090()
```

***

### Monitor My Orders

```python
import time
from clore_client import CloreClient

client = CloreClient(api_key="YOUR_API_KEY")

def monitor_orders(poll_interval_seconds=60):
    """Poll orders and print status updates."""
    print(f"Monitoring orders (polling every {poll_interval_seconds}s). Ctrl+C to stop.\n")
    
    while True:
        orders = client.get_orders(include_completed=False)
        active = [o for o in orders if not o.get("expired")]
        
        print(f"--- {len(active)} active order(s) ---")
        for order in active:
            cluster = order.get("pub_cluster", [])
            tcp = order.get("tcp_ports", [])
            spend = order.get("spend", 0)
            
            ssh_info = ""
            if cluster and tcp:
                port = tcp[0].split(":")[1]
                ssh_info = f" | SSH: {cluster[0]}:{port}"
            
            print(f"  Order {order['id']}: server {order['si']}"
                  f" | spent {spend:.8f} BTC{ssh_info}")
        
        if not active:
            print("  No active orders.")
        
        print()
        time.sleep(poll_interval_seconds)

monitor_orders()
```

***

### Auto-Rent When Price Drops Below X

```python
import time
from clore_client import CloreClient

client = CloreClient(api_key="YOUR_API_KEY")

def auto_rent_on_price_drop(
    gpu_model: str = "RTX 4090",
    max_price_btc: float = 0.00015,
    image: str = "cloreai/ubuntu20.04-jupyter",
    ssh_password: str = "SecurePass123",
    check_interval_seconds: int = 120,
):
    """
    Watch the marketplace and auto-rent a GPU when price drops below threshold.
    
    Args:
        gpu_model: GPU name to search for (case-insensitive)
        max_price_btc: Maximum acceptable price per day in BTC
        image: Docker image to deploy
        ssh_password: SSH password for the container
        check_interval_seconds: How often to check (respect rate limits!)
    """
    print(f"Watching for {gpu_model} at ≤ {max_price_btc:.8f} BTC/day...")
    
    while True:
        servers = client.list_servers()
        
        for server in servers:
            gpu = server["specs"].get("gpu", "")
            if gpu_model.lower() not in gpu.lower():
                continue
            if server["rented"]:
                continue
            
            price = server["price"]["on_demand"]["bitcoin"]
            if price <= max_price_btc:
                print(f"🎯 Found match! Server {server['id']}: {gpu} @ {price:.8f} BTC/day")
                
                try:
                    client.create_order(
                        server_id=server["id"],
                        image=image,
                        order_type="on-demand",
                        currency="bitcoin",
                        ports={"22": "tcp"},
                        ssh_password=ssh_password,
                        required_price=price,  # Lock in this price
                    )
                    print(f"✅ Order created for server {server['id']}!")
                    return server["id"]
                except Exception as e:
                    print(f"Failed to create order: {e}. Will retry...")
        
        print(f"No match yet. Checking again in {check_interval_seconds}s...")
        time.sleep(check_interval_seconds)

auto_rent_on_price_drop(gpu_model="4090", max_price_btc=0.00012)
```

***

## WebSocket

The Clore.ai REST API does not currently expose a WebSocket endpoint. For real-time monitoring, use polling with a reasonable interval (see rate limits below).

***

## Rate Limits

| Endpoint                           | Limit                    |
| ---------------------------------- | ------------------------ |
| Most endpoints                     | **1 request/second**     |
| `create_order`                     | **1 request/5 seconds**  |
| `set_spot_price` (price reduction) | Once per **600 seconds** |

**Rate limit response (code 5):**

```json
{ "code": 5 }
```

**Best practices:**

* Add `time.sleep(1)` between consecutive API calls
* For `create_order`, wait at least 5 seconds between requests
* Use polling intervals of 60+ seconds for monitoring loops
* Cache marketplace data locally if you need to query it frequently

**Python helper:**

```python
import time

def safe_api_call(fn, *args, delay=1.1, **kwargs):
    """Wrap an API call with rate limit safety."""
    result = fn(*args, **kwargs)
    time.sleep(delay)
    return result
```

***

## Error Handling

Every API response includes a `code` field. A value of `0` means success.

### Error Codes

| Code | Meaning                               | Action                               |
| ---- | ------------------------------------- | ------------------------------------ |
| `0`  | Success                               | —                                    |
| `1`  | Database error                        | Retry after a delay                  |
| `2`  | Invalid input data                    | Check your request body/parameters   |
| `3`  | Invalid API token                     | Verify your API key in the dashboard |
| `4`  | Invalid endpoint                      | Check the endpoint URL               |
| `5`  | Rate limit exceeded (1 req/sec)       | Add delays between requests          |
| `6`  | Application error (see `error` field) | Read the `error` field for details   |

### Code 6 Sub-errors

| `error` value                 | Meaning                                                              |
| ----------------------------- | -------------------------------------------------------------------- |
| `exceeded_max_step`           | Spot price reduction too large; check `max_step` field               |
| `can_lower_every_600_seconds` | Must wait before lowering spot price again; check `time_to_lowering` |

### Python Error Handling Example

```python
from clore_client import CloreClient, CloreAPIError
import time

client = CloreClient(api_key="YOUR_API_KEY")

def create_order_with_retry(server_id, image, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.create_order(
                server_id=server_id,
                image=image,
                order_type="on-demand",
                currency="bitcoin",
                ports={"22": "tcp"},
                ssh_password="SecurePass123",
            )
        except CloreAPIError as e:
            if e.code == 5:  # Rate limit
                print(f"Rate limited. Waiting 5s... (attempt {attempt+1}/{max_retries})")
                time.sleep(5)
            elif e.code == 3:  # Bad API key
                print("Invalid API key! Check your CLORE_API_KEY.")
                raise
            elif e.code == 2:  # Bad input
                print(f"Invalid request: {e}")
                raise
            else:
                print(f"API error: {e}. Retrying in 3s...")
                time.sleep(3)
    
    raise RuntimeError(f"Failed after {max_retries} attempts")
```

### JavaScript Error Handling Example

```javascript
async function createOrderWithRetry(client, serverConfig, maxRetries = 3) {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await client.createOrder(serverConfig);
    } catch (err) {
      const code = parseInt(err.message.match(/error (\d+)/)?.[1]);
      
      if (code === 5) {
        console.log(`Rate limited. Waiting 5s... (attempt ${attempt + 1}/${maxRetries})`);
        await new Promise(r => setTimeout(r, 5000));
      } else if (code === 3) {
        throw new Error('Invalid API key');
      } else {
        console.log(`API error: ${err.message}. Retrying in 3s...`);
        await new Promise(r => setTimeout(r, 3000));
      }
    }
  }
  throw new Error(`Failed after ${maxRetries} attempts`);
}
```

***

## Available Docker Images

Clore.ai provides pre-built images optimized for GPU workloads:

| Image                         | Description               |
| ----------------------------- | ------------------------- |
| `cloreai/ubuntu20.04-jupyter` | Ubuntu 20.04 + JupyterLab |
| `cloreai/ubuntu22.04-jupyter` | Ubuntu 22.04 + JupyterLab |

You can also use any public Docker Hub image. For GPU access, use CUDA-enabled images, e.g.:

```
nvidia/cuda:12.1.0-devel-ubuntu22.04
```

***

## Port Forwarding

When creating an order, specify ports to expose:

```json
{
  "ports": {
    "22": "tcp",
    "8888": "http",
    "6006": "http"
  }
}
```

* **`"tcp"`** — Direct TCP port forwarding (for SSH, custom servers)
* **`"http"`** — HTTPS proxy (for Jupyter, web UIs). Only one HTTP port allowed per `http_port` field.

After order creation, connection details appear in `my_orders`:

```json
{
  "pub_cluster": ["n1.c1.clorecloud.net", "n2.c1.clorecloud.net"],
  "tcp_ports": ["22:10000"],
  "http_port": "8888"
}
```

Connect via SSH:

```bash
ssh root@n1.c1.clorecloud.net -p 10000
```

Access Jupyter via browser: `https://n1.c1.clorecloud.net` (uses the `http_port`)


# New to Clore?

## Supported Currencies

Clore.ai supports the following currencies for deposits and withdrawals:

| Currency    | Network                               | Deposit | Withdrawal |
| ----------- | ------------------------------------- | ------- | ---------- |
| **CLORE**   | Ethereum (ERC-20)                     | ✅       | ✅          |
| **Bitcoin** | Bitcoin Mainnet                       | ✅       | ✅          |
| **USDT**    | Ethereum (ERC-20), BNB Chain (BEP-20) | ✅       | ✅          |
| **USDC**    | Ethereum (ERC-20), BNB Chain (BEP-20) | ✅       | ✅          |

> **Note:** USDT and USDC share a unified USD balance on the platform. You can deposit either stablecoin and they credit to the same balance.

***

## Deposits

### CLORE Token Deposits

| Parameter              | Value           |
| ---------------------- | --------------- |
| Minimum deposit        | 2,000 CLORE     |
| Maximum deposit        | 1,000,000 CLORE |
| Confirmations required | 5               |

### Bitcoin Deposits

| Parameter              | Value      |
| ---------------------- | ---------- |
| Minimum deposit        | No minimum |
| Confirmations required | 3          |

### USDT / USDC Deposits

| Parameter              | Value                                                                        |
| ---------------------- | ---------------------------------------------------------------------------- |
| Networks               | Ethereum (ERC-20), BNB Chain (BEP-20)                                        |
| Minimum deposit        | $10                                                                          |
| Maximum deposit        | $20,000,000                                                                  |
| Confirmations required | Up to 15, depending on the network (exact count shown in the deposit dialog) |

> **Important:** USDT and USDC deposits use the **same address** as your CLORE deposit address. Send on the **Ethereum (ERC-20)** or **BNB Chain (BEP-20)** network. Tokens sent on other networks (Polygon, Arbitrum, Tron, etc.) cannot be recovered.

> **Note:** Both USDT and USDC credit to a unified USD balance. You can deposit either stablecoin using the same address.

### How to Deposit

1. Log in to your Clore.ai account
2. Go to **Account** → **Billing**
3. Click **Deposit** next to your desired currency
4. Copy the deposit address
5. Send funds from your wallet or exchange
6. Wait for confirmations

> **Important:** Always verify the deposit address before sending. CLORE tokens must be sent on the Ethereum network.

***

## Withdrawals

### CLORE Token Withdrawals

| Parameter          | Value            |
| ------------------ | ---------------- |
| Minimum withdrawal | 200 CLORE        |
| Maximum withdrawal | 500,000 CLORE    |
| Maximum per day    | 500,000 CLORE    |
| Network fee        | \~$0.50 in CLORE |

### Bitcoin Withdrawals

| Parameter          | Value       |
| ------------------ | ----------- |
| Minimum withdrawal | 0.0001 BTC  |
| Maximum withdrawal | 0.1 BTC     |
| Maximum per day    | 1 BTC       |
| Network fee        | 0.00005 BTC |

### USDT Withdrawals

| Parameter          | Value                                 |
| ------------------ | ------------------------------------- |
| Networks           | Ethereum (ERC-20), BNB Chain (BEP-20) |
| Minimum withdrawal | $10                                   |
| Maximum withdrawal | $1,000                                |
| Maximum per day    | $10,000                               |
| Network fee        | $1                                    |

### USDC Withdrawals

| Parameter          | Value                                 |
| ------------------ | ------------------------------------- |
| Networks           | Ethereum (ERC-20), BNB Chain (BEP-20) |
| Minimum withdrawal | $10                                   |
| Maximum withdrawal | $1,000                                |
| Maximum per day    | $10,000                               |
| Network fee        | $1                                    |

### How to Withdraw

1. Go to **Account** → **Billing**
2. Click **Withdraw** next to your currency
3. Enter the destination wallet address
4. Enter the amount
5. Confirm the withdrawal

> **Note:** For USDT/USDC withdrawals, specify which stablecoin (USDT or USDC) you want to receive.

***

## Processing Times

| Transaction Type              | Typical Time |
| ----------------------------- | ------------ |
| Deposit (after confirmations) | Instant      |
| Withdrawal                    | Instant      |

***

## Tracking Transactions

Each balance card shows its recent transactions with block-explorer links. The full history lives on the Transactions page, with filters for network, asset, type, currency, direction, and date.

***

## Troubleshooting

### Deposit Not Showing

1. Verify you sent to the correct address
2. Check that the transaction has enough confirmations
3. For CLORE: Ensure you sent on the Ethereum network (ERC-20)
4. For USDT/USDC: Ensure you sent on **Ethereum (ERC-20)** or **BNB Chain (BEP-20)** (not Polygon, Arbitrum, Tron, etc.)
5. Wait up to 30 minutes after confirmations complete

### Stuck Withdrawal

If your withdrawal is pending for more than 2 hours:

1. Check the transaction status in your Account history
2. Contact support on [Discord](https://discord.gg/clore-ai)

***

## Security Notes

* Always double-check addresses before sending
* Clore.ai will never ask for your private keys
* Enable 2FA for additional account security
* Withdrawals to new addresses may require email confirmation

{% content-ref url="/pages/1MI8KnWG6NSNDCS7PiDM" %}
[Account Security](/getting-started/deposit-withdrawal/account-security)
{% endcontent-ref %}


# Supported Currencies

Clore.ai supports multiple currencies for deposits, withdrawals, and marketplace payments.

***

## Currency Overview

| Currency    | Ticker | Network                               | Deposit | Withdraw | Pay for Rentals | Receive Rewards |
| ----------- | ------ | ------------------------------------- | ------- | -------- | --------------- | --------------- |
| **CLORE**   | CLORE  | Ethereum (ERC-20)                     | ✅       | ✅        | ✅               | ✅ All rewards   |
| **Bitcoin** | BTC    | Bitcoin Mainnet                       | ✅       | ✅        | ✅               | Referrals only  |
| **USDT**    | USDT   | Ethereum (ERC-20), BNB Chain (BEP-20) | ✅       | ✅        | ✅               | Referrals only  |
| **USDC**    | USDC   | Ethereum (ERC-20), BNB Chain (BEP-20) | ✅       | ✅        | ✅               | Referrals only  |

***

## CLORE Token

CLORE is the primary platform token. It is used for all platform activities: renting GPUs, hosting rewards, MFP Lock rewards, PoH staking rewards, and referral earnings.

| Property         | Value                                                                              |
| ---------------- | ---------------------------------------------------------------------------------- |
| Token Type       | ERC-20                                                                             |
| Contract Address | `0xe60201989b8628f43dc0605f585a72bcf1f1e977`                                       |
| Decimals         | 18                                                                                 |
| Block Explorer   | [Etherscan](https://etherscan.io/token/0xe60201989b8628f43dc0605f585a72bcf1f1e977) |

**Benefits of paying with CLORE:**

* **No extra fees** — CLORE payments have 0% extra currency surcharge
* **Full platform access** — all features available
* **Receive all reward types** — MFP, PoH, hosting, and referrals

***

## Bitcoin (BTC)

Bitcoin can be used for deposits, withdrawals, and paying for GPU rentals.

| Property       | Value                                   |
| -------------- | --------------------------------------- |
| Network        | Bitcoin Mainnet                         |
| Decimals       | 8                                       |
| Block Explorer | [Blockstream](https://blockstream.info) |

**Note:** Bitcoin payments carry an extra hoster fee (currently 15%) which can be reduced to 0% through [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) tiers. See [Fee Structure](/for-renters/fee-structure) for details.

***

## USDT (Tether USD)

USDT is a stablecoin pegged to the US Dollar. Clore.ai accepts it on **two networks: Ethereum (ERC-20) and BNB Chain (BEP-20)**. It can be used for deposits, withdrawals, GPU rentals, and referral payouts.

| Property            | Value                                                                              |
| ------------------- | ---------------------------------------------------------------------------------- |
| Networks            | Ethereum (ERC-20), BNB Chain (BEP-20)                                              |
| Contract (Ethereum) | `0xdac17f958d2ee523a2206206994597c13d831ec7`                                       |
| Decimals (Ethereum) | 6                                                                                  |
| Block Explorer      | [Etherscan](https://etherscan.io/token/0xdac17f958d2ee523a2206206994597c13d831ec7) |

**Note:** USDT payments carry an extra hoster fee (currently 15%) which can be reduced to 0% through [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) tiers.

> **Important:** Send USDT only on the **Ethereum (ERC-20)** or **BNB Chain (BEP-20)** network. Tokens sent on other networks (Polygon, Arbitrum, Tron, etc.) cannot be recovered.

***

## USDC (USD Coin)

USDC is a stablecoin pegged to the US Dollar. Clore.ai accepts it on **two networks: Ethereum (ERC-20) and BNB Chain (BEP-20)**. It can be used for deposits, withdrawals, GPU rentals, and referral payouts.

| Property            | Value                                                                              |
| ------------------- | ---------------------------------------------------------------------------------- |
| Networks            | Ethereum (ERC-20), BNB Chain (BEP-20)                                              |
| Contract (Ethereum) | `0xa0b86991c6218b36c1d19d4a2e9eb0ce3606eb48`                                       |
| Decimals (Ethereum) | 6                                                                                  |
| Block Explorer      | [Etherscan](https://etherscan.io/token/0xa0b86991c6218b36c1d19d4a2e9eb0ce3606eb48) |

**Note:** USDC payments carry an extra hoster fee (currently 15%) which can be reduced to 0% through [MFP Lock](/for-hosts/mfp-lock-a-complete-breakdown-of-mechanics) tiers.

> **Important:** Send USDC only on the **Ethereum (ERC-20)** or **BNB Chain (BEP-20)** network. Tokens sent on other networks cannot be recovered.

***

## Unified USD Balance

USDT and USDC share a unified USD balance on the platform:

* Depositing **either** USDT or USDC credits your USD balance
* Both stablecoins use the **same deposit address** (which is also your CLORE deposit address)
* When withdrawing, you choose whether to receive **USDT** or **USDC**
* The USD balance is displayed with 2 decimal places

***

## Deposit & Withdrawal Limits

| Currency | Min Deposit | Max Deposit | Min Withdrawal | Max Withdrawal | Daily Limit | Withdrawal Fee   |
| -------- | ----------- | ----------- | -------------- | -------------- | ----------- | ---------------- |
| CLORE    | 2,000       | 1,000,000   | 200            | 500,000        | 500,000     | \~$0.50 in CLORE |
| BTC      | No minimum  | —           | 0.0001 BTC     | 0.1 BTC        | 1 BTC       | 0.00005 BTC      |
| USDT     | $10         | $20,000,000 | $10            | $1,000         | $10,000     | $1               |
| USDC     | $10         | $20,000,000 | $10            | $1,000         | $10,000     | $1               |

***

## Confirmations Required

| Currency | Confirmations                       | Approximate Time |
| -------- | ----------------------------------- | ---------------- |
| CLORE    | 5 blocks                            | \~1 minute       |
| BTC      | 3 blocks                            | \~30 minutes     |
| USDT     | Up to 15 blocks (network-dependent) | a few minutes    |
| USDC     | Up to 15 blocks (network-dependent) | a few minutes    |

***

## Related Pages

{% content-ref url="/pages/8fTxdLzWSsDxoko7LjtC" %}
[New to Clore?](/getting-started/deposit-withdrawal)
{% endcontent-ref %}

{% content-ref url="/pages/3Fy2rz08NzsQhx7vBl4g" %}
[Fee Structure](/for-renters/fee-structure)
{% endcontent-ref %}


# Account Security

Protect your Clore.ai account with these security features and best practices.

## Sign-In Methods

You can create an account and sign in with any of the following:

| Method               | Notes                                |
| -------------------- | ------------------------------------ |
| **Email + password** | Classic registration, works with 2FA |
| **Google**           | One-click OAuth sign-in              |
| **GitHub**           | One-click OAuth sign-in              |
| **Hugging Face**     | One-click OAuth sign-in              |
| **HiveOS**           | Sign in with your HiveOS account     |

All methods create a regular Clore.ai account. For OAuth sign-ins, keep the linked provider account (Google/GitHub/Hugging Face/HiveOS) secure: whoever controls it controls your Clore.ai login.

## Two-Factor Authentication (2FA)

Two-factor authentication adds an extra layer of security to your account.

### How to Enable 2FA

1. Go to **Account** → **Security**
2. Click **Enable 2FA**
3. Scan the QR code with your authenticator app:
   * Google Authenticator
   * Authy
   * Microsoft Authenticator
4. Enter the 6-digit code from your app
5. Save your backup codes securely

### Backup Codes

When enabling 2FA, you'll receive backup codes. **Store these safely!**

* Each code can only be used once
* Use them if you lose access to your authenticator
* Generate new codes if you run out

### Disabling 2FA

1. Go to **Account** → **Security**
2. Click **Disable 2FA**
3. Enter your current 2FA code
4. Confirm the action

> **Warning:** Disabling 2FA reduces your account security. Only do this if necessary.

## Password Security

### Strong Password Requirements

* Minimum 8 characters
* Mix of uppercase and lowercase letters
* Include numbers and special characters
* Avoid common words or personal information

### Changing Your Password

1. Go to **Account** → **Security**
2. Click **Change Password**
3. Enter current password
4. Enter new password twice
5. Save changes

### Password Recovery

If you forget your password:

1. Go to login page
2. Click **Forgot Password**
3. Enter your email address
4. Check email for reset link
5. Create a new password

## API Keys

API keys allow programmatic access to your account.

### Managing API Keys

1. Go to **Account** → **API Keys**
2. Click **Generate New Key**
3. Copy and save the key immediately (shown only once)
4. Set appropriate permissions

### API Key Best Practices

* **Never share** your API keys
* Use separate keys for different applications
* Revoke unused keys
* Monitor API usage regularly
* Limit key permissions to what's needed

### Key Limits

* Maximum **3 API keys** per account
* Keys can be revoked anytime

## SSH Keys

SSH keys provide secure, passwordless access to rented servers.

### Adding SSH Keys

1. Go to **Account** → **SSH Keys**
2. Click **Add Key**
3. Paste your public key
4. Give it a descriptive name
5. Save

### SSH Key Limits

* Maximum **3 SSH keys** per account
* Keys are automatically deployed to new rentals

## Login Sessions

### Active Sessions

View and manage your active login sessions:

1. Go to **Account** → **Security**
2. View **Active Sessions**
3. Revoke any suspicious sessions

### Session Limits

* Maximum **50 active sessions** per account
* Sessions auto-expire after **14 days** of inactivity

## Security Best Practices

### Do's

* Enable 2FA immediately after registration
* Use a unique password for Clore.ai
* Regularly review active sessions
* Keep your email account secure
* Log out on shared devices

### Don'ts

* Never share your password or 2FA codes
* Don't use the same password on multiple sites
* Never give API keys to untrusted applications
* Don't ignore suspicious login notifications

## If Your Account is Compromised

1. **Change password immediately**
2. **Revoke all active sessions**
3. **Regenerate API keys**
4. **Enable 2FA** if not already enabled
5. **Contact support** on [Discord](https://discord.gg/clore-ai)
6. **Review recent transactions** for unauthorized activity

## Contact Security Team

For security concerns or to report vulnerabilities:

* Discord: [discord.gg/clore-ai](https://discord.gg/clore-ai) (use #support channel)
* Email: Check official channels for security contact


# Referral Program

Invite friends to Clore.ai and earn commission from the platform's profits on every transaction they make on the marketplace.

## 🔥 Limited-Time Promotion: 50% Commission

For a limited promotional period, **all referrers earn 50% commission** — regardless of their normal tier. This promotion:

* Applies to **all users** automatically
* Works for **all currencies** (CLORE, BTC, USD)
* Is **time-limited** — the system automatically reverts to standard rates when the promotion ends
* Requires **no action** — if you have a referral link, you're already earning 50%

> Check the Referral page in your account to see if the promotion is currently active.

***

## Standard Commission Tiers

Outside of promotional periods, commission rates are:

| Tier     | Commission Rate | How to Get            |
| -------- | --------------- | --------------------- |
| Standard | 15%             | Default for all users |
| Premium  | 20%             | Assigned by admin     |
| VIP      | 30%             | Assigned by admin     |

***

## Who Earns You Commission?

You earn commission from **both parties** in a transaction:

* **Referred renters** (customers who rent GPUs)
* **Referred hosts** (server owners providing GPUs)

Each side's referrer earns independently from that side's actual fees paid.

***

## How Commission is Calculated

Your referral reward is calculated as a percentage of the actual fees paid by the referred user:

**For a referred renter:**

```
renter_referral = (renter_base_fee + renter_extra_fee) × commission_rate
```

**For a referred host:**

```
hoster_referral = (hoster_base_fee + hoster_extra_fee) × commission_rate
```

> **Note:** If the referred user has PoH holdings that reduce their fees, your referral earnings are based on their **reduced** fee amount.

***

## Example

Your friend rents a GPU for 100 CLORE/day using On-Demand (10% base fee), no PoH:

| Component                     | Amount   |
| ----------------------------- | -------- |
| Renter base fee (half of 10%) | 5 CLORE  |
| Hoster base fee (half of 10%) | 5 CLORE  |
| Platform total fees           | 10 CLORE |

**Your referral earnings (at 15% standard rate):**

| If You Referred | Your Earning             |
| --------------- | ------------------------ |
| The renter      | 5 × 15% = **0.75 CLORE** |
| The host        | 5 × 15% = **0.75 CLORE** |
| Both            | **1.50 CLORE**           |

**During 50% promotion:**

| If You Referred | Your Earning             |
| --------------- | ------------------------ |
| The renter      | 5 × 50% = **2.50 CLORE** |
| The host        | 5 × 50% = **2.50 CLORE** |
| Both            | **5.00 CLORE**           |

***

## Minimum Withdrawal Amounts

| Currency        | Minimum     |
| --------------- | ----------- |
| Bitcoin         | 0.00005 BTC |
| CLORE           | 50 CLORE    |
| USD (USDT/USDC) | $1          |

***

## How to Get Your Referral Link

1. Log in to your Clore.ai account
2. Go to the **Account** page
3. Find the **Referral Program** section
4. Copy your unique link: `https://clore.ai/login?ref_id=YOUR_CODE` (your code is applied automatically when your invitee signs up)

***

## How to Withdraw Earnings

1. Accumulate at least the minimum withdrawal amount
2. Click **"Withdraw"** next to your referral balance
3. Funds will be added to your main account balance

***

## Important Notes

* Referral earnings are accumulated in real-time as your referred users transact
* Earnings are calculated separately for each currency (BTC, CLORE, USD)
* Commission applies to marketplace transactions only (GPU rentals)
* Your referral code never expires
* PoH fee reductions affect the fee base used for referral calculations
* During promotions, all users earn the promotional rate regardless of their tier


# Changelog

{% hint style="info" %}
This changelog covers platform and marketplace updates. For token-related news, see [Tokenomics](/main/tokenomics).
{% endhint %}

## Coming Soon

* **Clore VPN** — Native iOS, Android, Windows, and macOS apps (currently under App Store / Google Play review).
* **Web3 Login** — Sign in via MetaMask or Solana wallet.

***

## 2026

### July 2026

* **Partial GPU Rental Rollout**: rent a subset of GPUs on supported multi-GPU servers and pay proportionally; available on homogeneous rigs with up-to-date hosting software, with a per-server opt-out for hosts.
* **Bare Metal: Up to 1,000 GPUs**: maximum order size extended to 1,000 GPUs per configuration; RTX 5080 and RTX 5070 added to the catalog; rental terms extended up to 5 years.
* **My Servers: Live stats**: real-time per-GPU and CPU/RAM stats column in the server list, enabled by default.
* **Hosting Installer Channels**: My Servers onboarding redesigned with Stable and Beta installer channels, one-line install and register commands, and a legacy fallback.

### June 2026

* **Marketplace Power Tools**: private server notes, pinned servers, favourites and blacklist, table view, per-GPU prices, and rating sort.
* **PoH Marketplace In-App**: PoH lending and renting integrated directly into the platform at clore.ai/poh-marketplace.
* **Interactive API Docs**: browsable API reference at clore.ai/api-docs.
* **Price Coach**: built-in pricing guidance that suggests competitive prices based on live marketplace data.
* **Bare Metal Checkout Simplified**: sign-in and account creation moved into the checkout modal (no login wall), SSH key optional with saved-key autofill.
* **Higher MFP Cap**: Maximum Fair Price upper limit raised to 200.

### May 2026

* **Marketplace V3 Live**: the rebuilt marketplace became the default: three-step rent flow (image, configure, pay), in-modal deposit and sign-in, expanded filters (price, disk, CPU, rating, GPU model), USD-first pricing.
* **New Sign-In Methods**: OAuth sign-in with Google, GitHub, and Hugging Face, alongside HiveOS and email.
* **Bare Metal Self-Serve Configurator**: multi-region catalog (USA, Japan, EU, and more), instant USD per-GPU-hour quotes, USD stablecoin payments, browse and configure without an account.
* **USD Deposit Cap Raised**: maximum USDT/USDC deposit increased to $20,000,000.
* **Simplified Host Pricing**: USD/day pricing view with automatic BTC/CLORE conversion; per-hour and per-day price display.
* **MFP Base Recalculation**: base prices raised across the GPU catalog and previously unmapped GPU models added.
* **Analytics: Yearly View**: revenue and spending chart gains a yearly view with monthly aggregation.

### April 2026

* **CLORE Withdrawal Fee Pegged to USD**: withdrawal fee now targets \~$0.50 worth of CLORE regardless of token price.

### March 2026

* **Transaction History Filters**: filter transactions by network, asset, type, currency, direction, and date, with fuller per-transaction details.

***

## 2025

### November 2025

* **New AI Models** — Confucius and Acemath models added to the base image catalog.
* **My Orders / My Servers Redesign** — Complete UX overhaul of both pages for improved usability and clarity.
* **Visible Forwarded Ports** — Forwarded ports are now shown directly in the My Orders view.
* **Bare Metal Cocoon Support** — Clore Partners integration adds Cocoon support for Bare Metal servers.
* **Rentals up to 3,000 Hours** — Maximum rental duration extended to 3,000 hours.
* **CMP170HX Fixes** — Compatibility and stability improvements for CMP170HX GPUs.
* **Lock MFP Link** — Hosts can now lock the Maximum Fair Price link.

### October 2025

* **Hive OS Integration** — Technical partnership with Hive OS brings native integration into the platform.
* **Redesigned Login & Registration** — Refreshed auth flow with improved UX and dedicated support page.
* **Richer Notifications** — Expanded notification system with more event types and detail.
* **Higher MFP Limits** — Maximum Fair Price caps increased for hosts.
* **Host–Renter Messaging Survey** — Public survey launched to shape upcoming messaging feature.

### September 2025

* **Aethir Integration** — Bare Metal partnership with Aethir adds decentralized GPU supply to the platform.
* **Bare Metal New Locations & Configs** — Additional data center locations and enterprise hardware configurations for Bare Metal.
* **PoH — Tangem Hardware Wallet Support** — Proof-of-Holding now supports Tangem hardware wallets.

### August 2025

* **Bare Metal by Clore Launch** — Non-virtualized GPU servers with full root access, targeting AI and HPC workloads.
* **New Base Images** — Added QWQ, Qwen, Gemma, and DBRX to the default image library.
* **My Orders Beta Layout** — New beta UI for the My Orders page.
* **2FA & Daily Withdrawal Limits** — Two-factor authentication hardened; daily withdrawal limits added for security.
* **Multilingual Documentation** — Official docs now available in multiple languages.

### July 2025

* **MyServers Page Overhaul** — Significant design and usability improvements to the server management page.

### June 2025

* **Reward Distribution Redesign** — Reward distribution system rebuilt for accuracy and transparency.
* **PoH Staking** — Proof-of-Holding staking mechanism introduced.
* **Testnet PoS Launch** — Proof-of-Stake testnet goes live.

### May 2025

* **Multi-Currency Deposits** — Users can now deposit in USDT, BTC, ETH, and other currencies.
* **Banned GPU IDs Published** — Public list of banned GPU IDs released for transparency.
* **Template Manager Update** — Container template manager updated with new capabilities.

### April 2025

* **Referral System Update** — Commission split revised to 50/50 between hosts and users; Ambassador program offers up to 30% for bloggers and companies.

### February 2025

* **Free Mobile App** — Released mobile app with PoH Marketplace, auto rig rentals, mining earnings tracker, and Hive OS integration.

### January 2025

* **Gigaspot Auction Platform** — Super-fast auction-based spot market launched, replacing the standard spot market.
* **Payments Switched to CLORE Only** — Bitcoin removed as a payment method; platform moves to CLORE-native payments.

***

## 2024

### December 2024

* **Clore Coin on Tangem Wallet** — CLORE added to Tangem hardware wallet.
* **Vast.ai Integration** — Cross-platform integration with Vast.ai marketplace.
* **User Statistics Dashboard** — New in-platform user statistics and analytics.
* **Main Page & Roadmap Redesign** — Rebuilt landing page and public roadmap.
* **Default Images Rebuilt** — Base container images rebuilt from scratch for improved performance and security.

### November 2024

* **Backend 5.2.8** — NVIDIA GPU compatibility fixes and stability improvements.
* **Clore VPN Beta** — VPN service enters beta; mobile app and VPN plans announced.
* **PoH 2.0** — Second generation Proof-of-Holding system with enhanced mechanics.
* **MPF 3.0** — Maximum Fair Price system updated to version 3.0.
* **PoH Calculator** — New on-platform tool for calculating PoH rewards.

### October 2024

* **Order Fee Reductions** — Lower fees for marketplace orders.
* **Mixed GPU Support** — Servers can now list mixed GPU configurations.
* **Bulk Price Changes** — Hosts can update pricing for multiple servers simultaneously.
* **NVIDIA NCCL Support** — Added NCCL library support for multi-GPU training workloads.
* **Clock Lock Management** — GPU clock lock controls available to renters.
* **Clore Notification Bot** — Telegram bot for real-time platform notifications.
* **Clore Fleet Miner (HiveOS)** — Fleet mining tool for HiveOS users launched.

### September 2024

* **Marketplace V2.0 Beta** — Second major version of the marketplace enters public beta.
* **Clore Encyclopedia** — Knowledge base launched for platform documentation and guides.
* **Automated Trading Nodes** — Infrastructure for automated spot market trading deployed.

### August 2024

* **PoH Marketplace Launch** — Proof-of-Holding marketplace goes live.
* **Team Doxxed + Certik KYC** — Core team publicly identified and verified through Certik KYC process.

### July 2024

* **iOS Wallet** — Native iOS wallet app released.
* **Marketplace V2.0 Beta (UI)** — New marketplace UI beta made available.

### June 2024

* **Android & Web Wallet** — Android app and web-based wallet released.

### May 2024

* **Backend v5.1 + v5.2** — Light background jobs, overclocking stability fixes, auto-updates.
* **Mainline / Power Efficient Split** — Marketplace divided into Mainline and Power Efficient categories.
* **Power Limit Transparency** — GPU power limit data now visible to renters.

### April 2024

* **2FA Authentication** — Two-factor authentication implemented platform-wide.

### March 2024

* **New Marketplace Design** — Refreshed marketplace UI with new filtering and discovery features.

### January 2024

* **Custom Container Templates** — Users can create and reuse custom templates for container deployments.

***

## 2023

### November 2023

* **Private Containers** — Support for private (non-public) container images added.
* **Improved Container Interaction** — Enhanced tools for managing and interacting with running containers.

### October 2023

* **Major MFP Updates** — Significant improvements to the Maximum Fair Price algorithm.
* **CLORE as Main Payment Method** — Platform transitions to CLORE as the primary payment currency.

### August 2023

* **Neural Network Auto MFP** — Machine learning model introduced to automatically set fair pricing for new servers.

### July 2023

* **New Base Images** — TensorFlow, PyTorch, Jupyter Lab, and tftrack added to default image catalog.
* **Review Function** — Users can now leave reviews for hosts and rental experiences.

### June 2023

* **Clore Mining** — Auto-mine top profitable coin when a server is not rented.
* **OC Profiles** — Overclocking profiles support added for GPU hosts.
* **More Forwarding Nodes** — Expanded port forwarding infrastructure.
* **Referral System** — First version of the referral system launched.

### April 2023

* **Dollar Top-Ups** — Fiat top-up support added to accounts.
* **Account Logs** — Detailed activity logs available in user accounts.
* **Maximum Fair Price Feature** — MFP system introduced to protect renters from overpriced listings.

### March 2023

* **Proof-of-Holding System** — PoH launched, rewarding CLORE holders with platform benefits.
* **Hive OS Support** — Integration with Hive OS for GPU rig management.

### February 2023

* **Bug Fixes & Stability** — Hundreds of platform bug fixes based on community feedback.

***

## 2022

### December 2022

* **Clore.ai Marketplace Launch** — Platform publicly launched with GPU rental marketplace and Bitcoin payments.


# CLORE Migration (completed)

<figure><img src="/files/dVLv4B1iA18T4OeY4nGB" alt=""><figcaption></figcaption></figure>

{% hint style="success" %}
**Migration complete.** The CLORE token migration from the legacy PoW chain to Ethereum (ERC-20) was completed in December 2025. All marketplace operations now run on ERC-20 CLORE only.
{% endhint %}

***

## Migration Summary

CLORE was migrated from its original proof-of-work blockchain to Ethereum mainnet (ERC-20) via a 1:1 snapshot and claim process.

### Timeline (completed)

| Event                                                                  | Date (UTC)             |
| ---------------------------------------------------------------------- | ---------------------- |
| Migration announced                                                    | 16 Dec 2025            |
| Legacy deposits/withdrawals closed                                     | 20 Dec 2025            |
| Snapshot taken                                                         | 21 Dec 2025, 00:00 UTC |
| ERC-20 deposits/withdrawals resumed                                    | 22 Dec 2025            |
| Claim window closed                                                    | \~21 Feb 2026          |
| In-platform migration window extended (+10 days, by community request) | 21 Feb 2026            |
| In-platform migration window closed                                    | 3 Mar 2026             |

***

## How it worked

There were exactly **two ways** to complete the migration:

1. **Claim Portal** — Submit a claim request through the official portal. Migrated ERC-20 CLORE would be sent to the wallet address you specified.
2. **Deposit to clore.ai** — Deposit legacy CLORE to your internal clore.ai marketplace wallet, then **confirm your consent to migrate** within the platform.

{% hint style="info" %}
**In-platform extension:** Due to community requests, the **Method 2** migration window inside the clore.ai dashboard was extended by 10 days and closed on **3 March 2026**. This extension applied only to the in-platform deposit method — the Claim Portal closed on its original schedule (\~21 Feb 2026).
{% endhint %}

{% hint style="danger" %}
If you did not complete either of these methods within the 2-month claim window (ending \~21 Feb 2026), your tokens are **permanently and irreversibly lost**. No exceptions, no extensions.
{% endhint %}

***

## Current state

* **Deposits and withdrawals** on clore.ai and exchanges operate exclusively via the **Ethereum (ERC-20) network**.
* The **legacy PoW chain is deprecated** and unsupported — no deposits, withdrawals, or operational support.
* The **claim portal is closed**. The 2-month claim window has ended and will not be reopened.

### Unclaimed tokens

All unclaimed CLORE from the migration will be subject to a **community DAO vote** to determine their fate. The community will decide how these tokens are used — whether they are burned, redistributed, allocated to ecosystem development, or handled in another way. Details on the DAO governance process and voting timeline will be announced through official channels.

***

## Token details

| Property   | Value                                        |
| ---------- | -------------------------------------------- |
| Network    | Ethereum mainnet (ERC-20)                    |
| Contract   | `0xe60201989b8628f43dc0605f585a72bcf1f1e977` |
| Symbol     | CLORE                                        |
| Max supply | 1,000,000,000 (hard-capped, cannot increase) |

{% embed url="<https://etherscan.io/token/0xe60201989b8628f43dc0605f585a72bcf1f1e977>" %}

***

## Anti-scam reminder

* We will **never** ask for your seed phrase or private key.
* We will **never** DM you support links.
* The claim portal is permanently closed — any site claiming to still offer migration is a scam.

***

{% content-ref url="/pages/eoZkyCjSDfQ9akLAkyHP" %}
[Tokenomics](/main/tokenomics)
{% endcontent-ref %}


# Claim Window (closed)

{% hint style="danger" %}
**Migration permanently closed.** The CLORE token migration from the legacy PoW chain to Ethereum (ERC-20) was completed in December 2025. The 2-month claim window closed in February 2026 and will not be reopened. The claim portal has been permanently deactivated.

**Unclaimed CLORE from the legacy PoW network is permanently unrecoverable. No further migration is possible.**
{% endhint %}

***

## How the migration worked

There were exactly **two ways** to migrate your legacy CLORE tokens to ERC-20:

### Method 1 — Claim Portal

Submit a claim request through the official claim portal. Your migrated ERC-20 CLORE would be sent to the wallet address you specified in the claim.

### Method 2 — Deposit to clore.ai

Deposit your legacy CLORE tokens to your internal clore.ai marketplace wallet. After the deposit was received, you had to **confirm your consent to migrate** within the marketplace.

{% hint style="info" %}
**Extension (Method 2 only):** Due to community requests, the in-platform migration window inside the clore.ai dashboard was extended by 10 days and closed on **3 March 2026**. This extension applied only to this method — the Claim Portal (Method 1) closed on its original schedule (\~21 Feb 2026).
{% endhint %}

{% hint style="danger" %}
If you completed **neither** of these two methods within the 2-month claim window, your tokens are **permanently and irreversibly lost**. No exceptions, no extensions.
{% endhint %}

***

## Current status

* The **claim portal is permanently closed** and has been deactivated.
* The **legacy PoW chain is deprecated** — no deposits, withdrawals, or operational support.
* All CLORE operations now run exclusively on the **Ethereum (ERC-20) network**.
* **No exceptions or extensions** will be made to the claim window.

{% hint style="warning" %}
**Scam warning:** Any website, person, or service claiming to offer a CLORE migration or claim portal is a **scam**. Do not connect your wallet or share your seed phrase. The official claim window ended in February 2026 and will never reopen.
{% endhint %}

***

{% content-ref url="/pages/zCxES2v08ivQd7WQNu7M" %}
[CLORE Migration (completed)](/clore-migration)
{% endcontent-ref %}


# Terms of Service

**Last Updated: March 2026**

{% hint style="info" %}
Please read these Terms of Service carefully before using the Clore.ai platform. By accessing or using Clore.ai, you agree to be bound by these Terms. If you do not agree, you must not use the Platform.
{% endhint %}

***

## 1. Introduction and Acceptance of Terms

### 1.1 Agreement

These Terms of Service ("Terms") constitute a legally binding agreement between you ("User," "you," or "your") and Clore.ai ("Clore.ai," "the Platform," "we," "us," or "our"), governing your access to and use of the Clore.ai website located at [clore.ai](https://clore.ai) and all related services, features, and content (collectively, the "Service").

### 1.2 Binding Nature

By creating an account, accessing, or using the Service in any manner, you acknowledge that you have read, understood, and agree to be bound by these Terms, as well as our Privacy Policy, which is incorporated herein by reference.

### 1.3 Modifications

We reserve the right to modify, amend, or update these Terms at any time, at our sole discretion, without prior notice. Any changes will be effective immediately upon posting. Your continued use of the Service following the posting of revised Terms constitutes your acceptance of such changes. It is your responsibility to check these Terms periodically for changes.

***

## 2. Eligibility and Account Requirements

### 2.1 Age Requirement

{% hint style="info" %}
You must be at least **18 years of age** to use Clore.ai. By using the Platform, you represent and warrant that you are 18 years of age or older.
{% endhint %}

### 2.2 Legal Capacity

You represent and warrant that you have the full legal capacity and authority to enter into these Terms and to use the Service. If you are using the Service on behalf of an organization, you represent that you have the authority to bind that organization to these Terms.

### 2.3 Prohibited Jurisdictions

You represent that you are not located in, nor a national or resident of, any country subject to a U.S., EU, or UN embargo or sanctions regime, and that you are not otherwise prohibited by applicable law from using the Service.

### 2.4 Account Accuracy

You agree to provide accurate, current, and complete information when creating your account and to keep such information up to date at all times. Providing false or misleading information is grounds for immediate account termination.

***

## 3. Nature of the Platform

### 3.1 Peer-to-Peer Marketplace

{% hint style="info" %}
Clore.ai is a **decentralized, peer-to-peer marketplace** that connects individuals who wish to rent GPU computing resources ("Renters") with individuals who wish to offer their GPU hardware for rent ("Hosts"). Clore.ai does **not** own, operate, control, or manage any GPU servers or computing hardware listed on the Platform.
{% endhint %}

### 3.2 No Provider Relationship

Clore.ai acts solely as a marketplace intermediary. We are not a cloud computing provider, data center operator, or internet service provider. Any computing resources you access through the Platform are provided directly by independent third-party Hosts.

### 3.3 No Guarantee of Availability or Uptime

**Clore.ai makes no representations, warranties, or guarantees of any kind regarding:**

* The availability, reliability, or performance of any GPU server listed on the Platform
* Uptime, latency, or connectivity of any Host's hardware
* The continued listing or availability of any particular server or Host
* The quality, fitness, or suitability of any computing resources for your intended use

The Platform provides access to a marketplace only. All service-level commitments, if any, are solely between Renters and Hosts.

### 3.4 Independent Contractors

Hosts are independent contractors and are not employees, agents, partners, or joint venturers of Clore.ai. Clore.ai is not responsible for the acts, omissions, errors, or negligence of any Host.

***

## 4. Fees and Payments

### 4.1 Platform Commission

Clore.ai charges a commission and/or transaction fee on transactions facilitated through the Platform ("Platform Fee"). The applicable fee rates are displayed on the Platform and are subject to change at any time at our sole discretion.

### 4.2 Accepted Payment Methods

Payments on the Platform are processed using the following cryptocurrencies:

* **CLORE** — the native utility token of the Clore.ai ecosystem
* **Bitcoin (BTC)**
* **Tether (USDT)**
* **USD Coin (USDC)**

### 4.3 Fee Changes

Clore.ai reserves the right to change its fee structure at any time without prior notice. Continued use of the Platform after a fee change constitutes acceptance of the new fee structure.

### 4.4 No Refund Policy

{% hint style="info" %}
**All fees and payments made on Clore.ai are generally non-refundable.** Due to the decentralized and peer-to-peer nature of the Platform, Clore.ai does not guarantee refunds for any transactions. Refunds, if any, are issued solely at the discretion of Clore.ai and on a case-by-case basis.
{% endhint %}

### 4.5 Cryptocurrency Volatility

You acknowledge and accept that cryptocurrency values are highly volatile and may fluctuate significantly. Clore.ai is not responsible for any losses arising from cryptocurrency price movements.

### 4.6 Tax Obligations

You are solely responsible for determining and fulfilling any tax obligations arising from your use of the Platform, including but not limited to income tax, capital gains tax, VAT, or GST.

***

## 5. CLORE Token Disclaimer

### 5.1 Utility Token

{% hint style="info" %}
The CLORE token is a **utility token** designed solely for use within the Clore.ai ecosystem. It is not an investment product, security, share, bond, or any other financial instrument.
{% endhint %}

### 5.2 No Investment Advice

Nothing on the Platform constitutes financial, investment, legal, or tax advice. Clore.ai makes no representations or guarantees regarding the future value, price, liquidity, or utility of the CLORE token.

### 5.3 No Securities Representation

The CLORE token has not been registered under any securities law in any jurisdiction. By acquiring or using CLORE tokens, you represent that you are not purchasing them as an investment and do not expect profits from the efforts of Clore.ai or any third party.

### 5.4 ERC-20 Risks

CLORE tokens are issued on the Ethereum blockchain as ERC-20 tokens. You acknowledge the inherent risks of blockchain technology, including but not limited to smart contract bugs, network congestion, loss of private keys, and regulatory changes.

### 5.5 No Price Guarantee

Clore.ai does not guarantee any specific price, value, or exchange rate for CLORE tokens at any time.

***

## 6. Prohibited Uses

### 6.1 General Prohibitions

You agree not to use the Platform, directly or indirectly, for any of the following purposes:

1. **Illegal Activities** — Any activity that violates applicable local, national, or international laws or regulations.
2. **Malware Distribution** — Distributing, deploying, or executing any ransomware, cryptojacking software, botnets, trojans, viruses, or other malicious code on rented resources or targeting third parties.
3. **DDoS and Network Attacks** — Conducting distributed denial-of-service (DDoS) attacks, port scanning, network intrusion, or any form of cyberattack against any system, network, or individual.
4. **Unauthorized Access** — Attempting to gain unauthorized access to any system, account, network, or data.
5. **Fee Circumvention** — Attempting to bypass, circumvent, or manipulate the Platform's fee structure, payment systems, or commission mechanisms.
6. **Spam and Phishing** — Sending unsolicited communications, conducting phishing campaigns, or engaging in social engineering attacks.
7. **Intellectual Property Infringement** — Infringing upon the intellectual property rights of Clore.ai or any third party.
8. **Money Laundering and Fraud** — Using the Platform to launder money, conduct fraud, or engage in any financial crime.
9. **Sanctions Violations** — Engaging in transactions prohibited by applicable sanctions programs.
10. **Child Exploitation** — Accessing, storing, or distributing child sexual abuse material (CSAM) or any content that exploits or endangers minors.
11. **Platform Manipulation** — Manipulating reviews, ratings, or any other Platform integrity mechanism.
12. **Reverse Engineering** — Reverse engineering, decompiling, or disassembling any portion of the Platform's software or systems.

### 6.2 Consequences of Violation

Violation of any of the above prohibitions may result in immediate account suspension or termination, forfeiture of funds, reporting to law enforcement authorities, and civil or criminal legal action.

***

## 7. Account Termination and Suspension

### 7.1 Termination by Clore.ai

{% hint style="info" %}
Clore.ai reserves the right to suspend, restrict, or permanently terminate your account at its **sole discretion**, at any time, with or without cause, and without prior notice or liability to you.
{% endhint %}

### 7.2 Termination by User

You may terminate your account at any time by following the account deletion process on the Platform. Termination does not entitle you to any refund of fees paid.

### 7.3 Effect of Termination

Upon termination, your right to use the Platform immediately ceases. Any outstanding obligations or liabilities incurred prior to termination shall survive.

### 7.4 Preservation of Data

Clore.ai may retain certain data following account termination as required by law or for legitimate business purposes, in accordance with our Privacy Policy.

***

## 8. KYC/AML Compliance

### 8.1 Right to Request Verification

Clore.ai reserves the right, at its sole discretion, to require you to complete Know Your Customer ("KYC") identity verification and/or Anti-Money Laundering ("AML") screening at any time, as a condition of accessing or continuing to use the Service.

### 8.2 Cooperation

You agree to cooperate fully with any KYC/AML verification process, including providing government-issued identification documents, proof of address, and any other documentation reasonably requested.

### 8.3 Suspension Pending Verification

Clore.ai may suspend your account or restrict your access to certain features pending completion of KYC/AML verification.

### 8.4 Right to Decline

Clore.ai reserves the right to decline service to any individual or entity that fails to meet its KYC/AML requirements or that Clore.ai believes may pose a compliance risk.

***

## 9. Intellectual Property

### 9.1 Ownership

All intellectual property rights in and to the Platform, including but not limited to the website, software, algorithms, interfaces, logos, trademarks, trade names, brand elements, documentation, and all related content, are the exclusive property of Clore.ai or its licensors.

### 9.2 Limited License

Subject to your compliance with these Terms, Clore.ai grants you a limited, non-exclusive, non-transferable, revocable license to access and use the Platform solely for your personal or internal business purposes.

### 9.3 Restrictions

You may not copy, reproduce, distribute, modify, create derivative works of, publicly display, publicly perform, republish, download, store, or transmit any part of the Platform's content without the prior written consent of Clore.ai.

### 9.4 User Content

By submitting any content to the Platform (including listings, descriptions, or feedback), you grant Clore.ai a worldwide, royalty-free, non-exclusive license to use, display, and distribute such content for Platform operational purposes.

***

## 10. Limitation of Liability

### 10.1 Disclaimer of Warranties

**THE PLATFORM AND ALL SERVICES ARE PROVIDED ON AN "AS IS" AND "AS AVAILABLE" BASIS, WITHOUT WARRANTIES OF ANY KIND, WHETHER EXPRESS, IMPLIED, STATUTORY, OR OTHERWISE. CLORE.AI EXPRESSLY DISCLAIMS ALL WARRANTIES, INCLUDING BUT NOT LIMITED TO WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, AND NON-INFRINGEMENT.**

### 10.2 Scope of Limitation

{% hint style="info" %}
**TO THE FULLEST EXTENT PERMITTED BY APPLICABLE LAW, CLORE.AI SHALL NOT BE LIABLE FOR ANY:**

* Loss or corruption of data
* Server downtime or hardware failures
* Host misconduct, negligence, or breach of obligations
* Renter misuse of computing resources
* Financial losses, including loss of profits or revenue
* Cryptocurrency or token price fluctuations
* Unauthorized access to or alteration of your data
* Service interruptions, bugs, or errors
* Actions or omissions of third parties
* Any indirect, incidental, special, consequential, punitive, or exemplary damages
  {% endhint %}

### 10.3 Aggregate Liability Cap

To the extent that any liability of Clore.ai cannot be fully excluded under applicable law, our total aggregate liability to you for any claims arising out of or relating to these Terms or the Service shall not exceed the greater of: (a) the total fees paid by you to Clore.ai in the three (3) months preceding the event giving rise to the claim, or (b) USD $100.

### 10.4 Essential Basis

You acknowledge that the limitations of liability in this Section reflect a reasonable allocation of risk and are an essential basis of the bargain between you and Clore.ai, without which Clore.ai would not provide the Service.

***

## 11. Indemnification

### 11.1 User Indemnification Obligation

You agree to defend, indemnify, and hold harmless Clore.ai, its affiliates, officers, directors, employees, contractors, agents, licensors, and successors (collectively, "Clore.ai Parties") from and against any and all claims, damages, losses, liabilities, costs, and expenses (including reasonable attorneys' fees) arising out of or relating to:

1. Your use of or access to the Platform
2. Your breach of these Terms
3. Your violation of any applicable law or regulation
4. Your infringement of any third-party rights
5. Any content you submit to the Platform
6. Any transaction you enter into through the Platform
7. Your actions as a Host or Renter, including any harm caused to third parties

### 11.2 Cooperation

Clore.ai reserves the right to assume the exclusive defense and control of any matter subject to indemnification by you, in which case you agree to cooperate fully with Clore.ai in asserting any available defenses.

***

## 12. Dispute Resolution

### 12.1 Informal Resolution

Before initiating any formal dispute resolution proceedings, you agree to first attempt to resolve any dispute informally by contacting Clore.ai and allowing us a reasonable period of not less than thirty (30) days to address the issue.

### 12.2 Arbitration

{% hint style="info" %}
Any dispute, controversy, or claim arising out of or relating to these Terms or the Service that cannot be resolved informally shall be submitted to and resolved by **binding arbitration** rather than in court, except as provided below.
{% endhint %}

### 12.3 Class Action Waiver

You agree that any dispute resolution proceedings will be conducted only on an individual basis and not in a class, consolidated, or representative action. You waive any right to participate in a class action lawsuit or class-wide arbitration.

### 12.4 Exceptions

Nothing in this Section prevents either party from seeking injunctive or other equitable relief in a court of competent jurisdiction where necessary to prevent irreparable harm or protect intellectual property rights.

### 12.5 Governing Law

These Terms shall be governed by and construed in accordance with applicable general principles of commercial law. The specific governing jurisdiction shall be determined and updated in a subsequent version of these Terms. Until such determination, disputes shall be resolved in accordance with internationally recognized arbitration principles.

***

## 13. Force Majeure

### 13.1 No Liability for Force Majeure Events

Clore.ai shall not be liable for any failure or delay in performance of its obligations under these Terms to the extent caused by events beyond its reasonable control, including but not limited to:

* Acts of God, natural disasters, or extreme weather events
* War, terrorism, civil unrest, or government actions
* Blockchain network failures, protocol changes, or attacks
* Power or internet outages
* Regulatory changes, sanctions, or embargoes
* Pandemic, epidemic, or public health emergencies
* Cyber attacks, hacking, or data breaches affecting critical infrastructure
* Failure of third-party service providers

### 13.2 Notice and Mitigation

In the event of a force majeure event, Clore.ai will use commercially reasonable efforts to notify affected users and to resume normal operations as soon as reasonably practicable.

***

## 14. Miscellaneous

### 14.1 Entire Agreement

These Terms, together with the Privacy Policy and any other policies or guidelines published by Clore.ai, constitute the entire agreement between you and Clore.ai with respect to the subject matter hereof.

### 14.2 Severability

If any provision of these Terms is found to be invalid, illegal, or unenforceable, the remaining provisions shall continue in full force and effect.

### 14.3 Waiver

Clore.ai's failure to enforce any provision of these Terms shall not be deemed a waiver of its right to enforce such provision in the future.

### 14.4 Assignment

You may not assign or transfer any of your rights or obligations under these Terms without the prior written consent of Clore.ai. Clore.ai may freely assign these Terms without restriction.

### 14.5 No Third-Party Beneficiaries

These Terms do not create any third-party beneficiary rights.

### 14.6 Contact

For questions regarding these Terms, please contact us through the official channels available on [clore.ai](https://clore.ai).

***

*These Terms of Service were last updated in March 2026 and supersede all prior versions.*


# Privacy Policy

**Last Updated: March 2026**

{% hint style="info" %}
This Privacy Policy explains how Clore.ai ("Clore.ai," "the Platform," "we," "us," or "our") collects, uses, stores, and shares information about you when you use our website at [clore.ai](https://clore.ai) and related services. Please read this policy carefully.
{% endhint %}

***

## 1. Introduction

### 1.1 Our Commitment

Clore.ai is committed to protecting your privacy and handling your personal data with transparency and care. This Privacy Policy describes our practices regarding the collection, use, disclosure, and protection of your information.

### 1.2 Scope

This Privacy Policy applies to all users of the Clore.ai platform, including Hosts, Renters, and visitors. It covers data collected through our website, APIs, and any related services.

### 1.3 Acceptance

By using the Clore.ai platform, you acknowledge that you have read and understood this Privacy Policy and consent to the data practices described herein. If you do not agree with this policy, please discontinue use of the Platform.

***

## 2. Data We Collect

### 2.1 Account Information

When you register for a Clore.ai account, we collect:

* **Email address** — used for account authentication and communications
* **Username or display name**
* **Password** (stored in hashed/encrypted form)
* **Account creation date and activity timestamps**

### 2.2 Cryptocurrency and Wallet Data

* **Blockchain wallet addresses** used for sending and receiving payments (CLORE, BTC, USDT, USDC)
* **Transaction history** associated with your account on the Platform
* **Payment amounts and timestamps**

{% hint style="info" %}
Wallet addresses and on-chain transactions are publicly visible on their respective blockchains. Clore.ai does not control the public nature of blockchain data.
{% endhint %}

### 2.3 Technical and Usage Data

We automatically collect the following when you use the Platform:

* **IP address** and approximate geolocation derived therefrom
* **Device type, operating system, and browser information**
* **Pages visited, features used, and click-through data**
* **Referral URLs and traffic sources**
* **Session duration and interaction patterns**
* **Error logs and performance data**

### 2.4 Communication Data

* **Support tickets and correspondence** submitted to Clore.ai
* **Feedback and survey responses**
* **Notifications and messages** sent through the Platform

### 2.5 KYC/AML Data (If Required)

If identity verification is required under our KYC/AML procedures:

* Government-issued identification documents
* Proof of address
* Selfies or liveness check data
* Any other documentation required for regulatory compliance

### 2.6 Cookies and Tracking Data

See Section 6 for full details on our use of cookies and similar tracking technologies.

***

## 3. How We Use Your Data

### 3.1 Account Management

* Creating and maintaining your user account
* Authenticating your identity upon login
* Enabling you to list, rent, or manage GPU resources
* Processing and recording transactions

### 3.2 Platform Operations

* Facilitating transactions between Hosts and Renters
* Displaying marketplace listings and pricing
* Processing payments and commissions
* Sending transaction confirmations and receipts

### 3.3 Security and Fraud Prevention

* Detecting, investigating, and preventing fraudulent activity
* Monitoring for prohibited conduct as described in our Terms of Service
* Enforcing KYC/AML requirements
* Protecting the integrity of the Platform and its users
* Responding to security incidents

### 3.4 Analytics and Improvement

* Analyzing usage patterns to improve the Platform's features and performance
* Conducting internal research and development
* Generating aggregated, anonymized statistics about Platform usage

### 3.5 Communications

* Sending service-related emails (account notifications, security alerts, policy updates)
* Providing customer support
* Notifying you of changes to terms, policies, or Platform features
* Sending marketing communications where you have consented

### 3.6 Legal Compliance

* Complying with applicable laws and regulations
* Responding to lawful requests from law enforcement or regulatory authorities
* Enforcing our Terms of Service and other agreements

***

## 4. Data Sharing and Disclosure

### 4.1 General Principle

{% hint style="info" %}
Clore.ai does **not** sell your personal data to third parties. We share data only as described in this Section.
{% endhint %}

### 4.2 Law Enforcement and Legal Requests

We may disclose your information to government authorities, law enforcement agencies, or regulatory bodies when:

* Required to do so by applicable law, court order, or legal process
* Responding to a lawful subpoena, warrant, or official government request
* Necessary to protect the rights, property, or safety of Clore.ai, our users, or the public
* Required for KYC/AML compliance or to prevent financial crime

### 4.3 Analytics and Service Providers

We may share data with third-party service providers who assist us in operating the Platform, subject to appropriate data processing agreements, including:

* **Analytics providers** — for understanding Platform usage (e.g., aggregated behavioral data)
* **Infrastructure providers** — for hosting and technical operations
* **KYC/AML verification providers** — for identity verification services
* **Security and fraud detection services**

These providers are contractually bound to use your data only for the purposes for which it was shared and to maintain appropriate security standards.

### 4.4 Business Transfers

In the event of a merger, acquisition, sale of assets, or other business restructuring, your data may be transferred as part of that transaction. We will provide notice of any such transfer and any material changes to this Privacy Policy.

### 4.5 Public Blockchain Data

Transaction data recorded on public blockchains (Ethereum, Bitcoin, etc.) is inherently public and is not subject to this Privacy Policy.

### 4.6 Aggregated and Anonymized Data

We may share aggregated or de-identified data that cannot reasonably be used to identify you with third parties for research, analytics, or marketing purposes.

***

## 5. Data Retention

### 5.1 Retention Period

We retain your personal data for as long as your account is active and for a reasonable period thereafter, typically no less than the following:

| Data Type              | Retention Period                                   |
| ---------------------- | -------------------------------------------------- |
| Account information    | Duration of account + 3 years                      |
| Transaction records    | Minimum 5 years (legal/tax requirements)           |
| IP and usage logs      | 12 months                                          |
| Support communications | 3 years                                            |
| KYC/AML documents      | As required by applicable law (typically 5+ years) |
| Cookie data            | Per cookie type (see Section 6)                    |

### 5.2 Legal Obligations

Notwithstanding your account status, we may retain data for longer periods where required by applicable law, regulation, or legitimate legal proceedings.

### 5.3 Deletion After Retention Period

Upon expiry of the applicable retention period, data will be deleted or anonymized in accordance with our internal data management procedures.

***

## 6. Cookies and Tracking Technologies

### 6.1 What Are Cookies

Cookies are small text files stored on your device that help us recognize you, maintain your session, and understand how you use the Platform. We also use similar technologies such as web beacons and local storage.

### 6.2 Types of Cookies We Use

| Cookie Type             | Purpose                                      | Duration          |
| ----------------------- | -------------------------------------------- | ----------------- |
| **Essential / Session** | Authentication, security, session management | Session / 30 days |
| **Functional**          | User preferences, language settings          | 1 year            |
| **Analytics**           | Usage statistics, performance monitoring     | 90 days           |
| **Security**            | Fraud detection, bot prevention              | Session           |

### 6.3 Third-Party Cookies

We may use third-party analytics services (such as Google Analytics or similar) that set their own cookies. These third parties have their own privacy policies governing their use of such data.

### 6.4 Cookie Control

{% hint style="info" %}
You can control or disable cookies through your browser settings. Note that disabling essential cookies may impair your ability to use certain features of the Platform, including account login.
{% endhint %}

### 6.5 Do Not Track

We currently do not respond to "Do Not Track" browser signals. We will update this policy if our practices change.

***

## 7. Third-Party Services

### 7.1 Payment Processors

Cryptocurrency transactions are processed on their respective blockchains (Ethereum, Bitcoin). We may integrate third-party payment infrastructure services for certain payment flows. These services operate under their own privacy policies.

### 7.2 Analytics Providers

We use analytics tools to understand Platform performance. Analytics data is typically collected in aggregate or pseudonymous form.

### 7.3 KYC/AML Providers

Where identity verification is required, we may share your personal data with licensed KYC/AML verification providers. These providers are required to process your data only for verification purposes and in compliance with applicable law.

### 7.4 Third-Party Links

The Platform may contain links to third-party websites or services. Clore.ai is not responsible for the privacy practices of those third parties. We encourage you to review their privacy policies before providing any personal information.

***

## 8. Data Security

### 8.1 Security Measures

{% hint style="info" %}
We implement industry-standard technical and organizational security measures to protect your data against unauthorized access, alteration, disclosure, or destruction.
{% endhint %}

These measures include:

* **Encryption in transit** — All data transmitted between your browser and our servers is encrypted using TLS/HTTPS
* **Encryption at rest** — Sensitive data stored in our systems is encrypted
* **Access controls** — Strict role-based access controls limit employee access to personal data
* **Password hashing** — Passwords are never stored in plaintext; they are hashed using industry-standard algorithms
* **Security monitoring** — We maintain systems to detect and respond to security incidents
* **Regular audits** — We conduct periodic security reviews of our infrastructure and practices

### 8.2 No Absolute Guarantee

Despite our security measures, no system is completely impenetrable. We cannot guarantee the absolute security of your data. In the event of a data breach that affects your rights, we will notify you and relevant authorities as required by applicable law.

### 8.3 Your Responsibility

You are responsible for maintaining the security of your account credentials, including your password and private keys. Do not share your account credentials with anyone.

***

## 9. Your Rights and Choices

### 9.1 Access

You have the right to request access to the personal data we hold about you, including a summary of the categories of data collected and how it is used.

### 9.2 Correction

You have the right to request correction of any inaccurate or incomplete personal data we hold about you.

### 9.3 Deletion (Right to Be Forgotten)

You may request deletion of your personal data. We will fulfill such requests subject to:

* Our legal obligations to retain certain data (e.g., transaction records for tax/AML purposes)
* Legitimate business interests in retaining data for fraud prevention or legal proceedings
* The inherently permanent nature of blockchain transaction records

### 9.4 Data Portability

You may request an export of your personal data in a structured, commonly used, machine-readable format where technically feasible.

### 9.5 Objection and Restriction

You may object to or request restriction of certain processing of your personal data, particularly where we process data based on legitimate interests.

### 9.6 Marketing Opt-Out

You may opt out of marketing communications at any time by following the unsubscribe link in any email or by contacting us directly. Opting out of marketing does not affect service-related communications.

### 9.7 How to Exercise Your Rights

To exercise any of the above rights, please contact us through the official support channels available on [clore.ai](https://clore.ai). We will respond to verified requests within the timeframe required by applicable law.

***

## 10. GDPR — Additional Rights for EU/EEA Users

{% hint style="info" %}
If you are located in the European Union (EU) or European Economic Area (EEA), the following additional provisions apply under the General Data Protection Regulation (GDPR).
{% endhint %}

### 10.1 Legal Basis for Processing

We process your personal data on one or more of the following legal bases:

| Processing Activity             | Legal Basis                         |
| ------------------------------- | ----------------------------------- |
| Account creation and management | Contract performance (Art. 6(1)(b)) |
| Transaction processing          | Contract performance (Art. 6(1)(b)) |
| Security and fraud prevention   | Legitimate interests (Art. 6(1)(f)) |
| Legal compliance (KYC/AML)      | Legal obligation (Art. 6(1)(c))     |
| Analytics and improvement       | Legitimate interests (Art. 6(1)(f)) |
| Marketing communications        | Consent (Art. 6(1)(a))              |

### 10.2 Data Controller

For the purposes of the GDPR, Clore.ai acts as the data controller for personal data collected through the Platform.

### 10.3 International Transfers

Your data may be transferred to and processed in countries outside the EU/EEA that may not provide the same level of data protection as your home country. Where such transfers occur, we will implement appropriate safeguards such as Standard Contractual Clauses (SCCs) or equivalent mechanisms.

### 10.4 Right to Lodge a Complaint

EU/EEA users have the right to lodge a complaint with the data protection supervisory authority in their country of residence if they believe our processing of their data violates the GDPR.

### 10.5 Data Protection Officer

If required by applicable law, Clore.ai will appoint a Data Protection Officer (DPO). Contact information for the DPO, if applicable, will be made available on the Platform.

***

## 11. Children's Privacy

### 11.1 No Services for Minors

Clore.ai does not knowingly collect personal data from individuals under the age of 18. If we become aware that we have inadvertently collected data from a minor, we will take prompt steps to delete such data.

### 11.2 Parental Reporting

If you believe that a minor has provided us with personal data without appropriate consent, please contact us immediately through the channels available on [clore.ai](https://clore.ai).

***

## 12. Changes to This Privacy Policy

### 12.1 Right to Modify

Clore.ai reserves the right to update, modify, or replace this Privacy Policy at any time. We will use commercially reasonable efforts to notify you of material changes through:

* **Email notification** to the address associated with your account
* **Prominent notice on the Platform** prior to the change taking effect

### 12.2 Continued Use as Acceptance

Your continued use of the Platform following the effective date of any updated Privacy Policy constitutes your acceptance of the revised terms. If you do not agree to the updated policy, you must discontinue using the Platform.

***

## 13. Contact Us

For questions, concerns, or requests regarding this Privacy Policy or your personal data, please contact us through the official support channels available at [clore.ai](https://clore.ai).

***

*This Privacy Policy was last updated in March 2026 and supersedes all prior versions.*


# Introduction

> **194 practical guides** for deploying AI models, GPU workloads, and AI platforms on [Clore.ai](https://clore.ai) — the decentralized GPU rental marketplace.

{% hint style="success" %}
All examples can be run on GPU servers rented through the [Clore.ai Marketplace](https://clore.ai/marketplace). Rent powerful GPUs starting from **$0.15/day**.
{% endhint %}

## What is Clore.ai?

[Clore.ai](https://clore.ai) is a peer-to-peer GPU marketplace where you rent GPUs directly from other people — like Airbnb for compute. Thousands of GPUs are available 24/7, from budget RTX 3060s to enterprise H100s. Pay with **CLORE**, **BTC**, **USDT**, or **USDC**.

### Why Clore.ai for AI?

* **Affordable** — RTX 4090 from $0.50/day (vs $2–4 on cloud providers)
* **No commitments** — rent by the hour, no contracts
* **Full root access** — Docker containers with GPU passthrough
* **Wide GPU selection** — 3,400+ machines, 12,800+ GPUs online
* **Pay your way** — crypto payments (CLORE, BTC, USDT/USDC)

## 📚 Guide Categories

| Category                                                                 | Guides | Highlights                                                                                                                                                                              |
| ------------------------------------------------------------------------ | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🤖 [Language Models](/guides/language-models/language-models)            | **45** | DeepSeek V4, MiMo-V2.5-Pro, Hy3 Preview, Qwen3.6-27B, MiniMax M2.7, Ling-2.6-flash, GLM-5.1, Llama 4, Gemma 3, Qwen3.5-Omni, GLM-5, DeepSeek-R1, Nemotron 3 Super, Ollama, vLLM, SGLang |
| 🤖 [AI Platforms & Agents](/guides/ai-platforms-and-agents/ai-platforms) | **18** | Dify, CrewAI, AutoGPT, OpenHands, MetaGPT, n8n, LibreChat, Open Interpreter, SuperAGI, SWE-agent                                                                                        |
| 🔊 [Audio & Voice](/guides/audio-and-voice/audio-voice)                  | **23** | MOSS-TTS, Voxtral TTS, Whisper, Qwen3-TTS, MiniMax Speech 2.6, Dia, ChatTTS, Kokoro, Fish Speech, MeloTTS, StyleTTS2                                                                    |
| 🎬 [Video Generation](/guides/video-generation/video-generation)         | **14** | Wan2.1, Wan 2.2 VBVR, FramePack, CubeComposer 4K 360°, LTX-2, CogVideoX, SkyReels, HunyuanVideo, Mochi-1, AnimateDiff                                                                   |
| 🎨 [Image Generation](/guides/image-generation/image-generation)         | **11** | FLUX.2 Klein, HunyuanImage 3.0, SD 3.5, ComfyUI, InvokeAI, SD WebUI Forge                                                                                                               |
| 🧠 [Training](/guides/training/training)                                 | **12** | Unsloth, Axolotl, LoRA, DreamBooth, DeepSpeed, LLaMA-Factory, TRL, LitGPT, Mergekit                                                                                                     |
| 🏁 [Getting Started](/guides/getting-started/getting-started)            | **8**  | GPU comparison, pricing, FAQ, troubleshooting                                                                                                                                           |
| 👁️ [Vision Models](/guides/vision-models/vision-models)                 | **6**  | Qwen2.5-VL, SAM2, LLaVA, Florence-2                                                                                                                                                     |
| 🖼️ [Image Processing](/guides/image-processing/image-processing)        | **6**  | Real-ESRGAN, ControlNet, Depth Anything, ICLight                                                                                                                                        |
| 🧊 [3D Generation](/guides/3d-generation/3d-generation)                  | **6**  | Hunyuan World 2.0, TRELLIS, Hunyuan3D 2.1, TripoSR, Gaussian Splatting, Nerfstudio                                                                                                      |
| 🎭 [Talking Heads](/guides/talking-heads/talking-heads)                  | **3**  | LivePortrait, SadTalker, Wav2Lip                                                                                                                                                        |
| 👤 [Face & Identity](/guides/face-and-identity/face-identity)            | **3**  | FaceFusion, InstantID, IP-Adapter                                                                                                                                                       |
| ⚙️ [Advanced](/guides/advanced/advanced)                                 | **5**  | Multi-GPU, API integration, batch processing                                                                                                                                            |
| 🔧 [Other Workloads](/guides/other-workloads/other-workloads)            | **3**  | Blender, Kandinsky, OpenClaw                                                                                                                                                            |
| 💻 [AI Coding Tools](/guides/ai-coding-tools/ai-coding)                  | **2**  | Aider, TabbyML (self-hosted Copilot)                                                                                                                                                    |
| 📊 [Comparisons](/guides/comparisons/comparisons)                        | **7**  | LLM serving, fine-tuning, video gen, TTS, RAG frameworks, vector DBs                                                                                                                    |
| 📹 [Video Processing](/guides/video-processing/video-processing)         | **2**  | FFmpeg NVENC, RIFE interpolation                                                                                                                                                        |
| 🎵 [Music Generation](/guides/music-generation/music-generation)         | **1**  | ACE-Step (open-source Suno alternative)                                                                                                                                                 |
| 🔍 [Computer Vision](/guides/computer-vision/computer-vision)            | **2**  | YOLOv9/v10 detection                                                                                                                                                                    |
| 🗄️ [RAG / Vector DBs](/guides/rag-and-vector-databases/rag-vectordb)    | **6**  | LlamaIndex, RAGFlow, ChromaDB, Qdrant, Milvus, Weaviate                                                                                                                                 |
| 🔄 [MLOps](/guides/mlops-and-deployment/mlops)                           | **4**  | MLflow, Triton Inference Server, BentoML, ClearML                                                                                                                                       |
| ⚡ [DevOps GPU](/guides/gpu-devops/devops-gpu)                            | **2**  | TensorRT-LLM, ONNX Runtime                                                                                                                                                              |
| 🔬 [Science](/guides/science-and-research/science)                       | **3**  | AlphaFold2, ESMFold, GROMACS molecular dynamics                                                                                                                                         |
| 🎮 [Gaming / Streaming](/guides/gaming-and-streaming/gaming-streaming)   | **1**  | Sunshine + Moonlight remote gaming                                                                                                                                                      |
| ₿ [Crypto Mining](/guides/crypto-and-mining/crypto-mining)               | **1**  | XMRig CPU/GPU mining                                                                                                                                                                    |

## 🔥 What's New (April 2026)

### Week of April 27, 2026 — 6 New Guides

* 🤖 **DeepSeek V4** 🆕 — The frontier MoE finally shipped April 22 with full open weights under MIT. Pro: 1.6T total / 49B active, 1M context. Flash: 284B / 13B active, runs quantized on a single 80GB GPU. Day-0 vLLM and SGLang support — [guide](/guides/language-models/deepseek-v4)
* 🤖 **MiMo-V2.5-Pro** 🆕 — Xiaomi's first open-weight Pro tier (April 27). 1.02T MoE / 42B active, 1M context, MIT, FP8 native. Hybrid attention 6:1 + 3-step MTP — [guide](/guides/language-models/mimo-v25-pro)
* 🤖 **Hy3 Preview** 🆕 — Tencent Hunyuan 3 (April 23). 295B MoE / 21B active, 256K context, agent + reasoning focus. Single-node 4× A100 80GB serves it FP8 — [guide](/guides/language-models/hy3-preview)
* 🤖 **Ling-2.6-flash** 🆕 — Ant Group inclusionAI (April 28). 104B MoE / 7.4B active, agent-tuned. Fits on a single RTX 4090 INT4 or single H100 FP8 — [guide](/guides/language-models/ling-26-flash)
* 🤖 **Qwen3.6-27B** 🆕 — Alibaba's dense 27B (April 21). Apache 2.0, 262K context (extensible to 1M with YaRN). The "one card that just works" coding LLM — [guide](/guides/language-models/qwen36-27b)
* 🤖 **MiniMax M2.7** 🆕 — 229B MoE coding model (April 9). Custom MiniMax license, FP8 native. Previously listed as proprietary — open weights are now public — [guide](/guides/language-models/minimax-m27)

#### Industry Notes (April 21–29, 2026)

*Closed-source or API-only this window — self-host alternatives linked.*

* **Claude Opus 4.7** (Anthropic, April 16) — closed weights. → Self-host alternatives: [DeepSeek V4](/guides/language-models/deepseek-v4), [MiMo-V2.5-Pro](/guides/language-models/mimo-v25-pro), [GLM-5.1](/guides/language-models/glm-5-1)
* **GPT-5.5 / GPT-5.5 Pro** (OpenAI, April 23) — closed, API-only. → Self-host alternatives: [DeepSeek V4](/guides/language-models/deepseek-v4), [MiMo-V2.5-Pro](/guides/language-models/mimo-v25-pro)
* **Wan 2.7** (Alibaba) — open weights still pending Q2 2026. Wan 2.5 broke tradition and stayed closed; treat 2.7 with skepticism. → Self-host today: [Wan 2.2 VBVR](/guides/video-generation/wan22-vbvr), [Wan2.1](/guides/video-generation/wan-video), [LTX-2](/guides/video-generation/ltx-video-2)
* **Qwen3.6-Plus** (Alibaba) — 1M-context coding agent, API-only on OpenRouter. → Self-host alternative: [Qwen3.6-27B](/guides/language-models/qwen36-27b), [Qwen3.5-Omni](/guides/language-models/qwen35-omni)
* **Hugging Face ml-intern** (April 21) — open-source autonomous post-training agent (research → dataset → train). Watching adoption before adding a full guide. → Use today: [Unsloth](/guides/training/unsloth-finetune), [LLaMA-Factory](/guides/training/llama-factory), [Axolotl](/guides/training/axolotl-training)

### Week of April 20, 2026 — 3 New Guides

* 🧊 **Hunyuan World 2.0** 🆕 — Tencent's first open-source 3D world model. Text/image/video → editable mesh + 3D Gaussian Splatting for Unity/Unreal/Blender. \~1.2B params BF16, 12–24 GB VRAM — [guide](/guides/3d-generation/hunyuan-world-2)
* 🤖 **GLM-5.1** 🆕 — Zhipu AI's 744B MoE / 40B active coding powerhouse. 200K context, MIT license, #1 on SWE-Bench Pro (58.4%) beating Claude Opus 4.6 and GPT-5.4 — [guide](/guides/language-models/glm-5-1)
* 🔊 **MOSS-TTS** 🆕 — OpenMOSS ultra-lightweight 100M-param TTS. 48kHz stereo, 20 languages, CPU-only inference (no GPU required), zero-shot voice cloning, Apache 2.0 — [guide](/guides/audio-and-voice/moss-tts)

#### Industry Notes (April 13–20, 2026)

*Closed-source or API-only launches this week — not self-hostable on Clore yet. For each, we've linked the closest open-weight alternative you can run on Clore today.*

* **Alibaba Happy Oyster** — interactive real-time 3D world model (Roaming + Directing modes). Waitlisted / closed beta. → Self-host alternative: [Hunyuan World 2.0](/guides/3d-generation/hunyuan-world-2)
* **ByteDance Seedance 2.0 API** — fully opened April 14 on Volcano Engine, API-only. → Self-host alternatives: [Wan 2.2 VBVR](/guides/video-generation/wan22-vbvr), [LTX-2](/guides/video-generation/ltx-video-2)
* **Qwen 3.6-Plus** — Alibaba's proprietary 1M-context coding agent. API-only. → Self-host alternatives: [Qwen3.5-Omni](/guides/language-models/qwen35-omni), [GLM-5.1](/guides/language-models/glm-5-1)
* **Meta Muse Spark** — Meta's first post-Llama proprietary model. → Self-host alternative: [Llama 4](/guides/language-models/llama4)
* **Wan 2.7** — open weights still pending Q2 2026 (currently API-only via Together AI). → Self-host today: [Wan 2.2 VBVR](/guides/video-generation/wan22-vbvr)

### Week of April 13, 2026 — 2 New Guides

* 🤖 **Qwen3.5-Omni** 🆕 — Alibaba's unified multimodal model: text, audio, image & video understanding + text & speech generation in one 30B MoE. INT4 fits in 24GB VRAM. Apache 2.0 — [guide](/guides/language-models/qwen35-omni)
* 🎬 **Wan 2.2 VBVR** 🆕 — Video-Based Video Reference for consistent motion control. Use a reference clip to drive animation in new video generation. FP8 runs on RTX 3090 (16-24GB). ComfyUI workflow included — [guide](/guides/video-generation/wan22-vbvr)

#### Industry Notes (April 6–13, 2026)

*API-only or pending open weights — self-host alternatives linked.*

* **Wan 2.7** (Alibaba) — First+last-frame generation, video-to-video editing, subject referencing. Open weights expected mid-Q2 2026. → Self-host today: [Wan 2.2 VBVR](/guides/video-generation/wan22-vbvr), [Wan2.1](/guides/video-generation/wan-video)
* **Qwen3.6-Plus** (Alibaba) — 1M context window, agentic coding. Open-weight release pending. → Self-host alternatives: [Qwen3.5](/guides/language-models/qwen35), [GLM-5.1](/guides/language-models/glm-5-1)
* **LTX 2.3** — New LTX video generation model, 22B params, 32GB VRAM baseline. → Use current open release: [LTX-2](/guides/video-generation/ltx-video-2)

## 🔥 What's New (March 2026)

### Week of March 30, 2026 — 1 New Guide

* 🔊 **Voxtral TTS** 🆕 — Mistral's open-weight 4B TTS model, 9 languages, zero-shot voice cloning from 3s reference, only 3 GB VRAM, Apache 2.0 — [guide](/guides/audio-and-voice/voxtral-tts)

#### Industry Notes (March 23–30)

*API-only / shut down — self-host alternatives linked.*

* **Seedance 2.0** (ByteDance) — next-gen video gen launched in CapCut/Dreamina. API-only. → Self-host alternatives: [Wan2.1](/guides/video-generation/wan-video), [LTX-2](/guides/video-generation/ltx-video-2)
* **Google Lyria 3 Pro** — music generation via Gemini API/Vertex AI. API-only. → Self-host alternative: [ACE-Step](/guides/music-generation/ace-step)
* **Gemini 3 Deep Think** — Google's frontier reasoning model, API-only. → Self-host alternatives: [DeepSeek-R1](/guides/language-models/deepseek-r1), [GLM-5.1](/guides/language-models/glm-5-1)
* **Sora shutdown** — OpenAI officially shut down the Sora video generation app. → Self-host alternative: [OpenSora](/guides/video-generation/opensora)
* **DeepSeek V4** ✅ — shipped April 22, 2026 with full open weights (MIT). See [updated guide](/guides/language-models/deepseek-v4).

### Week of March 16, 2026 — 3 New Guides

* 🤖 **NVIDIA Nemotron 3 Super** 🆕 — 120B MoE / 12B active, 5× throughput, 1M context, Apache 2.0, built for agentic AI — [guide](/guides/language-models/nvidia-nemotron-3-super)
* 🌐 **Gemini 3.1 Flash Lite** 🆕 — Google's cheapest/fastest model (March 3, 2026), API + open-source alternatives — [guide](/guides/language-models/gemini-3-1-flash-lite)
* 🎬 **CubeComposer 4K 360° Video** 🆕 — first model to natively generate 4K 360° panoramic video from standard footage (CVPR 2026) — [guide](/guides/video-generation/cubecomposer-360-video)

### Also This Week (March 9–16, 2026)

* **GPT-5.4** — released March 5, 2026; native computer use (75.0% OSWorld), 1M context. API-only, no local deployment. → Self-host alternatives: [GLM-5.1](/guides/language-models/glm-5-1), [DeepSeek-R1](/guides/language-models/deepseek-r1)
* **DeepSeek V4** ✅ — released April 22, 2026, MIT, 1.6T-Pro / 284B-Flash. [Guide](/guides/language-models/deepseek-v4).
* **Wan 2.2** — new version of Wan video generation foundation model. → Self-host: [Wan 2.2 VBVR](/guides/video-generation/wan22-vbvr), [Wan2.1](/guides/video-generation/wan-video)

### New in March 2026 — 6 New Categories, 57 New Guides

* 🗄️ **RAG / Vector DBs** — 6 guides: [LlamaIndex](/guides/rag-and-vector-databases/llamaindex), [RAGFlow](/guides/rag-and-vector-databases/ragflow), [ChromaDB](/guides/rag-and-vector-databases/chromadb), [Qdrant](/guides/rag-and-vector-databases/qdrant), [Milvus](/guides/rag-and-vector-databases/milvus), [Weaviate](/guides/rag-and-vector-databases/weaviate)
* 🔄 **MLOps** — 4 guides: [MLflow](/guides/mlops-and-deployment/mlflow), [Triton Inference Server](/guides/mlops-and-deployment/triton-inference-server), [BentoML](/guides/mlops-and-deployment/bentoml), [ClearML](/guides/mlops-and-deployment/clearml)
* ⚡ **DevOps GPU** — 2 guides: [TensorRT-LLM](/guides/gpu-devops/tensorrt-llm), [ONNX Runtime](/guides/gpu-devops/onnx-runtime)
* 🔬 **Science** — 3 guides: [AlphaFold2](/guides/science-and-research/alphafold2), [ESMFold](/guides/science-and-research/esmfold), [GROMACS](/guides/science-and-research/gromacs) molecular dynamics
* 🎮 **Gaming / Streaming** — [Sunshine + Moonlight](/guides/gaming-and-streaming/sunshine-moonlight) GPU-accelerated remote gaming
* ₿ **Crypto Mining** — [XMRig](/guides/crypto-and-mining/xmrig) CPU/GPU mining on Clore.ai

### Latest Models Added (March 4, 2026)

* **DeepSeek V4** ✅ — 1.6T-Pro / 284B-Flash MoE, MIT license, **shipped April 22, 2026** — [guide](/guides/language-models/deepseek-v4)
* **MiniMax Speech 2.6** 🆕 — ultra-low latency TTS for voice agents, < 300ms TTFB — [guide](/guides/audio-and-voice/minimax-speech)
* **SGLang** 🆕 — RadixAttention for KV cache sharing, 2–5× throughput vs vLLM on MoE — [guide](/guides/language-models/sglang)
* **TGI** 🆕 — HuggingFace's production LLM serving with Flash Attention 2 + PagedAttention — [guide](/guides/language-models/tgi)
* **LLaMA-Factory** 🆕 — fine-tune 100+ LLMs with WebUI, LoRA/QLoRA, RLHF — [guide](/guides/training/llama-factory)
* **Fish Speech** 🆕 — zero-shot voice cloning in 8+ languages from 10–15s reference audio — [guide](/guides/audio-and-voice/fish-speech)
* **Mochi-1** 🆕 — 10B parameter video diffusion, 848×480 @ 30fps, 24GB VRAM — [guide](/guides/video-generation/mochi-1)

### Previously Added (February 2026)

* **Qwen3.5** — Alibaba's 397B MoE, beat Claude 4.5 Opus on math
* **GLM-5** — 744B MoE from Zhipu AI, MIT license, #1 in open-source rankings
* **Ling-2.5-1T** — Ant Group's trillion-parameter model with linear attention
* **Kimi K2.5** — Moonshot AI's 1T MoE, MIT license, visual agentic
* **Mistral Large 3** — 675B MoE, Apache 2.0, frontier coding & reasoning
* **Llama 4 Scout/Maverick** — Meta's MoE revolution, 10M context window
* **Gemma 3** — Google's 27B that beats 405B models
* **FLUX.2 Klein** — sub-second image generation (< 0.5s on RTX 4090)
* **HunyuanImage 3.0** — 80B MoE, largest open-source image model
* **ACE-Step 1.5** — full song generation on < 4GB VRAM
* **FramePack** — AI video with just 6GB VRAM
* **Qwen3-TTS** — voice cloning in 10+ languages from 3 seconds of audio
* **Kani-TTS-2** — ultra-lightweight TTS, only 3GB VRAM
* **DeepSeek-R1** — reasoning model matching OpenAI o1

### Previously Added Categories (February 2026)

* 🤖 **AI Platforms & Agents** — 18 guides: Dify, CrewAI, AutoGPT, OpenHands, MetaGPT, n8n, LibreChat, Open Interpreter, SuperAGI, SWE-agent
* 🎵 **Music Generation** — AI-composed songs with vocals (ACE-Step)
* 💻 **AI Coding Tools** — self-hosted Copilot alternatives (Aider, TabbyML)
* 🔧 **OpenClaw on Clore** — run your AI assistant 24/7 on rented GPUs

## 💰 GPU Pricing (April 2026)

> Snapshot of typical Clore.ai marketplace ranges, late-April 2026. See [clore.ai/marketplace](https://clore.ai/marketplace) for live numbers — spot orders typically run 20–40% cheaper than on-demand.

| GPU         | VRAM  | Typical /hr | Best For                                           | Landing Page                                                                      |
| ----------- | ----- | ----------- | -------------------------------------------------- | --------------------------------------------------------------------------------- |
| RTX 3060    | 12GB  | $0.16–1.00  | TTS, small models, music gen                       | —                                                                                 |
| RTX 3070    | 8GB   | $0.17–3.33  | 7B models, Whisper, batch inference                | [Rent](https://clore.ai/rent-3070.html) / [Host](https://clore.ai/host-3070.html) |
| RTX 3080    | 10GB  | $0.20–3.50  | 7B–14B models, image gen                           | [Rent](https://clore.ai/rent-3080.html) / [Host](https://clore.ai/host-3080.html) |
| RTX 3090    | 24GB  | $0.33–4.00  | SDXL, 32B Q4, Ling-2.6-flash INT4                  | [Rent](https://clore.ai/rent-3090.html) / [Host](https://clore.ai/host-3090.html) |
| RTX 4070 Ti | 12GB  | $0.30–2.00  | SDXL, ComfyUI, 7B models                           | [Rent](https://clore.ai/rent-4070-ti.html)                                        |
| RTX 4080    | 16GB  | $0.45–3.50  | FLUX, 14B models, fine-tuning                      | [Rent](https://clore.ai/rent-4080.html)                                           |
| RTX 4090    | 24GB  | $0.70–4.50  | FLUX, Llama 4, Qwen3.6-27B Q4, Ling-2.6-flash      | [Rent](https://clore.ai/rent-4090.html) / [Host](https://clore.ai/host-4090.html) |
| RTX 5080    | 16GB  | $0.90–9.00  | Fast FLUX, 30B LLMs, Blackwell FP4                 | [Rent](https://clore.ai/rent-5080.html)                                           |
| RTX 5090    | 32GB  | $1.72–10.00 | 70B quantized, Qwen3.6-27B Q8, Nemotron 3 Super    | [Rent](https://clore.ai/rent-5090.html) / [Host](https://clore.ai/host-5090.html) |
| RTX A6000   | 48GB  | $1.50–4.00  | 70B BF16, Qwen3.6-27B BF16                         | [Rent](https://clore.ai/rent-a6000.html)                                          |
| L40S        | 48GB  | $2.00–9.00  | FP8 datacenter inference, single-GPU 27B           | [Rent](https://clore.ai/rent-l40s.html)                                           |
| A100 80GB   | 80GB  | $2.00–6.00  | Hy3 Preview FP8, MiniMax M2.7, training            | [Rent](https://clore.ai/rent-a100-80gb.html)                                      |
| H100 80GB   | 80GB  | $4.00–9.00  | DeepSeek V4 Pro, MiMo-V2.5-Pro, frontier inference | [Rent](https://clore.ai/rent-h100.html)                                           |
| H200 141GB  | 141GB | $6.00–14.00 | DeepSeek V4 single-node, full BF16                 | [Rent](https://clore.ai/rent-h200.html)                                           |
| B200 192GB  | 192GB | $8.00–18.00 | Frontier training, MiMo-V2.5-Pro 8× cluster        | [Rent](https://clore.ai/rent-b200.html)                                           |

*Looking to monetize idle GPUs? See the* [*host pages*](https://clore.ai/llms.txt) *for monthly earnings estimates per card type.*

## 🚀 Quick Start

New to Clore.ai? → [**Quickstart Guide**](/guides/quickstart)

Already know what you need?

| I want to...                  | Start here                                                                                                                                                |
| ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chat with AI locally          | [Ollama Guide](/guides/language-models/ollama)                                                                                                            |
| Run a frontier coding LLM     | [DeepSeek V4](/guides/language-models/deepseek-v4), [GLM-5.1](/guides/language-models/glm-5-1), or [MiMo-V2.5-Pro](/guides/language-models/mimo-v25-pro)  |
| Run a single-GPU coding LLM   | [Qwen3.6-27B](/guides/language-models/qwen36-27b) (RTX 4090 Q4) or [Ling-2.6-flash](/guides/language-models/ling-26-flash)                                |
| Generate images               | [FLUX.2 Klein](/guides/image-generation/flux2-klein) or [ComfyUI](/guides/image-generation/comfyui)                                                       |
| Generate videos               | [FramePack](/guides/video-generation/framepack) (6GB!) or [Wan2.1](/guides/video-generation/wan-video)                                                    |
| Clone a voice                 | [Voxtral TTS](/guides/audio-and-voice/voxtral-tts), [Qwen3-TTS](/guides/audio-and-voice/qwen3-tts), or [Zonos](/guides/audio-and-voice/zonos-tts)         |
| Transcribe audio              | [WhisperX](/guides/audio-and-voice/whisperx)                                                                                                              |
| Fine-tune a model             | [Unsloth](/guides/training/unsloth-finetune) (2x faster, 70% less VRAM)                                                                                   |
| Generate music                | [ACE-Step](/guides/music-generation/ace-step) (< 4GB VRAM!)                                                                                               |
| Self-host Copilot             | [TabbyML](/guides/ai-coding-tools/tabby) ($4.50/month)                                                                                                    |
| Run an AI agent platform      | [Dify](/guides/ai-platforms-and-agents/dify), [CrewAI](/guides/ai-platforms-and-agents/crewai), or [OpenHands](/guides/ai-platforms-and-agents/openhands) |
| Self-host ChatGPT alternative | [LibreChat](/guides/ai-platforms-and-agents/librechat) or [LobeChat](/guides/ai-platforms-and-agents/lobechat)                                            |
| Run AI assistant 24/7         | [OpenClaw on Clore](/guides/other-workloads/openclaw-on-clore)                                                                                            |
| Pick the right GPU            | [GPU Comparison](/guides/getting-started/gpu-comparison)                                                                                                  |

## 📖 Documentation & Support

* **Main Docs**: [docs.clore.ai](https://docs.clore.ai/)
* **These Guides**: [docs.clore.ai/guides](https://docs.clore.ai/guides/)
* **Marketplace**: [clore.ai/marketplace](https://clore.ai/marketplace)
* **Support**: [clore.ai/support](https://clore.ai/support) / <support@clore.ai>
* **Discord**: [discord.com/invite/clore-ai](https://discord.com/invite/clore-ai)
* **Telegram**: [@clorechat](https://t.me/clorechat)


# Quickstart

{% hint style="success" %}
No prior GPU or AI experience needed. This guide gets you from zero to running AI in 5 minutes.
{% endhint %}

## Step 1: Create Account & Add Funds

1. Go to [clore.ai](https://clore.ai) → **Sign Up**
2. Verify your email
3. Go to **Account** → **Deposit**
4. Add funds via **CLORE**, **BTC**, **USDT**, or **USDC** (minimum \~$5 to start)

## Step 2: Pick a GPU

Go to the [Marketplace](https://clore.ai/marketplace) and choose based on your task:

| What I Want To Do         | Minimum GPU   | Budget/Day |
| ------------------------- | ------------- | ---------- |
| Chat with AI (7B models)  | RTX 3060 12GB | \~$0.15    |
| Chat with AI (32B models) | RTX 4090 24GB | \~$0.50    |
| Generate images (FLUX)    | RTX 3090 24GB | \~$0.30    |
| Generate videos           | RTX 4090 24GB | \~$0.50    |
| Generate music            | Any GPU 4GB+  | \~$0.15    |
| Voice cloning / TTS       | RTX 3060 6GB+ | \~$0.15    |
| Transcribe audio          | RTX 3060 8GB+ | \~$0.15    |
| Fine-tune a model         | RTX 4090 24GB | \~$0.50    |
| Run 70B+ models           | A100 80GB     | \~$2.00    |

{% hint style="danger" %}
**Important — check more than just GPU!**

* **RAM:** 16GB+ minimum for most AI workloads
* **Network:** 500Mbps+ recommended (models download from HuggingFace)
* **Disk:** 50GB+ free space for model storage
  {% endhint %}

### Quick GPU Guide

| GPU           | VRAM | Price          | Sweet Spot For                       |
| ------------- | ---- | -------------- | ------------------------------------ |
| **RTX 3060**  | 12GB | $0.15–0.30/day | TTS, music, small models             |
| **RTX 3090**  | 24GB | $0.30–1.00/day | Image gen, 32B models                |
| **RTX 4090**  | 24GB | $0.50–2.00/day | Everything up to 35B, fast inference |
| **RTX 5090**  | 32GB | $1.50–3.00/day | 70B quantized, fastest               |
| **A100 80GB** | 80GB | $2.00–4.00/day | 70B FP16, serious training           |
| **H100 80GB** | 80GB | $3.00–6.00/day | 400B+ MoE models                     |

## Step 3: Deploy

Click **Rent** on your chosen server, then configure:

* **Order type:** On-Demand (guaranteed) or Spot (30–50% cheaper, can be interrupted)
* **Docker image:** See recipes below
* **Ports:** Always include `22/tcp` (SSH) + your app port
* **Environment:** Add any API keys needed

### 🚀 One-Click Recipes

#### Chat with AI (Ollama + Open WebUI)

The easiest way to run local AI — ChatGPT-like interface with any open model.

```
Image: ghcr.io/open-webui/open-webui:ollama
Ports: 22/tcp, 8080/http
```

After deploy, open the HTTP URL → create account → pick a model (Llama 4 Scout, Gemma 3, Qwen3.5) → chat!

#### Image Generation (ComfyUI)

Node-based workflow for FLUX, Stable Diffusion, and more.

```
Image: yanwk/comfyui-boot:cu126-slim
Ports: 22/tcp, 8188/http
Environment: CLI_ARGS=--listen 0.0.0.0
```

#### Image Generation (Stable Diffusion WebUI)

Classic UI for Stable Diffusion, SDXL, and SD 3.5.

```
Image: universonic/stable-diffusion-webui:latest
Ports: 22/tcp, 8080/http
```

#### LLM API Server (vLLM)

Production-grade serving with OpenAI-compatible API.

```
Image: vllm/vllm-openai:latest
Ports: 22/tcp, 8000/http
Command: vllm serve Qwen/Qwen3.5-9B-Instruct --host 0.0.0.0 --max-model-len 8192
```

#### Music Generation (ACE-Step)

Generate full songs with vocals — works on any 4GB+ GPU!

```
Ports: 22/tcp, 7860/http
```

SSH in, then:

```bash
git clone https://github.com/ACE-Step/ACE-Step-1.5.git && cd ACE-Step-1.5
pip install -r requirements.txt
python app.py --port 7860 --listen 0.0.0.0
```

## Step 4: Connect

After your order starts:

1. Go to **My Orders** → find your active order
2. **Web UI:** Click the HTTP URL (e.g., `https://xxx.clorecloud.net`)
3. **SSH:** `ssh -p <port> root@<proxy-address>`

{% hint style="warning" %}
**First launch takes 5–20 minutes** — the server downloads AI models from HuggingFace. HTTP 502 errors during this time are normal. Wait and refresh.
{% endhint %}

| Deploy              | Typical Startup                  |
| ------------------- | -------------------------------- |
| Ollama + Open WebUI | 3–5 min                          |
| ComfyUI             | 10–15 min                        |
| vLLM                | 5–15 min (depends on model size) |
| SD WebUI            | 10–20 min                        |

## Step 5: Start Creating

Once your service is running, explore the guides for your specific use case:

### 🤖 Language Models (Chat, Code, Reasoning)

* [**Ollama**](/guides/language-models/ollama) — easiest model management
* [**Llama 4 Scout**](/guides/language-models/llama4) — Meta's latest, 10M context
* [**Gemma 3**](/guides/language-models/gemma3) — Google's 27B that beats 405B models
* [**Qwen3.5**](/guides/language-models/qwen35) — beat Claude 4.5 on math (Feb 2026!)
* [**DeepSeek-R1**](/guides/language-models/deepseek-r1) — chain-of-thought reasoning
* [**vLLM**](/guides/language-models/vllm) — production API serving

### 🎨 Image Generation

* [**FLUX.2 Klein**](/guides/image-generation/flux2-klein) — < 0.5 sec per image!
* [**ComfyUI**](/guides/image-generation/comfyui) — node-based workflows
* [**FLUX.1**](/guides/image-generation/flux) — highest quality with LoRA + ControlNet
* [**Stable Diffusion 3.5**](/guides/image-generation/stable-diffusion-3-5) — best text rendering

### 🎬 Video Generation

* [**FramePack**](/guides/video-generation/framepack) — only 6GB VRAM needed!
* [**Wan2.1**](/guides/video-generation/wan-video) — high quality T2V + I2V
* [**LTX-2**](/guides/video-generation/ltx-video-2) — video WITH audio
* [**CogVideoX**](/guides/video-generation/cogvideox) — Zhipu AI's video model

### 🔊 Audio & Voice

* [**Qwen3-TTS**](/guides/audio-and-voice/qwen3-tts) — voice cloning, 10+ languages
* [**WhisperX**](/guides/audio-and-voice/whisperx) — transcription + speaker diarization
* [**Dia TTS**](/guides/audio-and-voice/dia-tts) — multi-speaker dialog
* [**Kokoro**](/guides/audio-and-voice/kokoro-tts) — tiny TTS, only 2GB VRAM

### 🎵 Music

* [**ACE-Step**](/guides/music-generation/ace-step) — full songs on < 4GB VRAM

### 💻 AI Coding

* [**TabbyML**](/guides/ai-coding-tools/tabby) — self-hosted Copilot for $4.50/month
* [**Aider**](/guides/ai-coding-tools/aider) — terminal AI coding assistant

### 🧠 Training

* [**Unsloth**](/guides/training/unsloth-finetune) — 2x faster, 70% less VRAM
* [**Axolotl**](/guides/training/axolotl-training) — YAML-based fine-tuning

## 💡 Tips for Beginners

1. **Start with Ollama** — it's the easiest way to try AI locally
2. **RTX 4090 is the sweet spot** — handles 90% of use cases at $0.50–2/day
3. **Use Spot orders** for experiments — 30–50% cheaper
4. **Use On-Demand** for important work — guaranteed, no interruptions
5. **Download your outputs** before the order ends — files are deleted after
6. **Pay with CLORE token** — often better rates than stablecoins
7. **Check RAM and network** — low RAM is the #1 cause of failures

## Troubleshooting

| Problem                  | Solution                                                         |
| ------------------------ | ---------------------------------------------------------------- |
| HTTP 502 for a long time | Wait 10–20 min for first startup; check RAM ≥ 16GB               |
| Service won't start      | RAM too low (need 16GB+) or VRAM too small for the model         |
| Slow model download      | Normal on first run; prefer 500Mbps+ servers                     |
| CUDA out of memory       | Use smaller model or bigger GPU; try quantized versions          |
| Can't SSH                | Check port is `22/tcp` in config; wait for server to fully start |

## 🐍 Python SDK & CLI (Recommended)

Prefer code over clicking? Install the official SDK:

```bash
pip install clore-ai
clore search --gpu "RTX 4090" --max-price 5.0
clore deploy 123 --image cloreai/ubuntu22.04-cuda12 --type on-demand --currency bitcoin --ssh-password mypass --port 22:tcp
clore ssh 456
```

Or use Python directly:

```python
from clore_ai import CloreAI

client = CloreAI()
servers = client.marketplace(gpu="RTX 4090", max_price_usd=5.0)
order = client.create_order(server_id=servers[0].id, image="cloreai/ubuntu22.04-cuda12", type="on-demand", currency="bitcoin")
```

→ [Full Python Quickstart](/guides/getting-started/python-quickstart) | [SDK Guide](/guides/advanced/python-sdk) | [CLI Automation](/guides/advanced/cli-automation)

## Need Help?

* 📖 [Full Troubleshooting Guide](/guides/getting-started/clore-troubleshooting)
* 📊 [GPU Comparison Chart](/guides/getting-started/gpu-comparison)
* 💰 [Pricing Reference](/guides/getting-started/pricing)
* 💬 [Discord](https://discord.com/invite/clore-ai)
* 💬 [Telegram](https://t.me/clorechat)
* 📧 <support@clore.ai>


# Overview

Essential guides for new users of CLORE.AI GPU marketplace.

## What You'll Learn

* How to rent GPUs on CLORE.AI
* Understanding pricing and payment options
* Troubleshooting common issues

## Quick Start

1. **Create Account** - Sign up at [clore.ai](https://clore.ai)
2. **Add Funds** - Deposit CLORE, BTC, or USDT/USDC
3. **Browse Marketplace** - Find GPUs at [clore.ai/marketplace](https://clore.ai/marketplace)
4. **Create Order** - Select GPU, configure Docker image, set ports
5. **Connect** - SSH or web interface to your server

> 📚 New to GPU cloud? Read our [Complete Beginner's Guide to Renting a GPU on Clore.ai](https://blog.clore.ai/how-to-rent-a-gpu-complete-beginners-guide/)

## Guides in This Section

| Guide                                                                     | Description                                   |
| ------------------------------------------------------------------------- | --------------------------------------------- |
| [Pricing](/guides/getting-started/pricing)                                | GPU rental rates and cost optimization        |
| [Troubleshooting](/guides/getting-started/clore-troubleshooting)          | Common issues and solutions                   |
| [Cost Calculator](/guides/getting-started/cost-calculator)                | Calculate and estimate GPU rental costs       |
| [Docker Images Catalog](/guides/getting-started/docker-images)            | Pre-configured Docker images for AI workloads |
| [Frequently Asked Questions](/guides/getting-started/faq)                 | Common questions and answers                  |
| [GPU Comparison Guide](/guides/getting-started/gpu-comparison)            | Compare GPU models and specifications         |
| [Model Compatibility Matrix](/guides/getting-started/model-compatibility) | Check which models work with which GPUs       |
| [Python SDK Quickstart](/guides/getting-started/python-quickstart)        | Get started with the Python SDK               |

## Payment Options

* **CLORE** - Native token, often discounted
* **BTC** - Bitcoin payments
* **USDT/USDC** - Stablecoin payments

## Order Types

* **On-Demand** - Fixed hourly rate, guaranteed availability
* **Spot** - Bid-based pricing, typically 30-50% cheaper

## Need Help?

* [Discord](https://discord.com/invite/clore-ai)
* [Telegram](https://t.me/clorechat)
* [Main Documentation](https://docs.clore.ai)


# Python SDK Quickstart

Find a GPU, rent it, and SSH in — all from Python or terminal in 5 minutes

{% hint style="success" %}
**What you'll build:** In 5 minutes, you'll search the GPU marketplace, rent a server, and connect via SSH — all from code or CLI.
{% endhint %}

## Prerequisites

* **Python 3.9+** installed
* **Clore.ai account** — [sign up here](https://clore.ai)
* **API key** — get it from your [Clore.ai dashboard](https://clore.ai)
* **Funds** — minimum \~$5 in BTC, CLORE, USDT, or USDC

***

## Step 1: Install the SDK

```bash
pip install clore-ai
```

Verify installation:

```bash
clore --version
```

***

## Step 2: Configure Your API Key

**Option A: Environment variable (recommended)**

```bash
export CLORE_API_KEY=your_api_key_here
```

**Option B: CLI config (persisted to `~/.clore/config.json`)**

```bash
clore config set api_key YOUR_API_KEY
```

***

## Step 3: Search for a GPU

### CLI

```bash
clore search --gpu "RTX 4090" --max-price 5.0 --sort price --limit 10
```

You'll see a table with server IDs, GPU specs, RAM, price, and location.

### Python

```python
from clore_ai import CloreAI

client = CloreAI()

servers = client.marketplace(gpu="RTX 4090", max_price_usd=5.0)
for s in servers[:5]:
    print(f"Server {s.id}: {s.gpu_count}x {s.gpu_model} — ${s.price_usd:.4f}/h — {s.location}")
```

**Output:**

```
Server 142: 1x NVIDIA GeForce RTX 4090 — $0.0800/h — EU
Server 305: 1x NVIDIA GeForce RTX 4090 — $0.0950/h — US
Server 891: 2x NVIDIA GeForce RTX 4090 — $0.1200/h — EU
```

{% hint style="info" %}
**Tip:** The `marketplace()` method filters client-side. You can also filter by `min_gpu_count`, `min_ram_gb`, and `available_only` (default: `True`).
{% endhint %}

***

## Step 4: Deploy (Rent a Server)

### CLI

```bash
clore deploy 142 \
  --image cloreai/ubuntu22.04-cuda12 \
  --type on-demand \
  --currency bitcoin \
  --ssh-password MySecurePass123 \
  --port 22:tcp \
  --port 8888:http
```

### Python

```python
order = client.create_order(
    server_id=142,
    image="cloreai/ubuntu22.04-cuda12",
    type="on-demand",
    currency="bitcoin",
    ssh_password="MySecurePass123",
    ports={"22": "tcp", "8888": "http"}
)

print(f"Order created! ID: {order.id}")
print(f"IP: {order.pub_cluster}")
print(f"Ports: {order.tcp_ports}")
```

{% hint style="warning" %}
**First boot takes 1–5 minutes.** The server pulls the Docker image and starts services. If `pub_cluster` is `None`, wait and check again with `my_orders()`.
{% endhint %}

***

## Step 5: Connect via SSH

### CLI (auto-connects)

```bash
clore ssh 38
```

The CLI looks up the order, finds the public IP and SSH port, and runs `ssh` for you.

### Manual SSH

```bash
ssh root@<pub_cluster> -p <ssh_port>
# Password: MySecurePass123
```

***

## Step 6: Cleanup

When you're done, cancel the order to stop billing:

### CLI

```bash
clore cancel 38
```

### Python

```python
client.cancel_order(order_id=38, issue="Done with my work")
print("Order cancelled")
```

***

## Full Script: Search → Deploy → Monitor → Cancel

```python
from clore_ai import CloreAI
import time

client = CloreAI()  # Uses CLORE_API_KEY env var

# 1. Find cheapest RTX 4090
servers = client.marketplace(gpu="RTX 4090", max_price_usd=5.0)
servers.sort(key=lambda s: s.price_usd or float("inf"))

if not servers:
    print("No RTX 4090 available under $5/h")
    exit(1)

best = servers[0]
print(f"Best deal: Server {best.id} — ${best.price_usd:.4f}/h — {best.location}")

# 2. Deploy
order = client.create_order(
    server_id=best.id,
    image="cloreai/ubuntu22.04-cuda12",
    type="on-demand",
    currency="bitcoin",
    ssh_password="MySecurePass123",
    ports={"22": "tcp"}
)
print(f"Order {order.id} created!")

# 3. Wait for IP assignment
for _ in range(12):
    orders = client.my_orders()
    current = next((o for o in orders if o.id == order.id), None)
    if current and current.pub_cluster:
        print(f"Ready! SSH: ssh root@{current.pub_cluster} -p {current.tcp_ports.get('22', 22)}")
        break
    print("Waiting for server to start...")
    time.sleep(10)

# 4. ... do your work ...

# 5. Cancel when done
client.cancel_order(order_id=order.id)
print("Order cancelled, billing stopped.")
```

***

## What's Next

| Guide                                                    | What You'll Learn                                                |
| -------------------------------------------------------- | ---------------------------------------------------------------- |
| [Python SDK Guide](/guides/advanced/python-sdk)          | Async operations, spot market, server management, error handling |
| [CLI Automation](/guides/advanced/cli-automation)        | Bash scripts, CI/CD integration, batch deploy                    |
| [API Integration](/guides/advanced/api-integration)      | Connect AI services running on Clore to your apps                |
| [GPU Comparison](/guides/getting-started/gpu-comparison) | Choose the right GPU for your workload                           |

***

**Happy GPU renting! 🚀**


# FAQ

Frequently asked questions about using Clore.ai for AI workloads

Common questions about using CLORE.AI for AI workloads.

{% hint style="success" %}
Can't find your answer? Join [Discord](https://discord.com/invite/clore-ai) or [Telegram](https://t.me/clorechat) for help.
{% endhint %}

## Getting Started

### How do I create an account?

1. Go to [clore.ai](https://clore.ai)
2. Click **Sign Up**
3. Enter email and password
4. Verify your email
5. Done! You can now deposit funds and rent GPUs

### What payment methods are accepted?

* **CLORE** - Native token (often gives best rates)
* **BTC** - Bitcoin
* **USDT** - Tether (ERC-20, TRC-20)
* **USDC** - USD Coin

### What's the minimum deposit?

There's no strict minimum, but we recommend $5-10 to start experimenting. This covers several hours on budget GPUs.

### How long until my deposit arrives?

| Currency  | Confirmations     | Typical Time |
| --------- | ----------------- | ------------ |
| CLORE     | 10                | \~10 minutes |
| BTC       | 2                 | \~20 minutes |
| USDT/USDC | Network dependent | 1-15 minutes |

***

## Choosing Hardware

### Which GPU should I choose?

It depends on your task:

| Task                     | Recommended GPU                                              |
| ------------------------ | ------------------------------------------------------------ |
| Chat with 7B models      | RTX 3060 12GB ($0.15–0.30/day)                               |
| Chat with 13B-30B models | RTX 3090 24GB ($0.30–1.00/day)                               |
| Chat with 70B models     | RTX 5090 32GB ($1.50–3.00/day) or A100 40GB ($1.50–3.00/day) |
| Image generation (SDXL)  | RTX 3090 24GB ($0.30–1.00/day)                               |
| Image generation (FLUX)  | RTX 4090/5090 ($0.50–3.00/day)                               |
| Video generation         | RTX 4090+ or A100 ($0.50–3.00/day)                           |
| 70B+ models (FP16)       | A100 80GB ($2.00–4.00/day)                                   |

See [GPU Comparison](/guides/getting-started/gpu-comparison) for detailed specs.

### What's the difference between On-Demand and Spot?

| Type          | Price          | Availability       | Best For               |
| ------------- | -------------- | ------------------ | ---------------------- |
| **On-Demand** | Fixed price    | Guaranteed         | Production, long tasks |
| **Spot**      | 30-50% cheaper | Can be interrupted | Testing, batch jobs    |

Spot orders may be terminated if someone bids higher. Save your work frequently!

### How much VRAM do I need?

Use this quick guide:

| Model Size | Minimum VRAM (Q4) | Recommended |
| ---------- | ----------------- | ----------- |
| 7B         | 6GB               | 12GB        |
| 13B        | 8GB               | 16GB        |
| 30B        | 16GB              | 24GB        |
| 70B        | 35GB              | 48GB        |

See [Model Compatibility](/guides/getting-started/model-compatibility) for full details.

### What does "Unverified" mean on a server?

Unverified servers haven't completed CLORE.AI's hardware verification. They may:

* Have slightly different specs than listed
* Be less reliable

Verified servers have confirmed specs and better reliability.

***

## Connecting to Servers

### How do I connect via SSH?

After your order starts:

1. Go to **My Orders**
2. Find your order
3. Copy SSH command: `ssh -p <port> root@<proxy-address>`
4. Use password shown in order details

**Example:**

```bash
ssh -p 10022 root@proxy.clore.ai
# Enter password when prompted
```

### SSH connection refused - what do I do?

1. **Wait 1-2 minutes** - Server may still be starting
2. **Check order status** - Must be "Running"
3. **Verify port** - Use the port from order details, not 22
4. **Check firewall** - Your network may block non-standard ports

### How do I access web interfaces (Gradio, Jupyter)?

1. In order details, find the HTTP port link
2. Click the link to open in browser
3. Or manually: `http://<proxy-address>:<http-port>`

### Can I use VS Code Remote SSH?

Yes! Add to your `~/.ssh/config`:

```
Host clore-gpu
    HostName <proxy-address>
    Port <port>
    User root
```

Then connect in VS Code: `Remote-SSH: Connect to Host` → `clore-gpu`

### How do I transfer files?

**Upload to server:**

```bash
scp -P <port> local_file.zip root@<proxy>:/root/
```

**Download from server:**

```bash
scp -P <port> root@<proxy>:/root/output.zip ./
```

**For large transfers, use rsync:**

```bash
rsync -avz -e "ssh -p <port>" ./data/ root@<proxy>:/root/data/
```

***

## Running AI Workloads

### How do I install Python packages?

```bash
pip install torch transformers accelerate
```

For persistent installations, include in your startup command or Docker image.

### Why is my model running out of memory?

1. **Use quantization** - Q4 uses 4x less VRAM than FP16
2. **Enable CPU offload** - Slower but works
3. **Reduce batch size** - Use batch\_size=1
4. **Lower context length** - Reduce max\_tokens
5. **Choose bigger GPU** - Sometimes necessary

### How do I use Hugging Face models that require login?

```bash
# Set token as environment variable
export HUGGING_FACE_HUB_TOKEN=hf_xxxxx

# Or login interactively
huggingface-cli login
```

For gated models (Llama, etc.), first accept terms on Hugging Face website.

### Can I run multiple models at once?

Yes, if you have enough VRAM:

* Each model needs its own VRAM allocation
* Use different ports for different services
* Consider vLLM for efficient multi-model serving

### How do I keep my work after the order ends?

**Before order ends:**

1. Download outputs: `scp -P <port> root@<proxy>:/root/outputs/* ./`
2. Push to cloud: `aws s3 sync /root/outputs s3://bucket/`
3. Commit to Git: `git push`

**Data is lost when order ends!** Always backup important files.

***

## Billing & Orders

### How is billing calculated?

* **Hourly rate** × **hours used**
* Billing starts when order status is "Running"
* Minimum billing: 1 minute
* Partial hours billed by the minute

### How do I stop an order?

1. Go to **My Orders**
2. Click **Stop** on your order
3. Confirm termination

You'll only be charged for time used.

### Can I extend my order?

Yes, if no one has bid higher on your spot order. Add funds to your balance and the order continues automatically.

### What happens if my balance runs out?

1. Warning notification sent
2. Grace period (\~5-10 minutes)
3. Order terminated
4. **All data lost!**

Keep sufficient balance or download work before running low.

### Can I get a refund?

Contact support for:

* Hardware issues (GPU not working)
* Significant downtime
* Billing errors

Refunds not provided for:

* User errors
* Normal order usage
* Spot order interruptions

***

## Docker & Images

### What Docker images are available?

See [Docker Images Catalog](/guides/getting-started/docker-images) for ready-to-use images:

* `ollama/ollama` - LLM runner
* `vllm/vllm-openai` - High-performance LLM API
* `universonic/stable-diffusion-webui` - Image generation
* `yanwk/comfyui-boot` - Node-based image gen

### Can I use my own Docker image?

Yes! Specify your image in the order:

```
Image: your-registry/your-image:tag
```

Ensure the image is publicly accessible or provide credentials.

### How do I persist data between restarts?

Use the provided volume mount points or specify custom volumes:

```
-v /data/models:/root/.cache/huggingface
```

***

## Troubleshooting

### "CUDA out of memory" error

1. **Check VRAM** - `nvidia-smi`
2. **Use smaller model** or quantization
3. **Enable offload** - `--cpu-offload`
4. **Reduce batch size** - `batch_size=1`
5. **Kill other processes** - `pkill python`

### "Connection timed out" error

1. **Check order status** - Must be "Running"
2. **Wait longer** - Large images take time to start
3. **Check ports** - Use correct port from order
4. **Try again** - Network issues are temporary

### Web UI not loading

1. **Wait for startup** - Some UIs take 2-5 minutes
2. **Check logs**: `docker logs <container>`
3. **Verify port** - Use HTTP port from order details
4. **Check command** - Must include `--listen 0.0.0.0`

### Model download stuck

1. **Check disk space**: `df -h`
2. **Use smaller model** - Start with 7B
3. **Pre-download** - Include in Docker image
4. **Check HF token** - Required for gated models

### Slow inference speed

1. **Check GPU usage**: `nvidia-smi`
2. **Enable GPU** - Model may be on CPU
3. **Use Flash Attention** - `--flash-attn`
4. **Optimize settings** - Reduce precision, batch

***

## Security

### Is my data secure?

* Each order runs in isolated container
* Data is deleted when order ends
* Network traffic goes through CLORE proxies
* Don't store sensitive data on rented GPUs

### Should I change the root password?

Optional but recommended for long orders:

```bash
passwd root
```

### How do I protect API keys?

1. **Use environment variables** - Not command line args
2. **Don't commit to Git** - Use `.gitignore`
3. **Delete after use** - Clear history: `history -c`

***

## Still Need Help?

* [Troubleshooting Guide](/guides/getting-started/clore-troubleshooting)
* [Discord Community](https://discord.com/invite/clore-ai)
* [Telegram Chat](https://t.me/clorechat)
* Email: <support@clore.ai>


# GPU Comparison

Complete GPU comparison guide for AI workloads on Clore.ai

Complete comparison of GPUs available on CLORE.AI for AI workloads.

{% hint style="success" %}
Find the right GPU for your task at [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Quick Recommendation

| Your Task                 | Budget Pick   | Best Value    | Maximum Performance |
| ------------------------- | ------------- | ------------- | ------------------- |
| Chat with AI (7B)         | RTX 3060 12GB | RTX 3090 24GB | RTX 5090 32GB       |
| Chat with AI (70B)        | RTX 3090 24GB | RTX 5090 32GB | A100 80GB           |
| Image Generation (SD 1.5) | RTX 3060 12GB | RTX 3090 24GB | RTX 5090 32GB       |
| Image Generation (SDXL)   | RTX 3090 24GB | RTX 4090 24GB | RTX 5090 32GB       |
| Image Generation (FLUX)   | RTX 3090 24GB | RTX 5090 32GB | A100 80GB           |
| Video Generation          | RTX 4090 24GB | RTX 5090 32GB | A100 80GB           |
| Model Training            | A100 40GB     | A100 80GB     | H100 80GB           |

## Consumer GPUs

### NVIDIA RTX 3060 12GB

**Best for:** Budget AI, SD 1.5, small LLMs

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 12GB GDDR6    |
| Memory Bandwidth | 360 GB/s      |
| FP16 Performance | 12.7 TFLOPS   |
| Tensor Cores     | 112 (3rd gen) |
| TDP              | 170W          |
| \~Price/hour     | $0.02-0.04    |

**Capabilities:**

* ✅ Ollama with 7B models (Q4)
* ✅ Stable Diffusion 1.5 (512x512)
* ✅ SDXL (768x768, slow)
* ⚠️ FLUX schnell (with CPU offload)
* ❌ Large models (>13B)
* ❌ Video generation

***

### NVIDIA RTX 3070/3070 Ti 8GB

**Best for:** SD 1.5, lightweight tasks

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 8GB GDDR6X    |
| Memory Bandwidth | 448-608 GB/s  |
| FP16 Performance | 20.3 TFLOPS   |
| Tensor Cores     | 184 (3rd gen) |
| TDP              | 220-290W      |
| \~Price/hour     | $0.02-0.04    |

**Capabilities:**

* ✅ Ollama with 7B models (Q4)
* ✅ Stable Diffusion 1.5 (512x512)
* ⚠️ SDXL (low resolution only)
* ❌ FLUX (insufficient VRAM)
* ❌ Models >7B
* ❌ Video generation

***

### NVIDIA RTX 3080/3080 Ti 10-12GB

**Best for:** General AI tasks, good balance

| Spec             | Value             |
| ---------------- | ----------------- |
| VRAM             | 10-12GB GDDR6X    |
| Memory Bandwidth | 760-912 GB/s      |
| FP16 Performance | 29.8-34.1 TFLOPS  |
| Tensor Cores     | 272-320 (3rd gen) |
| TDP              | 320-350W          |
| \~Price/hour     | $0.04-0.06        |

**Capabilities:**

* ✅ Ollama with 13B models
* ✅ Stable Diffusion 1.5/2.1
* ✅ SDXL (1024x1024)
* ⚠️ FLUX schnell (with offload)
* ❌ Large models (>13B)
* ❌ Video generation

***

### NVIDIA RTX 3090/3090 Ti 24GB

**Best for:** SDXL, 13B-30B LLMs, ControlNet

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 24GB GDDR6X   |
| Memory Bandwidth | 936 GB/s      |
| FP16 Performance | 35.6 TFLOPS   |
| Tensor Cores     | 328 (3rd gen) |
| TDP              | 350-450W      |
| \~Price/hour     | $0.05-0.08    |

**Capabilities:**

* ✅ Ollama with 30B models
* ✅ vLLM with 13B models
* ✅ All Stable Diffusion models
* ✅ SDXL + ControlNet
* ✅ FLUX schnell (1024x1024)
* ⚠️ FLUX dev (with offload)
* ⚠️ Video (short clips)

***

### NVIDIA RTX 4070 Ti 12GB

**Best for:** Fast SD 1.5, efficient inference

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 12GB GDDR6X   |
| Memory Bandwidth | 504 GB/s      |
| FP16 Performance | 40.1 TFLOPS   |
| Tensor Cores     | 184 (4th gen) |
| TDP              | 285W          |
| \~Price/hour     | $0.04-0.06    |

**Capabilities:**

* ✅ Ollama with 7B models (fast)
* ✅ Stable Diffusion 1.5 (very fast)
* ✅ SDXL (768x768)
* ⚠️ FLUX schnell (limited res)
* ❌ Large models (>13B)
* ❌ Video generation

***

### NVIDIA RTX 4080 16GB

**Best for:** SDXL production, 13B LLMs

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 16GB GDDR6X   |
| Memory Bandwidth | 717 GB/s      |
| FP16 Performance | 48.7 TFLOPS   |
| Tensor Cores     | 304 (4th gen) |
| TDP              | 320W          |
| \~Price/hour     | $0.06-0.09    |

**Capabilities:**

* ✅ Ollama with 13B models (fast)
* ✅ vLLM with 7B models
* ✅ All Stable Diffusion models
* ✅ SDXL + ControlNet
* ✅ FLUX schnell (1024x1024)
* ⚠️ FLUX dev (limited)
* ⚠️ Short video clips

***

### NVIDIA RTX 4090 24GB

**Best for:** High-end consumer performance, FLUX, video

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 24GB GDDR6X   |
| Memory Bandwidth | 1008 GB/s     |
| FP16 Performance | 82.6 TFLOPS   |
| Tensor Cores     | 512 (4th gen) |
| TDP              | 450W          |
| \~Price/hour     | $0.08-0.12    |

**Capabilities:**

* ✅ Ollama with 30B models (fast)
* ✅ vLLM with 13B models
* ✅ All image generation models
* ✅ FLUX dev (1024x1024)
* ✅ Video generation (short)
* ✅ AnimateDiff
* ⚠️ 70B models (Q4 only)

***

### NVIDIA RTX 5080 16GB *(New — Feb 2025)*

**Best for:** Fast SDXL/FLUX, 13B-30B LLMs, high-performance mid-range

| Spec                  | Value         |
| --------------------- | ------------- |
| VRAM                  | 16GB GDDR7    |
| Memory Bandwidth      | 960 GB/s      |
| FP16 Performance      | \~80 TFLOPS   |
| Tensor Cores          | 336 (5th gen) |
| TDP                   | 360W          |
| \~Clore.ai Price/hour | $1.50-2.00    |

**Capabilities:**

* ✅ Ollama with 13B models (fast)
* ✅ vLLM with 13B models
* ✅ All Stable Diffusion models
* ✅ SDXL + ControlNet (very fast)
* ✅ FLUX schnell/dev (1024x1024)
* ✅ Short video clips
* ⚠️ 30B models (Q4 only)
* ❌ 70B models

***

### NVIDIA RTX 5090 32GB *(Flagship — Feb 2025)*

**Best for:** Maximum consumer performance, 70B models, high-res video generation

| Spec                  | Value         |
| --------------------- | ------------- |
| VRAM                  | 32GB GDDR7    |
| Memory Bandwidth      | 1792 GB/s     |
| FP16 Performance      | \~120 TFLOPS  |
| Tensor Cores          | 680 (5th gen) |
| TDP                   | 575W          |
| \~Clore.ai Price/hour | $3.00-4.00    |

**Capabilities:**

* ✅ Ollama with 70B models (Q4, fast)
* ✅ vLLM with 30B models
* ✅ All image generation models
* ✅ FLUX dev (1536x1536)
* ✅ Video generation (longer clips)
* ✅ AnimateDiff + ControlNet
* ✅ Model training (LoRA, small fine-tunes)
* ✅ DeepSeek-R1 32B distill (FP16)

## Professional/Datacenter GPUs

### NVIDIA A100 40GB

**Best for:** Production LLMs, training, large models

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 40GB HBM2e    |
| Memory Bandwidth | 1555 GB/s     |
| FP16 Performance | 77.97 TFLOPS  |
| Tensor Cores     | 432 (3rd gen) |
| TDP              | 400W          |
| \~Price/hour     | $0.15-0.20    |

**Capabilities:**

* ✅ Ollama with 70B models (Q4)
* ✅ vLLM production serving
* ✅ All image generation
* ✅ FLUX dev (high quality)
* ✅ Video generation
* ✅ Model fine-tuning
* ⚠️ 70B FP16 (tight)

***

### NVIDIA A100 80GB

**Best for:** 70B+ models, video, production workloads

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 80GB HBM2e    |
| Memory Bandwidth | 2039 GB/s     |
| FP16 Performance | 77.97 TFLOPS  |
| Tensor Cores     | 432 (3rd gen) |
| TDP              | 400W          |
| \~Price/hour     | $0.20-0.30    |

**Capabilities:**

* ✅ All LLMs up to 70B (FP16)
* ✅ vLLM high-throughput serving
* ✅ All image generation
* ✅ Long video generation
* ✅ Model training
* ✅ DeepSeek-V3 (partial)
* ⚠️ 100B+ models

***

### NVIDIA H100 80GB

**Best for:** Maximum performance, largest models

| Spec             | Value         |
| ---------------- | ------------- |
| VRAM             | 80GB HBM3     |
| Memory Bandwidth | 3350 GB/s     |
| FP16 Performance | 267 TFLOPS    |
| Tensor Cores     | 528 (4th gen) |
| TDP              | 700W          |
| \~Price/hour     | $0.40-0.60    |

**Capabilities:**

* ✅ All models with maximum speed
* ✅ 100B+ parameter models
* ✅ Multi-model serving
* ✅ Large-scale training
* ✅ Real-time video generation
* ✅ DeepSeek-V3 (671B)

## Performance Comparisons

### LLM Inference (tokens/second)

| GPU           | Llama 3 8B | Llama 3 70B | Mixtral 8x7B | Clore.ai $/hr |
| ------------- | ---------- | ----------- | ------------ | ------------- |
| RTX 3060 12GB | 25         | -           | -            | $0.02-0.04    |
| RTX 3090 24GB | 45         | 8\*         | 20\*         | $0.15-0.25    |
| RTX 4090 24GB | 80         | 15\*        | 35\*         | $0.35-0.55    |
| RTX 5080 16GB | 95         | -           | 40\*         | $1.50-2.00    |
| RTX 5090 32GB | 150        | 30\*        | 65\*         | $3.00-4.00    |
| A100 40GB     | 100        | 25          | 45           | $0.80-1.20    |
| A100 80GB     | 110        | 40          | 55           | $1.20-1.80    |
| H100 80GB     | 180        | 70          | 90           | $2.50-3.50    |

\*With quantization (Q4/Q8)

### Image Generation Speed

| GPU           | SD 1.5 (512) | SDXL (1024) | FLUX schnell | Clore.ai $/hr |
| ------------- | ------------ | ----------- | ------------ | ------------- |
| RTX 3060 12GB | 4 sec        | 15 sec      | 25 sec\*     | $0.02-0.04    |
| RTX 3090 24GB | 2 sec        | 7 sec       | 12 sec       | $0.15-0.25    |
| RTX 4090 24GB | 1 sec        | 3 sec       | 5 sec        | $0.35-0.55    |
| RTX 5080 16GB | 0.8 sec      | 2.5 sec     | 4 sec        | $1.50-2.00    |
| RTX 5090 32GB | 0.6 sec      | 1.8 sec     | 3 sec        | $3.00-4.00    |
| A100 40GB     | 1.5 sec      | 4 sec       | 6 sec        | $0.80-1.20    |
| A100 80GB     | 1.5 sec      | 4 sec       | 5 sec        | $1.20-1.80    |

\*With CPU offload, lower resolution

### Video Generation (5 sec clip)

| GPU           | SVD     | Wan2.1  | Hunyuan |
| ------------- | ------- | ------- | ------- |
| RTX 3090 24GB | 3 min   | 5 min\* | -       |
| RTX 4090 24GB | 1.5 min | 3 min   | 8 min\* |
| RTX 5090 32GB | 1 min   | 2 min   | 5 min   |
| A100 40GB     | 1 min   | 2 min   | 5 min   |
| A100 80GB     | 45 sec  | 1.5 min | 3 min   |

\*Limited resolution

## Price/Performance Ratio

### Best Value by Task

**Chat/LLM (7B-13B models):**

1. 🥇 RTX 3090 24GB - Best price/performance
2. 🥈 RTX 3060 12GB - Lowest cost
3. 🥉 RTX 4090 24GB - Fastest

**Image Generation (SDXL/FLUX):**

1. 🥇 RTX 3090 24GB - Great balance
2. 🥈 RTX 4090 24GB - 2x faster
3. 🥉 A100 40GB - Production stability

**Large Models (70B+):**

1. 🥇 A100 40GB - Best value for 70B
2. 🥈 A100 80GB - Full precision
3. 🥉 RTX 4090 24GB - Budget option (Q4 only)

**Video Generation:**

1. 🥇 A100 40GB - Good balance
2. 🥈 RTX 4090 24GB - Consumer option
3. 🥉 A100 80GB - Longest clips

**Model Training:**

1. 🥇 A100 40GB - Standard choice
2. 🥈 A100 80GB - Large models
3. 🥉 RTX 4090 24GB - Small models/LoRA

## Multi-GPU Configurations

Some tasks benefit from multiple GPUs:

| Configuration | Use Case                | VRAM Total |
| ------------- | ----------------------- | ---------- |
| 2x RTX 3090   | 70B inference           | 48GB       |
| 2x RTX 4090   | Fast 70B, training      | 48GB       |
| 2x RTX 5090   | 70B FP16, fast training | 64GB       |
| 4x RTX 5090   | 100B+ models            | 128GB      |
| 4x A100 40GB  | 100B+ models            | 160GB      |
| 8x A100 80GB  | DeepSeek-V3, Llama 405B | 640GB      |

## Choosing Your GPU

### Decision Flowchart

```
What's your main task?
│
├─ Chat/LLM
│  ├─ Model size?
│  │  ├─ ≤7B → RTX 3060 ($0.15–0.30/day)
│  │  ├─ 7B-30B → RTX 3090 ($0.30–1.00/day)
│  │  ├─ 30B-70B → A100 40GB ($1.50–3.00/day)
│  │  └─ 70B+ → A100 80GB ($2.00–4.00/day)
│
├─ Image Generation
│  ├─ Model?
│  │  ├─ SD 1.5 → RTX 3060 ($0.15–0.30/day)
│  │  ├─ SDXL → RTX 3090 ($0.30–1.00/day)
│  │  └─ FLUX → RTX 4090 ($0.50–2.00/day)
│
├─ Video Generation
│  ├─ Length?
│  │  ├─ Short (2-5 sec) → RTX 4090 ($0.50–2.00/day)
│  │  └─ Longer → A100 40GB+ ($1.50–3.00+/day)
│
└─ Training
   ├─ LoRA/small → RTX 4090 ($0.50–2.00/day)
   └─ Full fine-tune → A100 40GB+ ($1.50–3.00+/day)
```

## Tips for Saving Money

1. **Use Spot Orders** - 30-50% cheaper than on-demand
2. **Start Small** - Test on cheaper GPUs first
3. **Quantize Models** - Q4/Q8 fits larger models in less VRAM
4. **Batch Processing** - Process multiple requests at once
5. **Off-peak Hours** - Better availability and sometimes lower prices

> 📚 See also: [Top 10 Cheapest GPUs for AI Training in 2025](https://blog.clore.ai/top-10-cheapest-gpus-for-ai-training/) | [Best GPU for AI Training — Detailed Guide](https://blog.clore.ai/best-gpu-for-ai-training/)

## Next Steps

* [Model Compatibility Matrix](/guides/getting-started/model-compatibility) - Which models run on which GPUs
* [Docker Images Catalog](/guides/getting-started/docker-images) - Ready-to-use images
* [Quickstart Guide](/guides/quickstart) - Get started in 5 minutes


# Model Compatibility

AI model and GPU compatibility matrix for Clore.ai

Complete guide to which AI models run on which GPUs on CLORE.AI.

{% hint style="success" %}
Find GPUs with the right VRAM at [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Quick Reference

### Language Models (LLM)

| Model                   | Parameters | Min VRAM   | Recommended            | Quantization             |
| ----------------------- | ---------- | ---------- | ---------------------- | ------------------------ |
| Llama 3.2               | 1B         | 2GB        | 4GB                    | Q4, Q8, FP16             |
| Llama 3.2               | 3B         | 4GB        | 6GB                    | Q4, Q8, FP16             |
| Llama 3.1/3             | 8B         | 6GB        | 12GB                   | Q4, Q8, FP16             |
| Mistral                 | 7B         | 6GB        | 12GB                   | Q4, Q8, FP16             |
| Qwen 2.5                | 7B         | 6GB        | 12GB                   | Q4, Q8, FP16             |
| Qwen 2.5                | 14B        | 12GB       | 16GB                   | Q4, Q8                   |
| Qwen 2.5                | 32B        | 20GB       | 24GB                   | Q4, Q8                   |
| Llama 3.1               | 70B        | 40GB       | 48GB                   | Q4, Q8                   |
| Qwen 2.5                | 72B        | 48GB       | 80GB                   | Q4, Q8                   |
| Mixtral                 | 8x7B       | 24GB       | 48GB                   | Q4                       |
| DeepSeek-V3             | 671B       | 320GB+     | 640GB                  | FP8                      |
| **DeepSeek-R1**         | **671B**   | **320GB+** | **8x H100**            | **FP8, reasoning model** |
| **DeepSeek-R1-Distill** | **32B**    | **20GB**   | **2x A100 / RTX 5090** | **Q4/Q8**                |

### Image Generation Models

| Model                | Min VRAM | Recommended         | Notes                           |
| -------------------- | -------- | ------------------- | ------------------------------- |
| SD 1.5               | 4GB      | 8GB                 | 512x512 native                  |
| SD 2.1               | 6GB      | 8GB                 | 768x768 native                  |
| SDXL                 | 8GB      | 12GB                | 1024x1024 native                |
| SDXL Turbo           | 8GB      | 12GB                | 1-4 steps                       |
| **SD3.5 Large (8B)** | **16GB** | **24GB**            | **1024x1024, advanced quality** |
| FLUX.1 schnell       | 12GB     | 16GB                | 4 steps, fast                   |
| FLUX.1 dev           | 16GB     | 24GB                | 20-50 steps                     |
| **TRELLIS**          | **16GB** | **24GB (RTX 4090)** | **3D generation from images**   |

### Video Generation Models

| Model                  | Min VRAM | Recommended              | Output                        |
| ---------------------- | -------- | ------------------------ | ----------------------------- |
| Stable Video Diffusion | 16GB     | 24GB                     | 4 sec, 576x1024               |
| AnimateDiff            | 12GB     | 16GB                     | 2-4 sec                       |
| **LTX-Video**          | **16GB** | **24GB (RTX 4090/3090)** | **5 sec, 768x512, very fast** |
| Wan2.1                 | 24GB     | 40GB                     | 5 sec, 480p-720p              |
| Hunyuan Video          | 40GB     | 80GB                     | 5 sec, 720p                   |
| OpenSora               | 24GB     | 40GB                     | Variable                      |

### Audio Models

| Model            | Min VRAM | Recommended | Task             |
| ---------------- | -------- | ----------- | ---------------- |
| Whisper tiny     | 1GB      | 2GB         | Transcription    |
| Whisper base     | 1GB      | 2GB         | Transcription    |
| Whisper small    | 2GB      | 4GB         | Transcription    |
| Whisper medium   | 4GB      | 6GB         | Transcription    |
| Whisper large-v3 | 6GB      | 10GB        | Transcription    |
| Bark             | 8GB      | 12GB        | Text-to-Speech   |
| Stable Audio     | 8GB      | 12GB        | Music Generation |

### Vision & Vision-Language Models

| Model                | Min VRAM | Recommended         | Task                      |
| -------------------- | -------- | ------------------- | ------------------------- |
| Llama 3.2 Vision 11B | 12GB     | 16GB                | Image Understanding       |
| Llama 3.2 Vision 90B | 48GB     | 80GB                | Image Understanding       |
| LLaVA 7B             | 8GB      | 12GB                | Visual QA                 |
| LLaVA 13B            | 16GB     | 24GB                | Visual QA                 |
| **Qwen2.5-VL 7B**    | **16GB** | **24GB (RTX 4090)** | **Image/Video/Doc OCR**   |
| **Qwen2.5-VL 72B**   | **48GB** | **2x A100 80GB**    | **Maximum VL capability** |

### Fine-tuning & Training Tools

| Tool / Method        | Min VRAM | Recommended GPU   | Task                            |
| -------------------- | -------- | ----------------- | ------------------------------- |
| **Unsloth QLoRA 7B** | **12GB** | **RTX 3090 24GB** | **2x faster QLoRA, low VRAM**   |
| Unsloth QLoRA 13B    | 16GB     | RTX 4090 24GB     | Fast fine-tuning                |
| LoRA (standard)      | 12GB     | RTX 3090          | Parameter-efficient fine-tuning |
| Full fine-tune 7B    | 40GB     | A100 40GB         | Maximum quality training        |

***

## Detailed Compatibility Tables

### LLM by GPU

| GPU              | Max Model (Q4) | Max Model (Q8) | Max Model (FP16) |
| ---------------- | -------------- | -------------- | ---------------- |
| RTX 3060 12GB    | 13B            | 7B             | 3B               |
| RTX 3070 8GB     | 7B             | 3B             | 1B               |
| RTX 3080 10GB    | 7B             | 7B             | 3B               |
| RTX 3090 24GB    | 30B            | 13B            | 7B               |
| RTX 4070 Ti 12GB | 13B            | 7B             | 3B               |
| RTX 4080 16GB    | 14B            | 7B             | 7B               |
| RTX 4090 24GB    | 30B            | 13B            | 7B               |
| RTX 5090 32GB    | 70B            | 14B            | 13B              |
| A100 40GB        | 70B            | 30B            | 14B              |
| A100 80GB        | 70B            | 70B            | 30B              |
| H100 80GB        | 70B            | 70B            | 30B              |

### Image Generation by GPU

| GPU              | SD 1.5 | SDXL   | FLUX schnell | FLUX dev |
| ---------------- | ------ | ------ | ------------ | -------- |
| RTX 3060 12GB    | ✅ 512  | ✅ 768  | ⚠️ 512\*     | ❌        |
| RTX 3070 8GB     | ✅ 512  | ⚠️ 512 | ❌            | ❌        |
| RTX 3080 10GB    | ✅ 512  | ✅ 768  | ⚠️ 512\*     | ❌        |
| RTX 3090 24GB    | ✅ 768  | ✅ 1024 | ✅ 1024       | ⚠️ 768\* |
| RTX 4070 Ti 12GB | ✅ 512  | ✅ 768  | ⚠️ 512\*     | ❌        |
| RTX 4080 16GB    | ✅ 768  | ✅ 1024 | ✅ 768        | ⚠️ 512\* |
| RTX 4090 24GB    | ✅ 1024 | ✅ 1024 | ✅ 1024       | ✅ 1024   |
| RTX 5090 32GB    | ✅ 1024 | ✅ 1024 | ✅ 1536       | ✅ 1536   |
| A100 40GB        | ✅ 1024 | ✅ 1024 | ✅ 1024       | ✅ 1024   |
| A100 80GB        | ✅ 2048 | ✅ 2048 | ✅ 1536       | ✅ 1536   |

\*With CPU offload or reduced batch size

### Video Generation by GPU

| GPU           | SVD    | AnimateDiff | Wan2.1  | Hunyuan  |
| ------------- | ------ | ----------- | ------- | -------- |
| RTX 3060 12GB | ❌      | ⚠️ short    | ❌       | ❌        |
| RTX 3090 24GB | ✅ 2-4s | ✅           | ⚠️ 480p | ❌        |
| RTX 4090 24GB | ✅ 4s   | ✅           | ✅ 480p  | ⚠️ short |
| RTX 5090 32GB | ✅ 6s   | ✅           | ✅ 720p  | ✅ 5s     |
| A100 40GB     | ✅ 4s   | ✅           | ✅ 720p  | ✅ 5s     |
| A100 80GB     | ✅ 8s   | ✅           | ✅ 720p  | ✅ 10s    |

***

## Quantization Guide

### What is Quantization?

Quantization reduces model precision to fit in less VRAM:

| Format   | Bits | VRAM Reduction | Quality Loss |
| -------- | ---- | -------------- | ------------ |
| FP32     | 32   | Baseline       | None         |
| FP16     | 16   | 50%            | Minimal      |
| BF16     | 16   | 50%            | Minimal      |
| FP8      | 8    | 75%            | Small        |
| Q8       | 8    | 75%            | Small        |
| Q6\_K    | 6    | 81%            | Small        |
| Q5\_K\_M | 5    | 84%            | Moderate     |
| Q4\_K\_M | 4    | 87%            | Moderate     |
| Q3\_K\_M | 3    | 91%            | Noticeable   |
| Q2\_K    | 2    | 94%            | Significant  |

### VRAM Calculator

**Formula:** `VRAM (GB) ≈ Parameters (B) × Bytes per Parameter`

| Model Size | FP16   | Q8    | Q4     |
| ---------- | ------ | ----- | ------ |
| 1B         | 2 GB   | 1 GB  | 0.5 GB |
| 3B         | 6 GB   | 3 GB  | 1.5 GB |
| 7B         | 14 GB  | 7 GB  | 3.5 GB |
| 8B         | 16 GB  | 8 GB  | 4 GB   |
| 13B        | 26 GB  | 13 GB | 6.5 GB |
| 14B        | 28 GB  | 14 GB | 7 GB   |
| 30B        | 60 GB  | 30 GB | 15 GB  |
| 32B        | 64 GB  | 32 GB | 16 GB  |
| 70B        | 140 GB | 70 GB | 35 GB  |
| 72B        | 144 GB | 72 GB | 36 GB  |

\*Add \~20% for KV cache and overhead

### Recommended Quantization by Use Case

| Use Case         | Recommended | Why                               |
| ---------------- | ----------- | --------------------------------- |
| Chat/General     | Q4\_K\_M    | Good balance of speed and quality |
| Coding           | Q5\_K\_M+   | Better accuracy for code          |
| Creative Writing | Q4\_K\_M    | Speed matters more                |
| Analysis         | Q6\_K+      | Higher precision needed           |
| Production       | FP16/BF16   | Maximum quality                   |

***

## Context Length vs VRAM

### How Context Affects VRAM

Each model has a context window (max tokens). Longer context = more VRAM:

| Model        | Default Context | Max Context | VRAM per 1K tokens |
| ------------ | --------------- | ----------- | ------------------ |
| Llama 3 8B   | 8K              | 128K        | \~0.3 GB           |
| Llama 3 70B  | 8K              | 128K        | \~0.5 GB           |
| Qwen 2.5 7B  | 8K              | 128K        | \~0.25 GB          |
| Mistral 7B   | 8K              | 32K         | \~0.25 GB          |
| Mixtral 8x7B | 32K             | 32K         | \~0.4 GB           |

### Context by GPU (Llama 3 8B Q4)

| GPU           | Comfortable Context | Maximum Context |
| ------------- | ------------------- | --------------- |
| RTX 3060 12GB | 16K                 | 32K             |
| RTX 3090 24GB | 64K                 | 96K             |
| RTX 4090 24GB | 64K                 | 96K             |
| RTX 5090 32GB | 96K                 | 128K            |
| A100 40GB     | 96K                 | 128K            |
| A100 80GB     | 128K                | 128K            |

***

## Multi-GPU Configurations

### Tensor Parallelism

Split one model across multiple GPUs:

| Configuration | Total VRAM | Max Model (FP16) |
| ------------- | ---------- | ---------------- |
| 2x RTX 3090   | 48GB       | 30B              |
| 2x RTX 4090   | 48GB       | 30B              |
| 2x RTX 5090   | 64GB       | 32B              |
| 4x RTX 5090   | 128GB      | 70B              |
| 2x A100 40GB  | 80GB       | 70B              |
| 4x A100 40GB  | 160GB      | 100B+            |
| 8x A100 80GB  | 640GB      | DeepSeek-V3      |

### vLLM Multi-GPU

```bash
# 2 GPUs
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 2

# 4 GPUs
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 4
```

***

## Specific Model Guides

### Llama 3.1 Family

| Variant        | Parameters | Min GPU      | Recommended Setup |
| -------------- | ---------- | ------------ | ----------------- |
| Llama 3.2 1B   | 1B         | Any 4GB      | RTX 3060          |
| Llama 3.2 3B   | 3B         | Any 6GB      | RTX 3060          |
| Llama 3.1 8B   | 8B         | RTX 3060     | RTX 3090          |
| Llama 3.1 70B  | 70B        | A100 40GB    | 2x A100 40GB      |
| Llama 3.1 405B | 405B       | 8x A100 80GB | 8x H100           |

### Mistral/Mixtral Family

| Variant       | Parameters | Min GPU      | Recommended Setup |
| ------------- | ---------- | ------------ | ----------------- |
| Mistral 7B    | 7B         | RTX 3060     | RTX 3090          |
| Mixtral 8x7B  | 46.7B      | RTX 3090     | A100 40GB         |
| Mixtral 8x22B | 141B       | 2x A100 80GB | 4x A100 80GB      |

### Qwen 2.5 Family

| Variant       | Parameters | Min GPU   | Recommended Setup |
| ------------- | ---------- | --------- | ----------------- |
| Qwen 2.5 0.5B | 0.5B       | Any 2GB   | Any 4GB           |
| Qwen 2.5 1.5B | 1.5B       | Any 4GB   | RTX 3060          |
| Qwen 2.5 3B   | 3B         | Any 6GB   | RTX 3060          |
| Qwen 2.5 7B   | 7B         | RTX 3060  | RTX 3090          |
| Qwen 2.5 14B  | 14B        | RTX 3090  | RTX 4090          |
| Qwen 2.5 32B  | 32B        | RTX 4090  | A100 40GB         |
| Qwen 2.5 72B  | 72B        | A100 40GB | A100 80GB         |

### DeepSeek Models

| Variant                          | Parameters | Min GPU           | Recommended Setup |
| -------------------------------- | ---------- | ----------------- | ----------------- |
| DeepSeek-Coder 6.7B              | 6.7B       | RTX 3060          | RTX 3090          |
| DeepSeek-Coder 33B               | 33B        | RTX 4090          | A100 40GB         |
| DeepSeek-V2-Lite                 | 15.7B      | RTX 3090          | A100 40GB         |
| DeepSeek-V3                      | 671B       | 8x A100 80GB      | 8x H100           |
| **DeepSeek-R1**                  | **671B**   | **8x A100 80GB**  | **8x H100 (FP8)** |
| **DeepSeek-R1-Distill-Qwen-32B** | **32B**    | **RTX 5090 32GB** | **2x A100 40GB**  |
| **DeepSeek-R1-Distill-Qwen-7B**  | **7B**     | **RTX 3090 24GB** | **RTX 4090**      |

***

## Troubleshooting

### "CUDA out of memory"

1. **Reduce quantization:** Q8 → Q4
2. **Lower context length:** Reduce max\_tokens
3. **Enable CPU offload:** `--cpu-offload` or `enable_model_cpu_offload()`
4. **Use smaller batch:** batch\_size=1
5. **Try different GPU:** Need more VRAM

### "Model too large"

1. **Use quantized version:** GGUF Q4 models
2. **Use multiple GPUs:** Tensor parallelism
3. **Offload to CPU:** Slower but works
4. **Choose smaller model:** 7B instead of 13B

### "Slow generation"

1. **Upgrade GPU:** More VRAM = less offloading
2. **Use faster quantization:** Q4 is faster than Q8
3. **Reduce context:** Shorter = faster
4. **Enable flash attention:** `--flash-attn`

## Next Steps

* [GPU Comparison Guide](/guides/getting-started/gpu-comparison) - Detailed GPU specs
* [Docker Images Catalog](/guides/getting-started/docker-images) - Ready-to-deploy images
* [Quickstart Guide](/guides/quickstart) - Get started in 5 minutes


# Cost Calculator

Estimate your AI workload costs on Clore.ai

Estimate your AI workload costs on CLORE.AI.

{% hint style="success" %}
Prices are estimates. Check [CLORE.AI Marketplace](https://clore.ai/marketplace) for current rates.
{% endhint %}

## Quick Cost Reference

### GPU Hourly Rates

| GPU              | On-Demand  | Spot (typical) |
| ---------------- | ---------- | -------------- |
| RTX 3060 12GB    | $0.03-0.05 | $0.02-0.03     |
| RTX 3070 8GB     | $0.03-0.04 | $0.02-0.03     |
| RTX 3080 10GB    | $0.04-0.06 | $0.03-0.04     |
| RTX 3090 24GB    | $0.05-0.08 | $0.04-0.06     |
| RTX 4070 Ti 12GB | $0.05-0.07 | $0.03-0.05     |
| RTX 4080 16GB    | $0.07-0.10 | $0.05-0.07     |
| RTX 4090 24GB    | $0.10-0.15 | $0.07-0.10     |
| RTX 5090 32GB    | $0.15-0.20 | $0.10-0.15     |
| A100 40GB        | $0.15-0.22 | $0.12-0.17     |
| A100 80GB        | $0.22-0.35 | $0.18-0.25     |
| H100 80GB        | $0.45-0.70 | $0.35-0.50     |

***

## Cost by Task

### Chat/LLM Inference

**Cost per 1 million tokens generated:**

| Model Size | GPU       | Speed (tok/s) | Cost/1M tokens |
| ---------- | --------- | ------------- | -------------- |
| 7B (Q4)    | RTX 3060  | 25            | $0.33          |
| 7B (Q4)    | RTX 3090  | 45            | $0.37          |
| 7B (Q4)    | RTX 4090  | 80            | $0.35          |
| 13B (Q4)   | RTX 3090  | 30            | $0.56          |
| 13B (Q4)   | RTX 4090  | 55            | $0.51          |
| 30B (Q4)   | RTX 4090  | 25            | $1.11          |
| 70B (Q4)   | A100 40GB | 25            | $1.89          |
| 70B (Q4)   | A100 80GB | 35            | $1.98          |

**Example costs:**

* 1 hour of chatting (\~50K tokens): $0.02-0.10
* Process 100 documents (1M tokens): $0.33-2.00
* Run chatbot for 24 hours: $0.72-4.00

***

### Image Generation

**Cost per 1000 images:**

| Model        | Resolution | GPU       | Time/image | Cost/1000 |
| ------------ | ---------- | --------- | ---------- | --------- |
| SD 1.5       | 512x512    | RTX 3060  | 4 sec      | $0.33     |
| SD 1.5       | 512x512    | RTX 3090  | 2 sec      | $0.28     |
| SD 1.5       | 512x512    | RTX 4090  | 1 sec      | $0.28     |
| SDXL         | 1024x1024  | RTX 3090  | 7 sec      | $0.97     |
| SDXL         | 1024x1024  | RTX 4090  | 3 sec      | $0.83     |
| FLUX schnell | 1024x1024  | RTX 4090  | 5 sec      | $1.39     |
| FLUX schnell | 1024x1024  | RTX 5090  | 3.5 sec    | $1.46     |
| FLUX dev     | 1024x1024  | RTX 4090  | 15 sec     | $4.17     |
| FLUX dev     | 1024x1024  | A100 40GB | 10 sec     | $4.72     |

**Example costs:**

* Generate 100 SDXL images: $0.08-0.10
* Generate 1000 SD 1.5 images: $0.28-0.33
* Batch of 500 FLUX images: $0.70-2.40

***

### Video Generation

**Cost per minute of video:**

| Model       | Length | Resolution | GPU       | Time    | Cost/min video |
| ----------- | ------ | ---------- | --------- | ------- | -------------- |
| SVD         | 4 sec  | 576x1024   | RTX 4090  | 1.5 min | $0.38          |
| SVD         | 4 sec  | 576x1024   | A100 40GB | 1 min   | $0.43          |
| AnimateDiff | 3 sec  | 512x512    | RTX 3090  | 2 min   | $0.53          |
| Wan2.1      | 5 sec  | 480p       | RTX 5090  | 2 min   | $0.80          |
| Wan2.1      | 5 sec  | 720p       | A100 40GB | 2 min   | $0.85          |
| Hunyuan     | 5 sec  | 720p       | A100 80GB | 5 min   | $3.13          |

**Example costs:**

* 1 minute of SVD clips: $0.38-0.50
* 5 minutes of AnimateDiff: $2.65
* 1 minute of Hunyuan video: $3.00-4.00

***

### Audio Processing

**Cost per hour of audio:**

| Task          | Model            | GPU      | Processing Time | Cost/hour audio |
| ------------- | ---------------- | -------- | --------------- | --------------- |
| Transcription | Whisper large-v3 | RTX 3060 | 10 min          | $0.05           |
| Transcription | Whisper large-v3 | RTX 4090 | 3 min           | $0.05           |
| TTS           | Bark             | RTX 3090 | Real-time       | $0.06           |
| TTS           | XTTS             | RTX 3090 | Real-time       | $0.06           |
| Music Gen     | Stable Audio     | RTX 3090 | 2x real-time    | $0.12           |

**Example costs:**

* Transcribe 10 hours of podcasts: $0.50
* Generate 1 hour of TTS: $0.06
* Create 30 minutes of music: $0.06

***

### Model Training

**Cost estimates for common training tasks:**

| Task                 | Dataset     | GPU       | Time        | Total Cost |
| -------------------- | ----------- | --------- | ----------- | ---------- |
| LoRA fine-tune (7B)  | 10K samples | RTX 4090  | 2-4 hours   | $0.20-0.50 |
| LoRA fine-tune (13B) | 10K samples | A100 40GB | 3-6 hours   | $0.50-1.00 |
| Full fine-tune (7B)  | 50K samples | A100 80GB | 24-48 hours | $5-12      |
| SD LoRA              | 100 images  | RTX 3090  | 1-2 hours   | $0.05-0.12 |
| DreamBooth           | 20 images   | RTX 4090  | 30-60 min   | $0.05-0.10 |

***

## Cost Calculators

### LLM Inference Calculator

```
Tokens to generate: ________
Model size: 7B / 13B / 30B / 70B
GPU choice: ________

Formula:
Time (hours) = Tokens ÷ (tokens_per_second × 3600)
Cost = Time × hourly_rate

Example: 1M tokens with 70B on A100 40GB
Time = 1,000,000 ÷ (25 × 3600) = 11.1 hours
Cost = 11.1 × $0.17 = $1.89
```

### Image Generation Calculator

```
Number of images: ________
Model: SD 1.5 / SDXL / FLUX
GPU choice: ________

Formula:
Time (hours) = Images × seconds_per_image ÷ 3600
Cost = Time × hourly_rate

Example: 1000 SDXL images on RTX 4090
Time = 1000 × 3 ÷ 3600 = 0.83 hours
Cost = 0.83 × $0.10 = $0.083
```

### Video Generation Calculator

```
Video length (seconds): ________
Model: SVD / AnimateDiff / Wan2.1
GPU choice: ________

Formula:
Clips needed = Total seconds ÷ clip_length
Time (hours) = Clips × generation_time ÷ 60
Cost = Time × hourly_rate

Example: 60 seconds with SVD (4-sec clips) on RTX 4090
Clips = 60 ÷ 4 = 15 clips
Time = 15 × 1.5 ÷ 60 = 0.375 hours
Cost = 0.375 × $0.10 = $0.04
```

***

## Budget Planning

### Small Project (\~$5)

What you can do:

* Generate 5,000+ SD 1.5 images
* Generate 500+ SDXL images
* Process 100+ hours of audio transcription
* Run 7B chatbot for 100+ hours
* Train 5+ LoRA models

### Medium Project (\~$25)

What you can do:

* Generate 25,000+ SDXL images
* Generate 5,000+ FLUX images
* Process 500+ hours of audio
* Run 70B model for 100+ hours
* Train multiple custom models
* Generate 30+ minutes of video

### Large Project (\~$100)

What you can do:

* Full fine-tune 7B model
* Generate 100,000+ images
* Run production LLM API for a week
* Generate hours of video content
* Train multiple DreamBooth models

***

## Cost Optimization Tips

### 1. Use Spot Orders

Save 30-50% with spot pricing:

* Great for batch jobs
* Save work frequently
* Avoid for time-critical tasks

### 2. Right-Size Your GPU

Don't overpay for unused power:

| If you need... | Don't use | Use instead |
| -------------- | --------- | ----------- |
| 7B chat        | A100      | RTX 3060    |
| SD 1.5 images  | RTX 4090  | RTX 3060    |
| SDXL images    | A100      | RTX 3090    |
| Quick tests    | A100      | RTX 3060    |

### 3. Use Quantization

Reduce GPU requirements:

* Q4: 4x less VRAM, slightly lower quality
* Q8: 2x less VRAM, minimal quality loss
* AWQ/GPTQ: Optimized for inference

### 4. Batch Processing

Process more in less time:

* Batch API requests
* Use larger batch sizes
* Run overnight jobs

### 5. Pre-download Models

Avoid paying while downloading:

* Use Docker images with pre-loaded models
* Cache models on persistent storage
* Download before starting paid order

***

## Comparison: CLORE.AI vs Alternatives

### vs Cloud Providers (AWS, GCP, Azure)

| Factor             | CLORE.AI   | Cloud Providers    |
| ------------------ | ---------- | ------------------ |
| A100 hourly        | $0.17-0.25 | $3-5+              |
| Setup time         | Minutes    | Hours              |
| Minimum commitment | None       | Often hourly       |
| Spot savings       | 30-50%     | 60-90% but complex |

**CLORE.AI advantage:** 10-20x cheaper for most AI workloads.

### vs RunPod, Vast.ai

| Factor          | CLORE.AI     | Others  |
| --------------- | ------------ | ------- |
| Price           | Competitive  | Similar |
| GPU variety     | Wide         | Wide    |
| Spot orders     | Yes          | Yes     |
| Crypto payments | Yes (native) | Limited |

### vs Local Hardware

| Factor            | CLORE.AI | Local RTX 4090      |
| ----------------- | -------- | ------------------- |
| Upfront cost      | $0       | $1,600+             |
| Monthly (8hr/day) | \~$24    | \~$15 (electricity) |
| Break-even        | -        | \~3 months          |
| Flexibility       | Any GPU  | One GPU             |
| Maintenance       | None     | You                 |

**When to rent:** Testing, variable workloads, need for high-end GPUs **When to buy:** Consistent daily use, 4+ hours/day for 3+ months

***

## Real-World Examples

### Startup: AI Chatbot

**Requirements:** 70B model, 8 hours/day, 30 days

```
GPU: A100 40GB
Hours: 8 × 30 = 240 hours
Cost: 240 × $0.17 = $40.80/month

With spot: ~$28/month
```

### Creator: Image Generation

**Requirements:** 1000 SDXL images/day, 30 days

```
GPU: RTX 4090
Time per day: 1000 × 3 sec = 50 minutes
Monthly hours: 25 hours
Cost: 25 × $0.10 = $2.50/month
```

### Researcher: Model Training

**Requirements:** Fine-tune 13B model on 100K samples

```
GPU: A100 80GB
Estimated time: 48 hours
Cost: 48 × $0.25 = $12.00

With spot: ~$8.00
```

### Agency: Video Production

**Requirements:** 10 minutes of AI video/week

```
GPU: A100 40GB
Time per minute: ~5 minutes generation
Weekly: 10 × 5 = 50 minutes
Monthly: 200 minutes = 3.3 hours
Cost: 3.3 × $0.17 = $0.56/week = $2.24/month
```

***

## Next Steps

* [GPU Comparison](/guides/getting-started/gpu-comparison) - Detailed GPU specs
* [Model Compatibility](/guides/getting-started/model-compatibility) - VRAM requirements
* [Quickstart Guide](/guides/quickstart) - Get started now


# Docker Images

Ready-to-deploy Docker images for AI workloads on Clore.ai

Ready-to-deploy Docker images for AI workloads on CLORE.AI.

{% hint style="success" %}
Deploy these images directly at [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Quick Deploy Reference

### Most Popular

| Task                 | Image                                | Ports     |
| -------------------- | ------------------------------------ | --------- |
| Chat with AI         | `ollama/ollama`                      | 22, 11434 |
| ChatGPT-like UI      | `ghcr.io/open-webui/open-webui`      | 22, 8080  |
| Image Generation     | `universonic/stable-diffusion-webui` | 22, 7860  |
| Node-based Image Gen | `yanwk/comfyui-boot`                 | 22, 8188  |
| LLM API Server       | `vllm/vllm-openai`                   | 22, 8000  |

***

## Language Models

### Ollama

**Universal LLM runner - easiest way to run any model.**

```
Image: ollama/ollama
Ports: 22/tcp, 11434/http
Command: ollama serve
```

**After deploy:**

```bash
# SSH into server
ssh -p <port> root@<proxy>

# Pull and run a model
ollama pull llama3.2
ollama run llama3.2
```

**Environment variables:**

```
OLLAMA_HOST=0.0.0.0
OLLAMA_MODELS=/root/.ollama/models
```

***

### Open WebUI

**ChatGPT-like interface for Ollama.**

```
Image: ghcr.io/open-webui/open-webui:ollama
Ports: 22/tcp, 8080/http
```

Includes Ollama built-in. Access via HTTP port.

**Standalone (connect to existing Ollama):**

```
Image: ghcr.io/open-webui/open-webui:main
Ports: 22/tcp, 8080/http
Environment: OLLAMA_BASE_URL=http://localhost:11434
```

***

### vLLM

**High-performance LLM serving with OpenAI-compatible API.**

```
Image: vllm/vllm-openai:latest
Ports: 22/tcp, 8000/http
Command: python -m vllm.entrypoints.openai.api_server --model meta-llama/Meta-Llama-3.1-8B-Instruct --host 0.0.0.0
```

**For larger models (multi-GPU):**

```bash
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3.1-70B-Instruct \
    --tensor-parallel-size 2 \
    --host 0.0.0.0
```

**Environment variables:**

```
HUGGING_FACE_HUB_TOKEN=<your-token>  # For gated models
```

***

### Text Generation Inference (TGI)

**HuggingFace's production LLM server.**

```
Image: ghcr.io/huggingface/text-generation-inference:latest
Ports: 22/tcp, 8080/http
Command: --model-id meta-llama/Meta-Llama-3.1-8B-Instruct
```

**Environment variables:**

```
HUGGING_FACE_HUB_TOKEN=<your-token>
MAX_INPUT_LENGTH=4096
MAX_TOTAL_TOKENS=8192
```

***

## Image Generation

### Stable Diffusion WebUI (AUTOMATIC1111)

**Most popular SD interface with extensions.**

```
Image: universonic/stable-diffusion-webui:latest
Ports: 22/tcp, 7860/http
```

**For low VRAM (8GB or less):**

```bash
./webui.sh --listen --medvram --xformers
```

**For API access:**

```bash
./webui.sh --listen --xformers --api
```

***

### ComfyUI

**Node-based workflow for advanced users.**

```
Image: yanwk/comfyui-boot:cu126-slim
Ports: 22/tcp, 8188/http
Environment: CLI_ARGS=--listen 0.0.0.0
```

**Alternative images:**

```
# With common extensions
Image: ai-dock/comfyui:latest

# Minimal
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
```

**Manual setup command:**

```bash
git clone https://github.com/comfyanonymous/ComfyUI && cd ComfyUI && pip install -r requirements.txt && python main.py --listen 0.0.0.0
```

***

### Fooocus

**Simplified SD interface, Midjourney-like.**

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp, 7865/http
Command: git clone https://github.com/lllyasviel/Fooocus && cd Fooocus && pip install -r requirements.txt && python launch.py --listen
```

***

### FLUX

**Latest high-quality image generation.**

Use ComfyUI with FLUX nodes:

```
Image: yanwk/comfyui-boot:cu126-slim
Ports: 22/tcp, 8188/http
```

Or via Diffusers:

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```python
# After SSH
pip install diffusers transformers accelerate
python << 'EOF'
from diffusers import FluxPipeline
pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-schnell")
pipe.enable_model_cpu_offload()
image = pipe("A cat", num_inference_steps=4).images[0]
image.save("output.png")
EOF
```

***

## Video Generation

### Stable Video Diffusion

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```bash
pip install diffusers transformers accelerate
python << 'EOF'
from diffusers import StableVideoDiffusionPipeline
from diffusers.utils import load_image, export_to_video
pipe = StableVideoDiffusionPipeline.from_pretrained(
    "stabilityai/stable-video-diffusion-img2vid-xt",
    variant="fp16"
)
pipe.to("cuda")
image = load_image("input.png")
frames = pipe(image, num_frames=25).frames[0]
export_to_video(frames, "output.mp4", fps=7)
EOF
```

***

### AnimateDiff

Use with ComfyUI:

```
Image: yanwk/comfyui-boot:cu126-slim
Ports: 22/tcp, 8188/http
```

Install AnimateDiff nodes via ComfyUI Manager.

***

## Audio & Voice

### Whisper (Transcription)

```
Image: onerahmet/openai-whisper-asr-webservice:latest
Ports: 22/tcp, 9000/http
Environment: ASR_MODEL=large-v3
```

**API usage:**

```bash
curl -X POST "http://localhost:9000/asr" \
    -F "audio_file=@audio.mp3" \
    -F "task=transcribe"
```

***

### Bark (Text-to-Speech)

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```bash
pip install bark
python << 'EOF'
from bark import SAMPLE_RATE, generate_audio, preload_models
from scipy.io.wavfile import write as write_wav
preload_models()
audio = generate_audio("Hello, this is a test.")
write_wav("output.wav", SAMPLE_RATE, audio)
EOF
```

***

### Stable Audio

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```bash
pip install stable-audio-tools
# Requires HF token for model access
```

***

## Vision Models

### LLaVA

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```bash
pip install llava
python -m llava.serve.cli --model-path liuhaotian/llava-v1.6-34b
```

***

### Llama 3.2 Vision

Use Ollama:

```
Image: ollama/ollama
Ports: 22/tcp, 11434/http
```

```bash
ollama pull llama3.2-vision
ollama run llama3.2-vision "describe this image" --images photo.jpg
```

***

## Development & Training

### PyTorch Base

**For custom setups and training.**

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

Includes: CUDA 12.1, cuDNN 8, PyTorch 2.1

***

### Jupyter Lab

**Interactive notebooks for ML.**

```
Image: jupyter/pytorch-notebook:cuda12-pytorch-2.1
Ports: 22/tcp, 8888/http
```

Or use PyTorch base with Jupyter:

```bash
pip install jupyterlab
jupyter lab --ip=0.0.0.0 --allow-root --no-browser
```

***

### Kohya Training

**For LoRA and model fine-tuning.**

```
Image: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
Ports: 22/tcp
```

```bash
git clone https://github.com/kohya-ss/sd-scripts
cd sd-scripts
pip install -r requirements.txt
# Use training scripts
```

***

## Base Images Reference

### NVIDIA Official

| Image                                    | CUDA | Use Case             |
| ---------------------------------------- | ---- | -------------------- |
| `nvidia/cuda:12.1.0-devel-ubuntu22.04`   | 12.1 | CUDA development     |
| `nvidia/cuda:12.1.0-runtime-ubuntu22.04` | 12.1 | CUDA runtime only    |
| `nvidia/cuda:11.8.0-devel-ubuntu22.04`   | 11.8 | Legacy compatibility |

### PyTorch Official

| Image                                          | PyTorch | CUDA |
| ---------------------------------------------- | ------- | ---- |
| `pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel`  | 2.5     | 12.4 |
| `pytorch/pytorch:2.0.1-cuda11.7-cudnn8-devel`  | 2.0     | 11.7 |
| `pytorch/pytorch:1.13.1-cuda11.6-cudnn8-devel` | 1.13    | 11.6 |

### HuggingFace

| Image                                           | Purpose                |
| ----------------------------------------------- | ---------------------- |
| `huggingface/transformers-pytorch-gpu`          | Transformers + PyTorch |
| `ghcr.io/huggingface/text-generation-inference` | TGI server             |

***

## Environment Variables

### Common Variables

| Variable                 | Description                   | Example        |
| ------------------------ | ----------------------------- | -------------- |
| `HUGGING_FACE_HUB_TOKEN` | HF API token for gated models | `hf_xxx`       |
| `CUDA_VISIBLE_DEVICES`   | GPU selection                 | `0,1`          |
| `TRANSFORMERS_CACHE`     | Model cache directory         | `/root/.cache` |

### Ollama Variables

| Variable              | Description       | Default            |
| --------------------- | ----------------- | ------------------ |
| `OLLAMA_HOST`         | Bind address      | `127.0.0.1`        |
| `OLLAMA_MODELS`       | Models directory  | `~/.ollama/models` |
| `OLLAMA_NUM_PARALLEL` | Parallel requests | `1`                |

### vLLM Variables

| Variable                 | Description                  |
| ------------------------ | ---------------------------- |
| `VLLM_ATTENTION_BACKEND` | Attention implementation     |
| `VLLM_USE_MODELSCOPE`    | Use ModelScope instead of HF |

***

## Port Reference

| Port  | Protocol | Service                    |
| ----- | -------- | -------------------------- |
| 22    | TCP      | SSH                        |
| 7860  | HTTP     | Gradio (SD WebUI, Fooocus) |
| 7865  | HTTP     | Fooocus alternative        |
| 8000  | HTTP     | vLLM API                   |
| 8080  | HTTP     | Open WebUI, TGI            |
| 8188  | HTTP     | ComfyUI                    |
| 8888  | HTTP     | Jupyter                    |
| 9000  | HTTP     | Whisper API                |
| 11434 | TCP      | Ollama API                 |

***

## Tips

### Persistent Storage

Mount volumes to keep data between restarts:

```bash
docker run -v /data/models:/root/.cache/huggingface ...
```

### GPU Selection

For multi-GPU systems:

```bash
docker run --gpus '"device=0,1"' ...
# or
CUDA_VISIBLE_DEVICES=0,1
```

### Memory Management

If running out of VRAM:

1. Use smaller models
2. Enable CPU offload
3. Reduce batch size
4. Use quantized models (GGUF Q4)

## Next Steps

* [GPU Comparison](/guides/getting-started/gpu-comparison) - Choose the right GPU
* [Model Compatibility](/guides/getting-started/model-compatibility) - What runs where
* [Quickstart Guide](/guides/quickstart) - Get started in 5 minutes


# GPU Pricing

Current GPU rental pricing on Clore.ai marketplace

This document contains current GPU rental prices for cost estimates in tutorials.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## GPU Prices (as of February 2026)

> Prices reflect current Clore.ai marketplace rates. Spot instances may be 20-40% cheaper. Pay with CLORE tokens for best value.

| GPU           | $/hour (range) | $/day (est.)  | Best For                          |
| ------------- | -------------- | ------------- | --------------------------------- |
| RTX 3060 12GB | $0.02-0.04     | $0.50-1.00    | Light inference, development      |
| RTX 3070 8GB  | $0.02-0.04     | $0.50-1.00    | Light inference                   |
| RTX 3080 10GB | $0.04-0.08     | $1.00-2.00    | Image generation, inference       |
| RTX 3090 24GB | **$0.15-0.25** | $3.60-6.00    | Training, large models, SDXL      |
| RTX 4070 12GB | $0.08-0.15     | $2.00-3.60    | Inference, image gen              |
| RTX 4080 16GB | $0.15-0.25     | $3.60-6.00    | Image gen, video, 13B LLMs        |
| RTX 4090 24GB | **$0.35-0.55** | $8.40-13.20   | Production, FLUX, 30B models      |
| RTX 5080 16GB | $1.50-2.00     | $36.00-48.00  | Fast FLUX, 13B-30B LLMs (new)     |
| RTX 5090 32GB | **$3.00-4.50** | $72.00-108.00 | Flagship, 70B models, video (new) |
| A100 40GB     | $0.80-1.20     | $19.20-28.80  | Enterprise training               |
| A100 80GB     | **$1.20-1.80** | $28.80-43.20  | 70B models, large training        |
| H100 80GB     | **$2.50-3.50** | $60.00-84.00  | Maximum performance               |
| 4x A100 80GB  | **$4.50-6.00** | $108-144.00   | Multi-GPU: 100B+ models           |

## Cost Estimate Template

When writing cost estimates for tutorials, use this format:

```markdown
## Cost Estimate

| Task | GPU | Duration | Cost |
|------|-----|----------|------|
| Testing | RTX 3090 | 30 min | ~$0.01 |
| Batch job | RTX 4090 | 2 hours | ~$0.08 |
| Production | A100 | 1 day | ~$2.00 |
```

## Example Calculations

### Image Generation (1000 images)

* RTX 3090: \~2 hours = $0.04
* RTX 4090: \~1.5 hours = $0.06
* A100: \~1 hour = $0.08

### LLM Inference (1M tokens)

* RTX 3090: \~30 min = $0.01
* RTX 4090: \~20 min = $0.015
* A100: \~10 min = $0.015

### Video Generation (10 videos, 6s each)

* RTX 4090: \~20 min = $0.015
* A100: \~15 min = $0.02

### Model Training (1 epoch, medium dataset)

* RTX 3090: \~4 hours = $0.08
* RTX 4090: \~3 hours = $0.12
* A100: \~2 hours = $0.16

## Notes

* Prices may vary based on availability and demand
* Spot instances may be cheaper
* Multi-GPU setups have different pricing
* Check [clore.ai](https://clore.ai/) for current prices

> 💡 See also: [GPU Cloud Pricing Comparison — Clore.ai vs Major Providers](https://blog.clore.ai/gpu-cloud-pricing-comparison/)


# Troubleshooting

Common issues and solutions for Clore.ai GPU rentals

Common issues and solutions when renting GPU servers on CLORE.AI marketplace.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

{% hint style="info" %}
This guide is based on CLORE.AI platform technical documentation.
{% endhint %}

## Table of Contents

* [Order Creation Issues](#order-creation-issues)
* [Connection Issues](#connection-issues)
* [Container Issues](#container-issues)
* [GPU Issues](#gpu-issues)
* [Payment Issues](#payment-issues)
* [Platform Limits](#platform-limits)

***

## Order Creation Issues

### Order fails: "Insufficient balance"

**Cause:** Not enough funds to cover creation fee and minimum deposit.

**Solution:**

* Check your balance in the selected currency (CLORE, BTC, or USDT/USDC)
* Creation fee is charged when order is created
* Top up your balance with enough for several hours of rental

### Order fails: "Server not available"

**Cause:** Server is already rented or offline.

**Solution:**

* Refresh the marketplace page
* Check server status (online/offline indicator)
* For Spot rentals - you may have been outbid

### Order stuck in "Creating" status

**Cause:** Container is deploying or an error occurred.

**Solution:**

1. Wait 2-5 minutes (Docker image is being pulled)
2. Check logs in **My Orders**
3. Large images (10GB+) take longer to download
4. If stuck for more than 10 minutes - cancel and retry

***

## Connection Issues

### Cannot connect via SSH

**Cause:** Port not configured or container not ready.

**Checklist:**

1. Port 22 must be set as **TCP** (not HTTP)
2. Container status must be **Active** (not Creating)
3. Use the correct mapped port from **My Orders**

**Correct SSH command:**

```bash
ssh -p <MAPPED_PORT> root@<PROXY_ADDRESS>
```

Where `<MAPPED_PORT>` is the public port (e.g., 45678), NOT port 22.

### SSH works but web interface doesn't open

**Cause:** Port set as TCP instead of HTTP, or service not running.

**Solution:**

1. Web interface ports must be set as **HTTP** (not TCP)
2. Service must listen on `0.0.0.0`, not `localhost`
3. Check logs - service may have crashed on startup

**Correct port configuration:**

```
22/tcp      - SSH access
7860/http   - Gradio/WebUI interface
8000/http   - API server
```

### "Connection refused" error

**Cause:** Service inside container is not running or listening on wrong address.

**Solution:**

1. SSH into container and check service status:

   ```bash
   ps aux | grep python
   netstat -tlnp
   ```
2. Service must listen on `0.0.0.0`, not `127.0.0.1`:

   ```bash
   # Wrong:
   python app.py --host 127.0.0.1

   # Correct:
   python app.py --host 0.0.0.0
   ```

### "Connection timed out" error

**Cause:** Wrong address/port or network issues.

**Checklist:**

1. Use Proxy address from **My Orders** (not server IP!)
2. Use Mapped port (public port, not container port)
3. Use correct protocol (http\:// for HTTP ports)

***

## Container Issues

### Container keeps restarting

**Cause:** Error in startup command or insufficient resources.

**Solution:**

1. Check logs in **My Orders**
2. Simplify startup command:

   ```bash
   # Bad - long command may fail:
   apt update && apt install -y ... && pip install ... && python ...

   # Better - start with simple command:
   sleep infinity
   ```
3. Then SSH in and configure manually

### Cannot reset container

**Cause:** Cooldown period between resets.

**Fact:** Reset container has a **120 second** cooldown.

**Solution:** Wait 2 minutes between reset attempts.

### Data lost after restart

**Cause:** Data not in persistent storage.

**Important:**

* Data inside container is **preserved** on Reset Container
* Data is **lost** when order is cancelled or expires
* Always download results before ending rental:

  ```bash
  scp -P <port> root@<proxy>:/workspace/results.tar.gz ./
  ```

### Startup command not executing

**Cause:** Syntax error or image issue.

**Common mistakes:**

```bash

# Error: extra space after \
apt update && \
apt install -y git   # <-- space before next line

# Correct:
apt update && \
apt install -y git && \
python app.py
```

**Solution:**

1. Use simple startup: `bash` or `sleep infinity`
2. Configure everything via SSH
3. Or create custom Docker image with pre-installed software

***

## GPU Issues

### GPU not visible in container

**Check:**

```bash
nvidia-smi
```

**If command not found:**

* Docker image must support CUDA
* Use CUDA-enabled images: `pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime`

**If GPU not displayed:**

* Verify server has GPU (check marketplace listing)
* Contact server provider

### CUDA version mismatch

**Error:** `CUDA driver version is insufficient for CUDA runtime version`

**Cause:** CUDA version in image incompatible with server driver.

**Solution:**

* Check driver version: `nvidia-smi` (top right corner)
* Use image with compatible CUDA version
* Safe choices: CUDA 11.8, CUDA 12.1

### Out of GPU memory

**Error:** `CUDA out of memory`

**Solutions:**

1. Use smaller model or quantization
2. Add memory optimization flags:
   * Stable Diffusion: `--medvram` or `--lowvram`
   * LLMs: `load_in_4bit=True` or `load_in_8bit=True`
3. Clear memory: `torch.cuda.empty_cache()`
4. Rent server with more VRAM

***

## Payment Issues

### Supported currencies

CLORE.AI supports three currencies:

* **CLORE** - platform's native token
* **BTC** - Bitcoin
* **USD** - stablecoins (if enabled by provider)

### Order cancelled: "Outbid"

**Cause:** Someone offered higher price on Spot market.

**Solution:**

* Use **On-Demand** for guaranteed rental
* Or increase your Spot bid price

### Balance charged but order not created

**Cause:** Creation fee is charged even if order fails.

**Solution:**

* Creation fee is usually minimal
* Check cancellation reason in history
* Contact support for recurring issues

***

## Platform Limits

Verified from CLORE.AI codebase:

| Parameter                   | Limit                        |
| --------------------------- | ---------------------------- |
| Ports per order             | **5**                        |
| Total environment variables | **12,288 characters** (12KB) |
| Single env var name         | 128 characters               |
| Single env var value        | 1,536 characters             |
| SSH key                     | **3,072 characters**         |
| SSH password                | **32 characters**            |
| Jupyter token               | **32 characters**            |
| Container reset cooldown    | **120 seconds**              |
| Port range                  | 1-65535                      |
| Port protocols              | TCP or HTTP only             |

***

## Environment Variables

Use environment variables for SSH and Jupyter access:

| Variable        | Purpose                | Max Length  |
| --------------- | ---------------------- | ----------- |
| `SSH_KEY`       | Your public SSH key    | 3,072 chars |
| `SSH_PASSWORD`  | SSH password           | 32 chars    |
| `JUPYTER_TOKEN` | Jupyter notebook token | 32 chars    |

**Example configuration:**

```
SSH_PASSWORD=mypassword123
JUPYTER_TOKEN=mysecrettoken
```

***

## Diagnostic Commands

```bash

# Check GPU
nvidia-smi

# Check memory usage
free -h

# Check disk space
df -h

# Check running processes
ps aux | grep python

# Check open ports
netstat -tlnp

# Check recent error logs
dmesg | tail -50

# Clear GPU memory (Python)
import torch
torch.cuda.empty_cache()
```

***

## Getting Help

If issue persists:

1. Check [CLORE.AI Documentation](https://docs.clore.ai/)
2. Describe issue with logs and screenshots
3. Include order ID and server ID


# Overview

Run large language models (LLMs) on CLORE.AI GPUs for inference and chat applications.

## Popular Tools

| Tool                                                                   | Use Case                           | Difficulty |
| ---------------------------------------------------------------------- | ---------------------------------- | ---------- |
| [Ollama](/guides/language-models/ollama)                               | Easiest LLM setup                  | Beginner   |
| [Open WebUI](/guides/language-models/open-webui)                       | ChatGPT-like interface             | Beginner   |
| [vLLM](/guides/language-models/vllm)                                   | High-throughput production serving | Medium     |
| [Llama.cpp Server](/guides/language-models/llamacpp-server)            | Efficient GGUF inference           | Easy       |
| [Text Generation WebUI](/guides/language-models/text-generation-webui) | Full-featured chat UI              | Easy       |
| [ExLlamaV2](/guides/language-models/exllamav2-fast)                    | Fastest EXL2 inference             | Medium     |
| [LocalAI](/guides/language-models/localai-openai-compatible)           | OpenAI-compatible API              | Medium     |
| [SGLang](/guides/language-models/sglang)                               | Fast structured generation         | Medium     |
| [Text Generation Inference (TGI)](/guides/language-models/tgi)         | HuggingFace serving solution       | Medium     |
| [LMDeploy](/guides/language-models/lmdeploy)                           | MMlab serving toolkit              | Medium     |
| [Aphrodite Engine](/guides/language-models/aphrodite-engine)           | vLLM fork with extra features      | Medium     |
| [MLC-LLM](/guides/language-models/mlc-llm)                             | Machine learning compilation       | Hard       |
| [LiteLLM](/guides/language-models/litellm)                             | Unified API proxy                  | Medium     |
| [PowerInfer](/guides/language-models/powerinfer)                       | Sparse model inference             | Hard       |
| [Mistral.rs](/guides/language-models/mistral-rs)                       | Rust-based inference engine        | Medium     |

## Model Guides

### Latest & Best Models

| Model                                              | Parameters | Best For                  |
| -------------------------------------------------- | ---------- | ------------------------- |
| [DeepSeek-V3](/guides/language-models/deepseek-v3) | 671B MoE   | Reasoning, code, math     |
| [DeepSeek-R1](/guides/language-models/deepseek-r1) | 671B MoE   | Advanced reasoning        |
| [DeepSeek V4](/guides/language-models/deepseek-v4) | TBA        | Next-generation DeepSeek  |
| [Qwen2.5](/guides/language-models/qwen25)          | 0.5B-72B   | Multilingual, code        |
| [Qwen3.5](/guides/language-models/qwen35)          | TBA        | Latest Qwen generation    |
| [Llama 3.3](/guides/language-models/llama33)       | 70B        | Meta's latest 70B         |
| [Llama 4](/guides/language-models/llama4)          | TBA        | Scout & Maverick variants |

### Specialized Models

| Model                                                    | Parameters | Best For                |
| -------------------------------------------------------- | ---------- | ----------------------- |
| [DeepSeek Coder](/guides/language-models/deepseek-coder) | 6.7B-33B   | Code generation         |
| [CodeLlama](/guides/language-models/codellama)           | 7B-34B     | Code completion         |
| [GLM-4.7-Flash](/guides/language-models/glm-47-flash)    | 4.7B       | Fast Chinese/English    |
| [GLM-5](/guides/language-models/glm5)                    | TBA        | Zhipu AI latest         |
| [Kimi K2.5](/guides/language-models/kimi-k2)             | TBA        | Moonshot AI model       |
| [Ling-2.5-1T](/guides/language-models/ling25)            | 1T         | Massive open-source LLM |
| [LFM2-24B](/guides/language-models/lfm2-24b)             | 24B        | Liquid AI model         |
| [MiMo-V2-Flash](/guides/language-models/mimo-v2-flash)   | TBA        | Fast inference model    |

### Efficient Models

| Model                                                      | Parameters | Best For                  |
| ---------------------------------------------------------- | ---------- | ------------------------- |
| [Gemma 2](/guides/language-models/gemma2)                  | 2B-27B     | Efficient inference       |
| [Gemma 3](/guides/language-models/gemma3)                  | TBA        | Google's latest compact   |
| [Phi-4](/guides/language-models/phi4)                      | 14B        | Small but capable         |
| [Mistral/Mixtral](/guides/language-models/mistral-mixtral) | 7B / 8x7B  | General purpose           |
| [Mistral Large 3](/guides/language-models/mistral-large3)  | 675B MoE   | Enterprise-grade          |
| [Mistral Small 3.1](/guides/language-models/mistral-small) | TBA        | Efficient Mistral variant |

## GPU Recommendations

| Model Size | Minimum GPU   | Recommended |
| ---------- | ------------- | ----------- |
| 7B (Q4)    | RTX 3060 12GB | RTX 3090    |
| 13B (Q4)   | RTX 3090 24GB | RTX 4090    |
| 34B (Q4)   | 2x RTX 3090   | A100 40GB   |
| 70B (Q4)   | A100 80GB     | 2x A100     |

## Quantization Guide

| Format   | VRAM Usage | Quality   | Speed   |
| -------- | ---------- | --------- | ------- |
| Q2\_K    | Lowest     | Poor      | Fastest |
| Q4\_K\_M | Low        | Good      | Fast    |
| Q5\_K\_M | Medium     | Great     | Medium  |
| Q8\_0    | High       | Excellent | Slower  |
| FP16     | Highest    | Best      | Slowest |

## See Also

* [Training & Fine-tuning](/guides/training/training)
* [Vision-Language Models](/guides/vision-models/vision-models)


# Ollama

Run LLMs locally with Ollama on Clore.ai GPUs

The easiest way to run LLMs locally on CLORE.AI GPUs.

{% hint style="info" %}
**Current Version: v0.6+** — This guide covers Ollama v0.6 and later. Key new features include structured outputs (JSON schema enforcement), OpenAI-compatible embeddings endpoint (`/api/embed`), and concurrent model loading (run multiple models simultaneously without swapping). See [New in v0.6+](#new-in-v06) for details.
{% endhint %}

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Server Requirements

| Parameter    | Minimum      | Recommended |
| ------------ | ------------ | ----------- |
| RAM          | 8GB          | 16GB+       |
| VRAM         | 6GB          | 8GB+        |
| Network      | 100Mbps      | 500Mbps+    |
| Startup Time | \~30 seconds | -           |

{% hint style="info" %}
Ollama is lightweight and works on most GPU servers. For larger models (13B+), choose servers with 16GB+ RAM and 12GB+ VRAM.
{% endhint %}

## Why Ollama?

* **One-command setup** - No Python, no dependencies
* **Model library** - Download models with `ollama pull`
* **OpenAI-compatible API** - Drop-in replacement
* **GPU acceleration** - Automatic CUDA detection
* **Multi-model** - Run multiple models simultaneously (v0.6+)

## Quick Deploy on CLORE.AI

**Docker Image:**

```
ollama/ollama
```

**Ports:**

```
22/tcp
11434/http
```

**Command:**

```bash
ollama serve
```

### Verify It's Working

After deployment, find your `http_pub` URL in **My Orders** and test:

```bash
# Replace with your actual http_pub URL
curl https://your-http-pub.clorecloud.net/

# Expected response: "Ollama is running"
```

{% hint style="warning" %}
If you get HTTP 502, wait 30-60 seconds - the service is still starting.
{% endhint %}

## Accessing Your Service

When deployed on CLORE.AI, access your Ollama instance via the `http_pub` URL:

```bash
# Find your http_pub in My Orders, then:
curl https://your-http-pub.clorecloud.net/api/tags

# For API calls, use your http_pub URL:
curl https://your-http-pub.clorecloud.net/api/chat -d '{
  "model": "llama3.2",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": false
}'
```

{% hint style="info" %}
All `localhost:11434` examples below work when connected via SSH. For external access, replace with your `https://your-http-pub.clorecloud.net/` URL.
{% endhint %}

## Installation

### Using Docker (Recommended)

```bash
docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
```

### Manual Installation

```bash
curl -fsSL https://ollama.com/install.sh | sh
```

This single command installs the latest version of Ollama, sets up the systemd service, and configures GPU detection automatically. Works on Ubuntu, Debian, Fedora, and most modern Linux distributions.

## Running Models

### Pull and Run

```bash
# Pull model
ollama pull llama3.2

# Run interactive chat
ollama run llama3.2

# Run with prompt
ollama run llama3.2 "Explain quantum computing"
```

### Popular Models

| Model               | Size    | Use Case              |
| ------------------- | ------- | --------------------- |
| `llama3.2`          | 3B      | Fast, general purpose |
| `llama3.1`          | 8B      | Better quality        |
| `llama3.1:70b`      | 70B     | Best quality          |
| `mistral`           | 7B      | Fast, good quality    |
| `mixtral`           | 47B     | MoE, high quality     |
| `codellama`         | 7-34B   | Code generation       |
| `deepseek-coder-v2` | 16B     | Best for code         |
| `deepseek-r1`       | 7B-671B | Reasoning model       |
| `deepseek-r1:32b`   | 32B     | Balanced reasoning    |
| `qwen2.5`           | 7B      | Multilingual          |
| `qwen2.5:72b`       | 72B     | Best Qwen quality     |
| `phi4`              | 14B     | Microsoft's latest    |
| `gemma2`            | 9B      | Google's model        |

### Model Variants

```bash
# Quantization variants
ollama pull llama3.1:8b-instruct-q4_K_M   # 4-bit (smaller, faster)
ollama pull llama3.1:8b-instruct-q8_0     # 8-bit (better quality)
ollama pull llama3.1:8b-instruct-fp16     # Full precision

# Size variants
ollama pull llama3.1:8b    # 8 billion parameters
ollama pull llama3.1:70b   # 70 billion parameters

# New models (v0.6+ era)
ollama pull deepseek-r1:7b      # Reasoning, budget
ollama pull deepseek-r1:14b     # Reasoning, efficient
ollama pull deepseek-r1:32b     # Reasoning, balanced
ollama pull deepseek-r1:70b     # Reasoning, high quality
ollama pull qwen2.5:72b         # Largest Qwen, top quality
ollama pull phi4                # Microsoft Phi-4 14B
```

## New in v0.6+

Ollama v0.6 introduced several major features for production workloads:

### Structured Outputs (JSON Schema)

Force model responses to match a specific JSON schema. Useful for building applications that need reliable, parseable output:

```bash
curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "messages": [{"role": "user", "content": "Tell me about Canada."}],
  "format": {
    "type": "object",
    "properties": {
      "name": {"type": "string"},
      "capital": {"type": "string"},
      "population": {"type": "integer"},
      "languages": {
        "type": "array",
        "items": {"type": "string"}
      }
    },
    "required": ["name", "capital", "population", "languages"]
  },
  "stream": false
}'
```

Python example with structured outputs:

```python
from openai import OpenAI
import json

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

response = client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "List 3 programming languages with their main use cases"}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "languages",
            "schema": {
                "type": "object",
                "properties": {
                    "languages": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "properties": {
                                "name": {"type": "string"},
                                "use_case": {"type": "string"},
                                "popularity_rank": {"type": "integer"}
                            }
                        }
                    }
                }
            }
        }
    }
)

data = json.loads(response.choices[0].message.content)
print(data)
```

### OpenAI-Compatible Embeddings Endpoint (`/api/embed`)

New in v0.6+: the `/api/embed` endpoint is fully OpenAI-compatible and supports batched inputs:

```bash
# Single text embedding
curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": "Hello world"
}'

# Batch embeddings (new in v0.6)
curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": ["First document", "Second document", "Third document"]
}'
```

OpenAI client works directly with `/v1/embeddings`:

```python
from openai import OpenAI
import numpy as np

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

# Pull embedding model first: ollama pull nomic-embed-text
response = client.embeddings.create(
    model="nomic-embed-text",
    input=["Hello world", "Goodbye world"]
)

emb1 = np.array(response.data[0].embedding)
emb2 = np.array(response.data[1].embedding)

# Cosine similarity
similarity = np.dot(emb1, emb2) / (np.linalg.norm(emb1) * np.linalg.norm(emb2))
print(f"Similarity: {similarity:.4f}")
```

Popular embedding models:

```bash
ollama pull nomic-embed-text      # 137M, fast, good quality
ollama pull mxbai-embed-large     # 335M, higher quality
ollama pull all-minilm            # 23M, fastest
```

### Concurrent Model Loading

Before v0.6, Ollama would unload one model to load another. V0.6+ supports running multiple models simultaneously, limited only by available VRAM:

```bash
# Load two models at the same time
ollama run llama3.2 &
ollama run deepseek-r1:7b &

# Check what's running
curl http://localhost:11434/api/ps
```

Configure concurrency:

```bash
# Allow up to 4 models loaded simultaneously
OLLAMA_MAX_LOADED_MODELS=4 ollama serve

# Each runner in a separate process (better isolation)
OLLAMA_NUM_PARALLEL=2 ollama serve
```

This is especially useful for:

* A/B testing different models
* Specialized models for different tasks (coding + chat)
* Keeping frequently-used models warm in VRAM

## API Usage

### Chat Completion

```bash
# Via http_pub (external access):
curl https://your-http-pub.clorecloud.net/api/chat -d '{
  "model": "llama3.2",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": false
}'

# Via SSH tunnel (localhost):
curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": false
}'
```

{% hint style="info" %}
Add `"stream": false` to get the complete response at once instead of streaming.
{% endhint %}

### OpenAI-Compatible Endpoint

```python
from openai import OpenAI

# For external access, use your http_pub URL:
client = OpenAI(
    base_url="https://your-http-pub.clorecloud.net/v1",
    api_key="ollama"  # any string works
)

# Or via SSH tunnel:
# client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

response = client.chat.completions.create(
    model="llama3.2",
    messages=[
        {"role": "user", "content": "What is machine learning?"}
    ]
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "Write a poem"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
```

### Embeddings

```bash
# Legacy endpoint (still works)
curl http://localhost:11434/api/embeddings -d '{
  "model": "nomic-embed-text",
  "prompt": "Hello world"
}'

# New v0.6+ endpoint (batch support, OpenAI-compatible)
curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": ["Hello world", "Another text"]
}'
```

### Text Generation (Non-Chat)

```bash
curl https://your-http-pub.clorecloud.net/api/generate -d '{
  "model": "llama3.2",
  "prompt": "The meaning of life is",
  "stream": false
}'
```

## Complete API Reference

All endpoints work with both `http://localhost:11434` (via SSH) and `https://your-http-pub.clorecloud.net` (external).

### Model Management

| Endpoint       | Method | Description                   |
| -------------- | ------ | ----------------------------- |
| `/api/tags`    | GET    | List all downloaded models    |
| `/api/show`    | POST   | Get model details             |
| `/api/pull`    | POST   | Download a model              |
| `/api/delete`  | DELETE | Remove a model                |
| `/api/ps`      | GET    | List currently running models |
| `/api/version` | GET    | Get Ollama version            |

#### List Models

```bash
curl https://your-http-pub.clorecloud.net/api/tags
```

Response:

```json
{
  "models": [
    {"name": "llama3.2:latest", "size": 2019393189, "digest": "...", "modified_at": "..."}
  ]
}
```

#### Show Model Details

```bash
curl https://your-http-pub.clorecloud.net/api/show -d '{"name": "llama3.2"}'
```

#### Pull Model via API

```bash
curl https://your-http-pub.clorecloud.net/api/pull -d '{
  "name": "mistral:7b",
  "stream": false
}'
```

Response:

```json
{"status": "success"}
```

{% hint style="warning" %}
Large models may take several minutes to download. For very large models (30GB+), consider using SSH and the CLI: `ollama pull model-name`
{% endhint %}

#### Delete Model

```bash
curl -X DELETE https://your-http-pub.clorecloud.net/api/delete -d '{"name": "mistral:7b"}'
```

#### List Running Models

```bash
curl https://your-http-pub.clorecloud.net/api/ps
```

Response:

```json
{
  "models": [
    {"name": "llama3.2:latest", "size": 2019393189, "expires_at": "2025-01-25T12:00:00Z"}
  ]
}
```

#### Get Version

```bash
curl https://your-http-pub.clorecloud.net/api/version
```

Response:

```json
{"version": "0.6.8"}
```

### Inference Endpoints

| Endpoint               | Method | Description                                          |
| ---------------------- | ------ | ---------------------------------------------------- |
| `/api/generate`        | POST   | Text completion                                      |
| `/api/chat`            | POST   | Chat completion                                      |
| `/api/embeddings`      | POST   | Generate embeddings (legacy)                         |
| `/api/embed`           | POST   | Generate embeddings v0.6+ (batch, OpenAI-compatible) |
| `/v1/chat/completions` | POST   | OpenAI-compatible chat                               |
| `/v1/embeddings`       | POST   | OpenAI-compatible embeddings                         |

### Custom Model Creation

Create custom models with specific system prompts via API:

```bash
curl https://your-http-pub.clorecloud.net/api/create -d '{
  "name": "my-assistant",
  "modelfile": "FROM llama3.2\nSYSTEM You are a helpful coding assistant."
}'
```

## GPU Configuration

### Check GPU Usage

```bash
# In container or server
nvidia-smi

# Ollama shows GPU in logs
ollama run llama3.2 --verbose
```

### Multi-GPU

Ollama automatically uses available GPUs. For specific GPU:

```bash
CUDA_VISIBLE_DEVICES=0 ollama serve
```

### Memory Management

```bash
# Set GPU memory limit
OLLAMA_GPU_MEMORY=8GiB ollama serve

# Keep model loaded
OLLAMA_KEEP_ALIVE=24h ollama serve

# Allow concurrent models (v0.6+)
OLLAMA_MAX_LOADED_MODELS=3 ollama serve
```

## Custom Models (Modelfile)

Create custom models with system prompts:

```dockerfile
# Modelfile
FROM llama3.2

SYSTEM You are a helpful coding assistant. Always provide code examples.

PARAMETER temperature 0.7
PARAMETER top_p 0.9
```

```bash
ollama create coding-assistant -f Modelfile
ollama run coding-assistant
```

## Running as Service

### Systemd

```ini
# /etc/systemd/system/ollama.service
[Unit]
Description=Ollama Service
After=network.target

[Service]
Type=simple
ExecStart=/usr/local/bin/ollama serve
Restart=always
Environment="OLLAMA_HOST=0.0.0.0"

[Install]
WantedBy=multi-user.target
```

```bash
systemctl enable ollama
systemctl start ollama
```

## Performance Tips

1. **Use appropriate quantization**
   * Q4\_K\_M for speed
   * Q8\_0 for quality
   * fp16 for maximum quality
2. **Match model to VRAM**
   * 8GB: 7B models (Q4)
   * 16GB: 13B models or 7B (Q8)
   * 24GB: 34B models (Q4)
   * 48GB+: 70B models
3. **Keep model loaded**

   ```bash
   OLLAMA_KEEP_ALIVE=1h ollama serve
   ```
4. **Fast SSD improves performance**
   * Model loading and KV cache benefit from fast storage
   * Servers with NVMe SSD can achieve 2-3x better performance

## Benchmarks

### Generation Speed (tokens/sec)

| Model                | RTX 3060 | RTX 3090 | RTX 4090 | A100 40GB |
| -------------------- | -------- | -------- | -------- | --------- |
| Llama 3.2 3B (Q4)    | 120      | 160      | 200      | 220       |
| Llama 3.1 8B (Q4)    | 60       | 100      | 130      | 150       |
| Llama 3.1 8B (Q8)    | 45       | 80       | 110      | 130       |
| Mistral 7B (Q4)      | 70       | 110      | 140      | 160       |
| Mixtral 8x7B (Q4)    | -        | 35       | 55       | 75        |
| Llama 3.1 70B (Q4)   | -        | -        | 18       | 35        |
| DeepSeek-R1 7B (Q4)  | 65       | 105      | 135      | 155       |
| DeepSeek-R1 32B (Q4) | -        | -        | 22       | 42        |
| Qwen2.5 72B (Q4)     | -        | -        | 15       | 30        |
| Phi-4 14B (Q4)       | -        | 50       | 75       | 90        |

*Benchmarks updated January 2026. Actual speeds may vary based on server configuration.*

### Time to First Token (ms)

| Model | RTX 3090 | RTX 4090 | A100 |
| ----- | -------- | -------- | ---- |
| 3B    | 50       | 35       | 25   |
| 7-8B  | 120      | 80       | 60   |
| 13B   | 250      | 150      | 100  |
| 34B   | 600      | 350      | 200  |
| 70B   | -        | 1200     | 500  |

### Context Length vs VRAM (Q4)

| Model | 2K ctx | 4K ctx | 8K ctx | 16K ctx |
| ----- | ------ | ------ | ------ | ------- |
| 7B    | 5GB    | 6GB    | 8GB    | 12GB    |
| 13B   | 8GB    | 10GB   | 14GB   | 22GB    |
| 34B   | 20GB   | 24GB   | 32GB   | 48GB    |
| 70B   | 40GB   | 48GB   | 64GB   | 96GB    |

## GPU Requirements

| Model | Q4 VRAM | Q8 VRAM |
| ----- | ------- | ------- |
| 3B    | 3GB     | 5GB     |
| 7-8B  | 5GB     | 9GB     |
| 13B   | 8GB     | 15GB    |
| 34B   | 20GB    | 38GB    |
| 70B   | 40GB    | 75GB    |

## Cost Estimate

Typical CLORE.AI marketplace rates:

| GPU                                                                                                | VRAM | Price/day  | Good For         |
| -------------------------------------------------------------------------------------------------- | ---- | ---------- | ---------------- |
| [RTX 3090](https://clore.ai/rent-3090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=ollama) | 24GB | $0.30–1.00 | 13B-34B models   |
| [RTX 4090](https://clore.ai/rent-4090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=ollama) | 24GB | $0.50–2.00 | 34B models, fast |
| RTX 4070                                                                                           | 12GB | $0.20–0.50 | 7B models        |
| A100 80GB                                                                                          | 80GB | $1.50–3.00 | 70B models       |

*Prices in USD/day. Rates vary by provider — check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

## Troubleshooting

### Model won't load

```bash
# Check available memory
nvidia-smi

# Try smaller quantization
ollama pull llama3.1:8b-q4_0
```

### Slow generation

```bash
# Check if GPU is used
ollama run llama3.2 --verbose

# Ensure CUDA is available
nvidia-smi
```

### Connection refused

```bash
# Make sure server is running
ollama serve

# Check if binding to all interfaces
OLLAMA_HOST=0.0.0.0 ollama serve
```

### HTTP 502 on http\_pub URL

This means the service is still starting. Wait 30-60 seconds and retry:

```bash
# Check if service is ready
curl https://your-http-pub.clorecloud.net/

# Expected: "Ollama is running"
# If 502: wait and retry
```

## Next Steps

* [Open WebUI](/guides/language-models/open-webui) - Beautiful chat interface for Ollama
* [vLLM](/guides/language-models/vllm) - High-throughput production serving
* [DeepSeek-R1](/guides/language-models/deepseek-r1) - Reasoning model
* [DeepSeek-V3](/guides/language-models/deepseek-v3) - Best general model
* [Qwen2.5](/guides/language-models/qwen25) - Multilingual alternative
* [Text Generation WebUI](/guides/language-models/text-generation-webui) - Advanced features


# Open WebUI

ChatGPT-like interface for running LLMs on Clore.ai GPUs

Beautiful ChatGPT-like interface for running LLMs on CLORE.AI GPUs.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Why Open WebUI?

* **ChatGPT-like UI** - Familiar, polished interface
* **Multi-model** - Switch between models easily
* **RAG built-in** - Upload documents for context
* **User management** - Multi-user support
* **History** - Conversation persistence
* **Ollama integration** - Works out of the box

## Quick Deploy on CLORE.AI

**Docker Image:**

```
ghcr.io/open-webui/open-webui:cuda
```

**Ports:**

```
22/tcp
8080/http
```

**Command:**

```bash
# Start Ollama in background
ollama serve &
sleep 5
ollama pull llama3.2

# Start Open WebUI (connects to Ollama automatically)
# Note: The Docker image handles this
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

### Verify It's Working

```bash
# Check health
curl https://your-http-pub.clorecloud.net/health

# Get version
curl https://your-http-pub.clorecloud.net/api/version
```

Response:

```json
{"version": "0.7.2"}
```

{% hint style="warning" %}
If you get HTTP 502, wait 1-2 minutes - the service is still starting.
{% endhint %}

## Installation

### With Ollama (Recommended)

```bash
# Start Ollama first
docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

# Pull a model
docker exec -it ollama ollama pull llama3.2

# Start Open WebUI
docker run -d -p 8080:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main
```

### All-in-One (Bundled Ollama)

```bash
docker run -d -p 8080:8080 \
  --gpus all \
  -v ollama:/root/.ollama \
  -v open-webui:/app/backend/data \
  --name open-webui \
  ghcr.io/open-webui/open-webui:ollama
```

## First Setup

1. Open `http://your-server:8080`
2. Create admin account (first user becomes admin)
3. Go to Settings → Models → Pull a model
4. Start chatting!

## Features

### Chat Interface

* Markdown rendering
* Code highlighting
* Image generation (with compatible models)
* Voice input/output
* File attachments

### Model Management

* Pull models directly from UI
* Create custom models
* Set default model
* Model-specific settings

### RAG (Document Chat)

1. Click "+" in chat
2. Upload PDF, TXT, or other documents
3. Ask questions about the content

### User Management

* Multiple users
* Role-based access
* API key management
* Usage tracking

## Configuration

### Environment Variables

```bash
docker run -d \
  -e OLLAMA_BASE_URL=http://ollama:11434 \
  -e WEBUI_AUTH=True \
  -e WEBUI_NAME="My AI Chat" \
  -e DEFAULT_MODELS="llama3.2" \
  ghcr.io/open-webui/open-webui:main
```

### Key Settings

| Variable                | Description           | Default                  |
| ----------------------- | --------------------- | ------------------------ |
| `OLLAMA_BASE_URL`       | Ollama API URL        | `http://localhost:11434` |
| `WEBUI_AUTH`            | Enable authentication | `True`                   |
| `WEBUI_NAME`            | Instance name         | `Open WebUI`             |
| `DEFAULT_MODELS`        | Default model         | -                        |
| `ENABLE_RAG_WEB_SEARCH` | Web search in RAG     | `False`                  |

### Connect to Remote Ollama

```bash
docker run -d -p 8080:8080 \
  -e OLLAMA_BASE_URL=http://remote-server:11434 \
  ghcr.io/open-webui/open-webui:main
```

## Docker Compose

```yaml
version: '3.8'

services:
  ollama:
    image: ollama/ollama
    container_name: ollama
    volumes:
      - ollama:/root/.ollama
    ports:
      - "11434:11434"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    volumes:
      - open-webui:/app/backend/data
    ports:
      - "8080:8080"
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
    depends_on:
      - ollama

volumes:
  ollama:
  open-webui:
```

```bash
docker-compose up -d
```

## API Reference

Open WebUI provides several API endpoints:

| Endpoint           | Method | Description                  |
| ------------------ | ------ | ---------------------------- |
| `/health`          | GET    | Health check                 |
| `/api/version`     | GET    | Get Open WebUI version       |
| `/api/config`      | GET    | Get configuration            |
| `/ollama/api/tags` | GET    | List Ollama models (proxied) |
| `/ollama/api/chat` | POST   | Chat with Ollama (proxied)   |

### Check Health

```bash
curl https://your-http-pub.clorecloud.net/health
```

Response: `true`

### Get Version

```bash
curl https://your-http-pub.clorecloud.net/api/version
```

Response:

```json
{"version": "0.7.2"}
```

### List Models (via Ollama proxy)

```bash
curl https://your-http-pub.clorecloud.net/ollama/api/tags
```

{% hint style="info" %}
Most API operations require authentication. Use the web UI to create an account and manage API keys.
{% endhint %}

## Tips

### Faster Responses

1. Use quantized models (Q4\_K\_M)
2. Enable streaming in settings
3. Reduce context length if needed

### Better Quality

1. Use larger models (13B+)
2. Use Q8 quantization
3. Adjust temperature in model settings

### Save Resources

1. Set `OLLAMA_KEEP_ALIVE=5m`
2. Unload unused models
3. Use smaller models for testing

## GPU Requirements

Same as [Ollama](/guides/language-models/ollama#gpu-requirements).

Open WebUI itself uses minimal resources (\~500MB RAM).

## Troubleshooting

### Can't connect to Ollama

```bash
# Check Ollama is running
curl http://localhost:11434/api/tags

# If using Docker, use host networking or correct URL
docker run --network=host ghcr.io/open-webui/open-webui:main
```

### Models not showing

1. Check Ollama connection in Settings
2. Refresh model list
3. Pull models via CLI: `ollama pull modelname`

### Slow performance

1. Check GPU is being used: `nvidia-smi`
2. Try smaller/quantized models
3. Reduce concurrent users

## Cost Estimate

| Setup            | GPU      | Hourly  |
| ---------------- | -------- | ------- |
| Basic (7B)       | RTX 3060 | \~$0.03 |
| Standard (13B)   | RTX 3090 | \~$0.06 |
| Advanced (34B)   | RTX 4090 | \~$0.10 |
| Enterprise (70B) | A100     | \~$0.17 |

## Next Steps

* [Ollama](/guides/language-models/ollama) - CLI usage
* [LocalAI](/guides/language-models/localai-openai-compatible) - More backends
* [RAG + LangChain](/guides/training/finetune-llm) - Advanced RAG


# vLLM

High-throughput LLM inference with vLLM on Clore.ai GPUs

High-throughput LLM inference server for production workloads on CLORE.AI GPUs.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

{% hint style="info" %}
**Current Version: v0.7.x** — This guide covers vLLM v0.7.3+. New features include DeepSeek-R1 support, structured outputs with automatic tool choice, multi-LoRA serving, and improved memory efficiency.
{% endhint %}

## Server Requirements

| Parameter    | Minimum      | Recommended |
| ------------ | ------------ | ----------- |
| RAM          | **16GB**     | 32GB+       |
| VRAM         | 16GB (7B)    | 24GB+       |
| Network      | 500Mbps      | 1Gbps+      |
| Startup Time | 5-15 minutes | -           |

{% hint style="info" %}
Need a matching GPU? An [RTX 4090 (24GB)](https://clore.ai/rent-4090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=vllm) handles 7B–13B comfortably; for 70B+ models an [H100 (80GB)](https://clore.ai/rent-h100.html?utm_source=docs\&utm_medium=guide\&utm_campaign=vllm) is the safest choice.
{% endhint %}

{% hint style="danger" %}
**Important:** vLLM requires significant RAM and VRAM. Servers with less than 16GB RAM will fail to run even 7B models.
{% endhint %}

{% hint style="warning" %}
**Startup Time:** The first launch downloads the model from HuggingFace (5-15 minutes depending on model size and network speed). HTTP 502 during this time is normal.
{% endhint %}

## Why vLLM?

* **Fastest throughput** - PagedAttention for 24x higher throughput
* **Production ready** - OpenAI-compatible API out of the box
* **Continuous batching** - Efficient multi-user serving
* **Streaming** - Real-time token generation
* **Multi-GPU** - Tensor parallelism for large models
* **Multi-LoRA** - Serve multiple fine-tuned adapters simultaneously (v0.7+)
* **Structured outputs** - JSON schema enforcement and tool calling (v0.7+)

## Quick Deploy on CLORE.AI

**Docker Image:**

```
vllm/vllm-openai:v0.7.3
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
vllm serve mistralai/Mistral-7B-Instruct-v0.2 --host 0.0.0.0 --port 8000
```

### Verify It's Working

After deployment, find your `http_pub` URL in **My Orders**:

```bash
# Check health (may take 5-15 min on first run)
curl https://your-http-pub.clorecloud.net/health

# List models (only works after model loads)
curl https://your-http-pub.clorecloud.net/v1/models
```

{% hint style="warning" %}
If you get HTTP 502 for more than 15 minutes, check:

1. Server has 16GB+ RAM
2. Server has enough VRAM for the model
3. HuggingFace token is set for gated models
   {% endhint %}

## Accessing Your Service

When deployed on CLORE.AI, access vLLM via the `http_pub` URL:

```bash
# Chat completion
curl https://your-http-pub.clorecloud.net/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
```

{% hint style="info" %}
All `localhost:8000` examples below work when connected via SSH. For external access, replace with your `https://your-http-pub.clorecloud.net/` URL.
{% endhint %}

## Installation

### Using Docker (Recommended)

```bash
docker run -d --gpus all \
    -p 8000:8000 \
    --ipc=host \
    vllm/vllm-openai:v0.7.3 \
    --model mistralai/Mistral-7B-Instruct-v0.2 \
    --host 0.0.0.0
```

### Using pip

```bash
pip install vllm==0.7.3

# Run server
python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-Instruct-v0.2
```

## Supported Models

| Model                         | Parameters | VRAM Required     | RAM Required |
| ----------------------------- | ---------- | ----------------- | ------------ |
| Mistral 7B                    | 7B         | 14GB              | 16GB+        |
| Llama 3.1 8B                  | 8B         | 16GB              | 16GB+        |
| Llama 3.1 70B                 | 70B        | 140GB (or 2x80GB) | 64GB+        |
| Mixtral 8x7B                  | 47B        | 90GB              | 32GB+        |
| Qwen2.5 7B                    | 7B         | 14GB              | 16GB+        |
| Qwen2.5 72B                   | 72B        | 145GB             | 64GB+        |
| DeepSeek-V3                   | 236B MoE   | Multi-GPU         | 128GB+       |
| DeepSeek-R1-Distill-Qwen-7B   | 7B         | 14GB              | 16GB+        |
| DeepSeek-R1-Distill-Qwen-32B  | 32B        | 64GB              | 32GB+        |
| DeepSeek-R1-Distill-Llama-70B | 70B        | 140GB             | 64GB+        |
| Phi-4                         | 14B        | 28GB              | 32GB+        |
| Gemma 2 9B                    | 9B         | 18GB              | 16GB+        |
| CodeLlama 34B                 | 34B        | 68GB              | 32GB+        |

## Server Options

### Basic Server

```bash
vllm serve mistralai/Mistral-7B-Instruct-v0.2 \
    --host 0.0.0.0 \
    --port 8000
```

### Production Server

```bash
vllm serve mistralai/Mistral-7B-Instruct-v0.2 \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 1 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9 \
    --max-num-seqs 256 \
    --enable-prefix-caching
```

### With Quantization (Lower VRAM)

```bash
# AWQ quantized model (uses less VRAM)
vllm serve TheBloke/Mistral-7B-Instruct-v0.2-AWQ \
    --host 0.0.0.0 \
    --quantization awq
```

### Structured Outputs and Tool Calling (v0.7+)

Enable automatic tool choice and structured JSON outputs:

```bash
vllm serve mistralai/Mistral-7B-Instruct-v0.2 \
    --host 0.0.0.0 \
    --enable-auto-tool-choice \
    --tool-call-parser mistral
```

Use in Python:

```python
from openai import OpenAI
import json

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name"},
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
                },
                "required": ["city"]
            }
        }
    }
]

response = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[{"role": "user", "content": "What is the weather in Paris?"}],
    tools=tools,
    tool_choice="auto"
)

# Parse tool call
tool_call = response.choices[0].message.tool_calls[0]
args = json.loads(tool_call.function.arguments)
print(f"Tool: {tool_call.function.name}, Args: {args}")
```

Structured JSON output via response format:

```python
response = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[{"role": "user", "content": "Extract: John Smith, 30 years old, software engineer"}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "person",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "age": {"type": "integer"},
                    "occupation": {"type": "string"}
                },
                "required": ["name", "age", "occupation"]
            }
        }
    }
)
print(response.choices[0].message.content)
```

### Multi-LoRA Serving (v0.7+)

Serve a base model with multiple LoRA adapters simultaneously:

```bash
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
    --host 0.0.0.0 \
    --enable-lora \
    --lora-modules \
        sql-adapter=path/to/sql-lora \
        code-adapter=path/to/code-lora \
        chat-adapter=path/to/chat-lora \
    --max-lora-rank 64
```

Query a specific LoRA adapter by model name:

```python
# Use the SQL adapter
response = client.chat.completions.create(
    model="sql-adapter",
    messages=[{"role": "user", "content": "Write a SQL query to find top 10 customers"}]
)

# Use the code adapter
response = client.chat.completions.create(
    model="code-adapter",
    messages=[{"role": "user", "content": "Write a Python function to sort a list"}]
)
```

## DeepSeek-R1 Support (v0.7+)

vLLM v0.7+ has native support for DeepSeek-R1 distill models. These reasoning models produce `<think>` tags showing their reasoning process.

### DeepSeek-R1-Distill-Qwen-7B (Single GPU)

```bash
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-7B \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 16384
```

### DeepSeek-R1-Distill-Qwen-32B (Dual GPU)

```bash
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 2 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.90
```

### DeepSeek-R1-Distill-Llama-70B (Quad GPU)

```bash
vllm serve deepseek-ai/DeepSeek-R1-Distill-Llama-70B \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 4 \
    --max-model-len 32768
```

### Querying DeepSeek-R1

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
    messages=[
        {
            "role": "user",
            "content": "Solve: If a train travels 120km in 1.5 hours, what is its speed in m/s?"
        }
    ],
    max_tokens=2048,
    temperature=0.6
)

content = response.choices[0].message.content
# Response includes <think>...</think> reasoning block followed by the answer
print(content)
```

Parsing think tags:

```python
import re

def parse_deepseek_r1_response(content: str) -> dict:
    """Extract thinking and answer from DeepSeek-R1 response."""
    think_match = re.search(r'<think>(.*?)</think>', content, re.DOTALL)
    thinking = think_match.group(1).strip() if think_match else ""
    answer = re.sub(r'<think>.*?</think>', '', content, flags=re.DOTALL).strip()
    return {"thinking": thinking, "answer": answer}

result = parse_deepseek_r1_response(content)
print("Thinking:", result["thinking"][:200], "...")
print("Answer:", result["answer"])
```

## API Usage

### Chat Completions (OpenAI Compatible)

```python
from openai import OpenAI

# For external access, use your http_pub URL:
client = OpenAI(
    base_url="https://your-http-pub.clorecloud.net/v1",
    api_key="not-needed"
)

# Or via SSH tunnel:
# client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[
        {"role": "user", "content": "Explain quantum computing"}
    ],
    max_tokens=500,
    temperature=0.7
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[{"role": "user", "content": "Write a poem"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
```

### cURL

```bash
curl https://your-http-pub.clorecloud.net/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 100
  }'
```

### Text Completions

```bash
curl https://your-http-pub.clorecloud.net/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "prompt": "The capital of France is",
    "max_tokens": 50
  }'
```

## Complete API Reference

vLLM provides OpenAI-compatible endpoints plus additional utility endpoints.

### Standard Endpoints

| Endpoint               | Method | Description                     |
| ---------------------- | ------ | ------------------------------- |
| `/v1/models`           | GET    | List available models           |
| `/v1/chat/completions` | POST   | Chat completion                 |
| `/v1/completions`      | POST   | Text completion                 |
| `/health`              | GET    | Health check (may return empty) |

### Additional Endpoints

| Endpoint      | Method | Description              |
| ------------- | ------ | ------------------------ |
| `/tokenize`   | POST   | Tokenize text            |
| `/detokenize` | POST   | Convert tokens to text   |
| `/version`    | GET    | Get vLLM version         |
| `/docs`       | GET    | Swagger UI documentation |
| `/metrics`    | GET    | Prometheus metrics       |

#### Tokenize Text

Useful for counting tokens before sending requests:

```bash
curl https://your-http-pub.clorecloud.net/tokenize \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "prompt": "Hello world"
  }'
```

Response:

```json
{"count": 2, "max_model_len": 32768, "tokens": [9707, 1879]}
```

#### Detokenize

Convert token IDs back to text:

```bash
curl https://your-http-pub.clorecloud.net/detokenize \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "tokens": [9707, 1879]
  }'
```

Response:

```json
{"prompt": "Hello world"}
```

#### Get Version

```bash
curl https://your-http-pub.clorecloud.net/version
```

Response:

```json
{"version": "0.7.3"}
```

#### Swagger Documentation

Open in browser for interactive API documentation:

```
https://your-http-pub.clorecloud.net/docs
```

#### Prometheus Metrics

For monitoring:

```bash
curl https://your-http-pub.clorecloud.net/metrics
```

{% hint style="info" %}
**Reasoning Models:** DeepSeek-R1 and similar models include `<think>` tags in responses showing the model's reasoning process before the final answer.
{% endhint %}

## Benchmarks

### Throughput (tokens/sec per user)

| Model              | RTX 3090 | RTX 4090 | A100 40GB | A100 80GB |
| ------------------ | -------- | -------- | --------- | --------- |
| Mistral 7B         | 100      | 170      | 210       | 230       |
| Llama 3.1 8B       | 95       | 150      | 200       | 220       |
| Llama 3.1 8B (AWQ) | 130      | 190      | 260       | 280       |
| Mixtral 8x7B       | -        | 45       | 70        | 85        |
| Llama 3.1 70B      | -        | -        | 25 (2x)   | 45 (2x)   |
| DeepSeek-R1 7B     | 90       | 145      | 190       | 210       |
| DeepSeek-R1 32B    | -        | -        | 40        | 70 (2x)   |

*Benchmarks updated January 2026.*

### Context Length vs VRAM

| Model    | 4K ctx | 8K ctx | 16K ctx | 32K ctx |
| -------- | ------ | ------ | ------- | ------- |
| 8B FP16  | 18GB   | 22GB   | 30GB    | 46GB    |
| 8B AWQ   | 8GB    | 10GB   | 14GB    | 22GB    |
| 70B FP16 | 145GB  | 160GB  | 190GB   | 250GB   |
| 70B AWQ  | 42GB   | 50GB   | 66GB    | 98GB    |

## Hugging Face Authentication

For gated models (Llama, etc.):

```bash
# Set token in command
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
    --host 0.0.0.0 \
    --env HUGGING_FACE_HUB_TOKEN=hf_xxxxx
```

Or set as environment variable:

```bash
export HUGGING_FACE_HUB_TOKEN=hf_xxxxx
```

## GPU Requirements

| Model | Min VRAM | Min RAM  | Recommended         |
| ----- | -------- | -------- | ------------------- |
| 7-8B  | 16GB     | **16GB** | 24GB VRAM, 32GB RAM |
| 13B   | 26GB     | 32GB     | 40GB VRAM           |
| 34B   | 70GB     | 32GB     | 80GB VRAM           |
| 70B   | 140GB    | 64GB     | 2x80GB              |

## Cost Estimate

Typical CLORE.AI marketplace rates:

| GPU      | VRAM | Price/day  | Best For      |
| -------- | ---- | ---------- | ------------- |
| RTX 3090 | 24GB | $0.30–1.00 | 7-8B models   |
| RTX 4090 | 24GB | $0.50–2.00 | 7-13B, fast   |
| A100     | 40GB | $1.50–3.00 | 13-34B models |
| A100     | 80GB | $2.00–4.00 | 34-70B models |

*Prices in USD/day. Rates vary by provider — check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

## Troubleshooting

### HTTP 502 for a long time

1. **Check RAM:** Server must have 16GB+ RAM
2. **Check VRAM:** Must fit the model
3. **Model downloading:** First run downloads from HuggingFace (5-15 min)
4. **HF Token:** Gated models require authentication

### Out of Memory

```bash
# Reduce memory usage
--gpu-memory-utilization 0.8
--max-model-len 4096
--max-num-seqs 64

# Or use quantization
--quantization awq
```

### Model Download Fails

```bash
# Check HF token
echo $HUGGING_FACE_HUB_TOKEN

# Pre-download model
huggingface-cli download mistralai/Mistral-7B-Instruct-v0.2
```

## vLLM vs Others

| Feature      | vLLM        | llama.cpp | Ollama  |
| ------------ | ----------- | --------- | ------- |
| Throughput   | Best        | Good      | Good    |
| VRAM Usage   | High        | Low       | Medium  |
| Ease of Use  | Medium      | Medium    | Easy    |
| Startup Time | 5-15 min    | 1-2 min   | 30 sec  |
| Multi-GPU    | Native      | Limited   | Limited |
| Tool Calling | Yes (v0.7+) | Limited   | Limited |
| Multi-LoRA   | Yes (v0.7+) | No        | No      |

**Use vLLM when:**

* High throughput is priority
* Serving multiple users
* Have enough VRAM and RAM
* Production deployment
* Need tool calling / structured outputs

**Use Ollama when:**

* Quick setup needed
* Single user
* Less resources available

## Next Steps

* [Ollama](/guides/language-models/ollama) - Simpler alternative with faster startup
* [DeepSeek-R1](/guides/language-models/deepseek-r1) - Reasoning model guide
* [DeepSeek-V3](/guides/language-models/deepseek-v3) - Best general model
* [Qwen2.5](/guides/language-models/qwen25) - Multilingual models
* [Llama.cpp](/guides/language-models/llamacpp-server) - Lower VRAM option


# Llama.cpp Server

Efficient LLM inference with llama.cpp server on Clore.ai GPUs

Run LLMs efficiently with llama.cpp server on GPU.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Server Requirements

| Parameter    | Minimum       | Recommended |
| ------------ | ------------- | ----------- |
| RAM          | 8GB           | 16GB+       |
| VRAM         | 6GB           | 8GB+        |
| Network      | 200Mbps       | 500Mbps+    |
| Startup Time | \~2-5 minutes | -           |

{% hint style="info" %}
Llama.cpp is memory-efficient due to GGUF quantization. 7B models can run on 6-8GB VRAM.
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## What is Llama.cpp?

Llama.cpp is the fastest CPU/GPU inference engine for LLMs:

* Supports GGUF quantized models
* Low memory usage
* OpenAI-compatible API
* Multi-user support

## Quantization Levels

| Format   | Size (7B) | Speed   | Quality   |
| -------- | --------- | ------- | --------- |
| Q2\_K    | 2.8GB     | Fastest | Low       |
| Q4\_K\_M | 4.1GB     | Fast    | Good      |
| Q5\_K\_M | 4.8GB     | Medium  | Great     |
| Q6\_K    | 5.5GB     | Slower  | Excellent |
| Q8\_0    | 7.2GB     | Slowest | Best      |

## Quick Deploy

**Docker Image:**

```
ghcr.io/ggerganov/llama.cpp:server-cuda
```

**Ports:**

```
22/tcp
8080/http
```

**Command:**

```bash

# Download model
wget https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf

# Run server
./llama-server \
    -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
    --host 0.0.0.0 \
    --port 8080 \
    -ngl 35 \
    -c 4096
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

### Verify It's Working

```bash
# Check health
curl https://your-http-pub.clorecloud.net/health

# Get server info
curl https://your-http-pub.clorecloud.net/props
```

{% hint style="warning" %}
If you get HTTP 502, the service may still be starting or downloading the model. Wait 2-5 minutes and retry.
{% endhint %}

## Complete API Reference

### Standard Endpoints

| Endpoint               | Method | Description                         |
| ---------------------- | ------ | ----------------------------------- |
| `/health`              | GET    | Health check                        |
| `/v1/models`           | GET    | List models                         |
| `/v1/chat/completions` | POST   | Chat (OpenAI compatible)            |
| `/v1/completions`      | POST   | Text completion (OpenAI compatible) |
| `/v1/embeddings`       | POST   | Generate embeddings                 |
| `/completion`          | POST   | Native completion endpoint          |
| `/tokenize`            | POST   | Tokenize text                       |
| `/detokenize`          | POST   | Detokenize tokens                   |
| `/props`               | GET    | Server properties                   |
| `/metrics`             | GET    | Prometheus metrics                  |

#### Tokenize Text

```bash
curl https://your-http-pub.clorecloud.net/tokenize \
    -H "Content-Type: application/json" \
    -d '{"content": "Hello world"}'
```

Response:

```json
{"tokens": [15496, 1917]}
```

#### Server Properties

```bash
curl https://your-http-pub.clorecloud.net/props
```

Response:

```json
{
  "total_slots": 1,
  "chat_template": "...",
  "default_generation_settings": {...}
}
```

## Build from Source

```bash

# Clone repo
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp

# Build with CUDA
make LLAMA_CUDA=1

# Or with CMake
mkdir build && cd build
cmake .. -DLLAMA_CUDA=ON
cmake --build . --config Release
```

## Download Models

```bash

# Llama 3.1 8B
wget https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf

# Mistral 7B
wget https://huggingface.co/bartowski/Mistral-7B-Instruct-v0.3-GGUF/resolve/main/Mistral-7B-Instruct-v0.3-Q4_K_M.gguf

# Mixtral 8x7B
wget https://huggingface.co/bartowski/Mixtral-8x7B-Instruct-v0.1-GGUF/resolve/main/Mixtral-8x7B-Instruct-v0.1-Q4_K_M.gguf

# Phi-2
wget https://huggingface.co/bartowski/Phi-4-GGUF/resolve/main/Phi-4-Q4_K_M.gguf

# CodeLlama 7B
wget https://huggingface.co/bartowski/CodeLlama-7B-Instruct-GGUF/resolve/main/CodeLlama-7B-Instruct-Q4_K_M.gguf
```

## Server Options

### Basic Server

```bash
./llama-server \
    -m model.gguf \
    --host 0.0.0.0 \
    --port 8080
```

### Full GPU Offload

```bash
./llama-server \
    -m model.gguf \
    --host 0.0.0.0 \
    --port 8080 \
    -ngl 99 \           # GPU layers (99 = all)
    -c 4096 \           # Context size
    -t 8 \              # CPU threads
    --parallel 4        # Concurrent requests
```

### All Options

```bash
./llama-server \
    -m model.gguf \           # Model file
    --host 0.0.0.0 \          # Bind address
    --port 8080 \             # Port
    -ngl 35 \                 # GPU layers
    -c 4096 \                 # Context size
    -t 8 \                    # Threads
    -b 512 \                  # Batch size
    --parallel 4 \            # Parallel requests
    --mlock \                 # Lock memory
    --no-mmap \               # Disable mmap
    --cont-batching \         # Continuous batching
    --flash-attn \            # Flash attention
    --metrics                 # Enable metrics endpoint
```

## API Usage

### Chat Completions (OpenAI Compatible)

```python
import openai

client = openai.OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="llama-3.1-8b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is machine learning?"}
    ],
    temperature=0.7,
    max_tokens=500
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="llama-3.1-8b",
    messages=[{"role": "user", "content": "Write a story"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

### Text Completion

```python
response = client.completions.create(
    model="llama-3.1-8b",
    prompt="The future of AI is",
    max_tokens=100,
    temperature=0.8
)

print(response.choices[0].text)
```

### Embeddings

```python
response = client.embeddings.create(
    model="llama-3.1-8b",
    input="Hello, world!"
)

print(f"Embedding: {response.data[0].embedding[:5]}...")
```

## cURL Examples

### Chat

```bash
curl http://localhost:8080/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "llama-3.1-8b",
        "messages": [
            {"role": "user", "content": "Hello!"}
        ]
    }'
```

### Completion

```bash
curl http://localhost:8080/completion \
    -H "Content-Type: application/json" \
    -d '{
        "prompt": "Building a website requires",
        "n_predict": 128,
        "temperature": 0.7
    }'
```

### Health Check

```bash
curl http://localhost:8080/health
```

### Metrics

```bash
curl http://localhost:8080/metrics
```

## Multi-GPU

```bash

# Split across GPUs
./llama-server \
    -m model.gguf \
    -ngl 99 \
    --tensor-split 0.5,0.5 \  # Split between 2 GPUs
    --main-gpu 0              # Primary GPU
```

## Memory Optimization

### For Limited VRAM

```bash

# Partial offload
./llama-server -m model.gguf -ngl 20 -c 2048

# Use smaller quantization

# Download Q2_K or Q3_K instead of Q4_K
```

### For Maximum Speed

```bash
./llama-server \
    -m model.gguf \
    -ngl 99 \
    --flash-attn \
    --cont-batching \
    --parallel 8 \
    -b 1024
```

## Model-Specific Templates

### Llama 2 Chat

```bash
./llama-server -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
    --chat-template llama2
```

### Mistral Instruct

```bash
./llama-server -m mistral-7b-instruct.gguf \
    --chat-template mistral
```

### ChatML (Many Models)

```bash
./llama-server -m model.gguf \
    --chat-template chatml
```

## Python Server Wrapper

```python
import subprocess
import requests
import time

class LlamaCppServer:
    def __init__(self, model_path, port=8080, gpu_layers=35):
        self.port = port
        self.process = subprocess.Popen([
            "./llama-server",
            "-m", model_path,
            "--host", "0.0.0.0",
            "--port", str(port),
            "-ngl", str(gpu_layers),
            "-c", "4096"
        ])
        self._wait_for_ready()

    def _wait_for_ready(self, timeout=60):
        start = time.time()
        while time.time() - start < timeout:
            try:
                r = requests.get(f"http://localhost:{self.port}/health")
                if r.status_code == 200:
                    return
            except:
                pass
            time.sleep(1)
        raise TimeoutError("Server didn't start")

    def chat(self, messages, **kwargs):
        response = requests.post(
            f"http://localhost:{self.port}/v1/chat/completions",
            json={"messages": messages, **kwargs}
        )
        return response.json()

    def stop(self):
        self.process.terminate()

# Usage
server = LlamaCppServer("llama-3.1-8b.gguf")
result = server.chat([{"role": "user", "content": "Hello!"}])
print(result["choices"][0]["message"]["content"])
server.stop()
```

## Benchmarking

```bash

# Built-in benchmark
./llama-bench -m model.gguf -ngl 99

# Output includes:

# - Tokens per second

# - Memory usage

# - Load time
```

## Performance Comparison

| Model        | GPU      | Quantization | Tokens/sec |
| ------------ | -------- | ------------ | ---------- |
| Llama 3.1 8B | RTX 3090 | Q4\_K\_M     | \~100      |
| Llama 3.1 8B | RTX 4090 | Q4\_K\_M     | \~150      |
| Llama 3.1 8B | RTX 3090 | Q4\_K\_M     | \~60       |
| Mistral 7B   | RTX 3090 | Q4\_K\_M     | \~110      |
| Mixtral 8x7B | A100     | Q4\_K\_M     | \~50       |

## Troubleshooting

### CUDA Not Detected

```bash

# Rebuild with CUDA
make clean
make LLAMA_CUDA=1

# Check CUDA
nvidia-smi
```

### Out of Memory

```bash

# Reduce GPU layers
-ngl 20  # Instead of 99

# Reduce context
-c 2048  # Instead of 4096

# Use smaller quant

# Q4_K_S instead of Q4_K_M
```

### Slow Generation

```bash

# Increase batch size
-b 1024

# Enable flash attention
--flash-attn

# Enable continuous batching
--cont-batching
```

## Production Setup

### Systemd Service

```ini

# /etc/systemd/system/llama.service
[Unit]
Description=Llama.cpp Server
After=network.target

[Service]
Type=simple
ExecStart=/opt/llama.cpp/llama-server -m /models/model.gguf -ngl 99 --host 0.0.0.0 --port 8080
Restart=always

[Install]
WantedBy=multi-user.target
```

### With nginx

```nginx
upstream llama {
    server localhost:8080;
}

server {
    listen 80;

    location / {
        proxy_pass http://llama;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
    }
}
```

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* vLLM Inference - Higher throughput
* [ExLlamaV2](/guides/language-models/exllamav2-fast) - Faster inference
* [Text Generation WebUI](/guides/language-models/text-generation-webui) - Web interface


# Text Generation WebUI

Run text-generation-webui for LLM inference on Clore.ai GPUs

Run the most popular LLM interface with support for all model formats.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## Why Text Generation WebUI?

* Supports GGUF, GPTQ, AWQ, EXL2, HF formats
* Built-in chat, notebook, and API modes
* Extensions: voice, characters, multimodal
* Fine-tuning support
* Model switching on the fly

## Requirements

| Model Size | Min VRAM | Recommended |
| ---------- | -------- | ----------- |
| 7B (Q4)    | 6GB      | RTX 3060    |
| 13B (Q4)   | 10GB     | RTX 3080    |
| 30B (Q4)   | 20GB     | RTX 4090    |
| 70B (Q4)   | 40GB     | A100        |

## Quick Deploy

**Docker Image:**

```
atinoda/text-generation-webui:default-nvidia
```

**Ports:**

```
22/tcp
7860/http
5000/http
5005/http
```

**Environment:**

```
EXTRA_LAUNCH_ARGS=--listen --api
```

## Manual Installation

**Image:**

```
nvidia/cuda:12.1.0-devel-ubuntu22.04
```

**Ports:**

```
22/tcp
7860/http
5000/http
```

**Command:**

```bash
apt-get update && apt-get install -y git python3 python3-pip && \
cd /workspace && \
git clone https://github.com/oobabooga/text-generation-webui.git && \
cd text-generation-webui && \
pip install -r requirements.txt && \
python server.py --listen --api
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Access WebUI

1. Wait for deployment
2. Find port 7860 mapping in **My Orders**
3. Open: `http://<proxy>:<port>`

## Download Models

### From HuggingFace (in WebUI)

1. Go to **Model** tab
2. Enter model name: `bartowski/Meta-Llama-3.1-8B-Instruct-GGUF`
3. Click **Download**

### Via Command Line

```bash
cd /workspace/text-generation-webui

# Download GGUF model
python download-model.py bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

# Download specific file
python download-model.py bartowski/Meta-Llama-3.1-8B-Instruct-GGUF --specific-file Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf
```

### Recommended Models

**For Chat:**

```bash

# Llama 2 Chat (7B, fast)
python download-model.py bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

# Mistral Instruct (excellent)
python download-model.py bartowski/Mistral-7B-Instruct-v0.3-GGUF

# OpenHermes (great all-rounder)
python download-model.py bartowski/OpenHermes-2.5-Mistral-7B-GGUF
```

**For Coding:**

```bash

# CodeLlama
python download-model.py bartowski/CodeLlama-13B-Instruct-GGUF

# DeepSeek Coder
python download-model.py bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF
```

**For Roleplay:**

```bash

# MythoMax
python download-model.py bartowski/MythoMax-L2-13B-GGUF
```

## Loading Models

### GGUF (Recommended for most users)

1. **Model** tab → Select model folder
2. **Model loader:** llama.cpp
3. Set **n-gpu-layers:**
   * RTX 3090: 35-40
   * RTX 4090: 45-50
   * A100: 80+
4. Click **Load**

### GPTQ (Fast, quantized)

1. Download GPTQ model
2. **Model loader:** ExLlama\_HF or AutoGPTQ
3. Load model

### EXL2 (Best speed)

```bash

# Install exllamav2
pip install exllamav2
```

1. Download EXL2 model
2. **Model loader:** ExLlamav2\_HF
3. Load

## Chat Configuration

### Character Setup

1. Go to **Parameters** → **Character**
2. Create or load character card
3. Set:
   * Name
   * Context/persona
   * Example dialogue

### Instruct Mode

For instruction-tuned models:

1. **Parameters** → **Instruction template**
2. Select template matching your model:
   * Llama-2-chat
   * Mistral
   * ChatML
   * Alpaca

## API Usage

### Enable API

Start with `--api` flag (default port 5000)

### OpenAI-compatible API

```python
import openai

openai.api_base = "http://localhost:5000/v1"
openai.api_key = "not-needed"

response = openai.ChatCompletion.create(
    model="any",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
```

### Native API

```python
import requests

response = requests.post(
    "http://localhost:5000/api/v1/generate",
    json={
        "prompt": "Write a story about",
        "max_new_tokens": 200,
        "temperature": 0.7
    }
)
print(response.json()["results"][0]["text"])
```

## Extensions

### Installing Extensions

```bash
cd /workspace/text-generation-webui/extensions

# Silero TTS (voice)
git clone https://github.com/oobabooga/text-generation-webui-extensions

# SuperBoogav2 (RAG/long-term memory)

# Already included, enable in UI
```

### Enable Extensions

1. **Session** tab → **Extensions**
2. Check boxes for desired extensions
3. Click **Apply and restart**

### Popular Extensions

| Extension         | Purpose             |
| ----------------- | ------------------- |
| silero\_tts       | Voice output        |
| whisper\_stt      | Voice input         |
| superbooga        | Document Q\&A       |
| sd\_api\_pictures | Image generation    |
| multimodal        | Image understanding |

## Performance Tuning

### GGUF Settings

```
n_gpu_layers: 35    # GPU layers (more = faster)
n_ctx: 4096         # Context length
n_batch: 512        # Batch size
threads: 8          # CPU threads
```

### Memory Optimization

For limited VRAM:

```bash
python server.py --listen --n-gpu-layers 20 --no-mmap
```

### Speed Optimization

```bash

# Use llama.cpp with cuBLAS
python server.py --listen --loader llama.cpp --n-gpu-layers 50 --threads 8
```

## Fine-tuning (LoRA)

### Training Tab

1. Go to **Training** tab
2. Load base model
3. Upload dataset (JSON format)
4. Configure:
   * LoRA rank: 8-32
   * Learning rate: 1e-4
   * Epochs: 3-5
5. Start training

### Dataset Format

```json
[
  {"instruction": "Summarize this:", "input": "Long text...", "output": "Summary..."},
  {"instruction": "Translate to French:", "input": "Hello", "output": "Bonjour"}
]
```

## Saving Your Work

```bash

# Save models
rsync -avz /workspace/text-generation-webui/models/ backup-server:/models/

# Save characters
rsync -avz /workspace/text-generation-webui/characters/ backup-server:/characters/

# Save LoRAs
rsync -avz /workspace/text-generation-webui/loras/ backup-server:/loras/
```

## Troubleshooting

### Model won't load

* Check VRAM usage: `nvidia-smi`
* Reduce `n_gpu_layers`
* Use smaller quantization (Q4\_K\_M → Q4\_K\_S)

### Slow generation

* Increase `n_gpu_layers`
* Use EXL2 instead of GGUF
* Enable `--no-mmap`

{% hint style="danger" %}
**Out of memory**
{% endhint %}

during generation - Reduce \`n\_ctx\` (context length) - Use \`--n-gpu-layers 0\` for CPU-only - Try smaller model

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers


# ExLlamaV2

Maximum speed LLM inference with ExLlamaV2 on Clore.ai GPUs

Run LLMs at maximum speed with ExLlamaV2.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## What is ExLlamaV2?

ExLlamaV2 is the fastest inference engine for large language models:

* 2-3x faster than other engines
* Excellent quantization (EXL2)
* Low VRAM usage
* Supports speculative decoding

## Requirements

| Model Size | Min VRAM | Recommended |
| ---------- | -------- | ----------- |
| 7B         | 6GB      | RTX 3060    |
| 13B        | 10GB     | RTX 3090    |
| 34B        | 20GB     | RTX 4090    |
| 70B        | 40GB     | A100        |

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-devel
```

**Ports:**

```
22/tcp
8080/http
```

**Command:**

```bash
pip install exllamav2 && \
huggingface-cli download turboderp/Llama2-7B-exl2 --local-dir ./model && \
python -m exllamav2.server --model_dir ./model --host 0.0.0.0 --port 8080
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Installation

```bash

# Install from PyPI
pip install exllamav2

# Or from source (latest features)
git clone https://github.com/turboderp/exllamav2
cd exllamav2
pip install .
```

## Download Models

### EXL2 Quantized Models

```bash

# Llama 3.1 8B (4.0 bpw)
huggingface-cli download turboderp/Llama2-7B-exl2 \
    --revision 4.0bpw \
    --local-dir ./llama2-7b-exl2

# Llama 3.1 8B (4.0 bpw)
huggingface-cli download turboderp/Llama2-13B-exl2 \
    --revision 4.0bpw \
    --local-dir ./llama2-13b-exl2

# Mistral 7B (4.0 bpw)
huggingface-cli download turboderp/Mistral-7B-instruct-exl2 \
    --revision 4.0bpw \
    --local-dir ./mistral-7b-exl2

# Mixtral 8x7B
huggingface-cli download turboderp/Mixtral-8x7B-instruct-exl2 \
    --revision 4.0bpw \
    --local-dir ./mixtral-exl2
```

### Bits Per Weight (bpw)

| BPW | Quality   | VRAM (7B) |
| --- | --------- | --------- |
| 2.0 | Low       | \~3GB     |
| 3.0 | Good      | \~4GB     |
| 4.0 | Great     | \~5GB     |
| 5.0 | Excellent | \~6GB     |
| 6.0 | Near-FP16 | \~7GB     |

## Python API

### Basic Generation

```python
from exllamav2 import ExLlamaV2, ExLlamaV2Config, ExLlamaV2Cache, ExLlamaV2Tokenizer
from exllamav2.generator import ExLlamaV2StreamingGenerator, ExLlamaV2Sampler

# Load model
config = ExLlamaV2Config()
config.model_dir = "./llama2-7b-exl2"
config.prepare()

model = ExLlamaV2(config)
model.load()

tokenizer = ExLlamaV2Tokenizer(config)
cache = ExLlamaV2Cache(model, lazy=True)

# Create generator
generator = ExLlamaV2StreamingGenerator(model, cache, tokenizer)

# Set sampling settings
settings = ExLlamaV2Sampler.Settings()
settings.temperature = 0.7
settings.top_k = 50
settings.top_p = 0.9

# Generate
prompt = "The future of artificial intelligence is"
output = generator.generate_simple(prompt, settings, num_tokens=200)
print(output)
```

### Streaming Generation

```python
from exllamav2.generator import ExLlamaV2StreamingGenerator

generator = ExLlamaV2StreamingGenerator(model, cache, tokenizer)

prompt = "Write a short story about a robot:"
input_ids = tokenizer.encode(prompt)

generator.set_stop_conditions([tokenizer.eos_token_id])
generator.begin_stream(input_ids, settings)

while True:
    chunk, eos, _ = generator.stream()
    if eos:
        break
    print(chunk, end="", flush=True)
```

### Chat Format

```python
def format_chat(messages):
    text = ""
    for msg in messages:
        role = msg["role"]
        content = msg["content"]
        if role == "system":
            text += f"[INST] <<SYS>>\n{content}\n<</SYS>>\n\n"
        elif role == "user":
            text += f"{content} [/INST]"
        elif role == "assistant":
            text += f" {content}</s><s>[INST] "
    return text

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is Python?"}
]

prompt = format_chat(messages)
output = generator.generate_simple(prompt, settings, num_tokens=300)
```

## Server Mode

### Start Server

```bash
python -m exllamav2.server \
    --model_dir ./llama2-7b-exl2 \
    --host 0.0.0.0 \
    --port 8080 \
    --max_seq_len 4096 \
    --cache_size 4096
```

### API Usage

```python
import requests

response = requests.post(
    "http://localhost:8080/v1/completions",
    json={
        "prompt": "Hello, how are you?",
        "max_tokens": 100,
        "temperature": 0.7
    }
)

print(response.json()["choices"][0]["text"])
```

### Chat Completions

```python
import openai

client = openai.OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="llama2-7b",
    messages=[{"role": "user", "content": "Hello!"}],
    temperature=0.7
)

print(response.choices[0].message.content)
```

## TabbyAPI (Recommended Server)

TabbyAPI provides a feature-rich ExLlamaV2 server:

```bash

# Clone TabbyAPI
git clone https://github.com/theroyallab/tabbyAPI
cd tabbyAPI

# Install
pip install -r requirements.txt

# Configure

# Edit config.yml with your model path

# Run
python main.py
```

### TabbyAPI Features

* OpenAI-compatible API
* Multiple model support
* LoRA hot-swapping
* Streaming
* Function calling
* Admin API

## Speculative Decoding

Use a smaller model to accelerate generation:

```python
from exllamav2 import ExLlamaV2, ExLlamaV2Config, ExLlamaV2Cache

# Load main model (13B)
main_config = ExLlamaV2Config()
main_config.model_dir = "./llama2-13b-exl2"
main_config.prepare()
main_model = ExLlamaV2(main_config)
main_model.load()

# Load draft model (7B)
draft_config = ExLlamaV2Config()
draft_config.model_dir = "./llama2-7b-exl2"
draft_config.prepare()
draft_model = ExLlamaV2(draft_config)
draft_model.load()

# Create speculative generator
from exllamav2.generator import ExLlamaV2DraftGenerator

generator = ExLlamaV2DraftGenerator(
    main_model, draft_model,
    cache_main, cache_draft,
    tokenizer
)

# Generate (faster with speculation)
output = generator.generate_simple(prompt, settings, num_tokens=500)
```

## Quantize Your Own Models

### Convert to EXL2

```python
from exllamav2 import ExLlamaV2, ExLlamaV2Config
from exllamav2.conversion import convert_model

# Source: HuggingFace model

# Target: EXL2 quantized

convert_model(
    input_dir="./llama-3.1-8b-hf",
    output_dir="./llama-3.1-8b-exl2-4bpw",
    cal_dataset="wikitext",  # Calibration dataset
    bits=4.0,  # Bits per weight
    head_bits=6,  # Higher precision for attention
)
```

### Command Line

```bash
python convert.py \
    -i ./llama-3.1-8b-hf \
    -o ./llama-3.1-8b-exl2 \
    -cf ./llama-3.1-8b-exl2 \
    -b 4.0 \
    -hb 6
```

## Memory Management

### Cache Allocation

```python

# Fixed cache size
cache = ExLlamaV2Cache(model, max_seq_len=4096)

# Dynamic cache
cache = ExLlamaV2Cache(model, lazy=True)
cache.current_seq_len = 0  # Grows as needed
```

### Multi-GPU

```python
config = ExLlamaV2Config()
config.model_dir = "./large-model"

# Split across GPUs
config.set_auto_split([0.5, 0.5])  # 50% each GPU

model = ExLlamaV2(config)
model.load()
```

## Performance Comparison

| Model        | Engine    | GPU      | Tokens/sec |
| ------------ | --------- | -------- | ---------- |
| Llama 3.1 8B | ExLlamaV2 | RTX 3090 | \~150      |
| Llama 3.1 8B | llama.cpp | RTX 3090 | \~100      |
| Llama 3.1 8B | vLLM      | RTX 3090 | \~120      |
| Llama 3.1 8B | ExLlamaV2 | RTX 3090 | \~90       |
| Mixtral 8x7B | ExLlamaV2 | A100     | \~70       |

## Advanced Settings

### Sampling Parameters

```python
settings = ExLlamaV2Sampler.Settings()
settings.temperature = 0.7
settings.top_k = 50
settings.top_p = 0.9
settings.token_repetition_penalty = 1.1
settings.token_frequency_penalty = 0.0
settings.token_presence_penalty = 0.0
settings.mirostat = False
settings.mirostat_tau = 5.0
settings.mirostat_eta = 0.1
```

### Batch Generation

```python
prompts = [
    "The meaning of life is",
    "Artificial intelligence will",
    "Climate change is"
]

outputs = []
for prompt in prompts:
    output = generator.generate_simple(prompt, settings, num_tokens=100)
    outputs.append(output)
```

## Troubleshooting

### CUDA Out of Memory

```python

# Use smaller cache
cache = ExLlamaV2Cache(model, max_seq_len=2048)

# Or lower bpw model (3.0 instead of 4.0)
```

### Slow Loading

```python

# Enable fast loading
config.fasttensors = True
```

### Model Not Found

```bash

# Check model files exist
ls ./model/

# Should contain: config.json, *.safetensors, tokenizer.json
```

## Integration with LangChain

```python
from langchain.llms.base import LLM
from typing import Optional, List

class ExLlamaV2LLM(LLM):
    model: ExLlamaV2
    tokenizer: ExLlamaV2Tokenizer
    generator: ExLlamaV2StreamingGenerator
    settings: ExLlamaV2Sampler.Settings

    @property
    def _llm_type(self) -> str:
        return "exllamav2"

    def _call(self, prompt: str, stop: Optional[List[str]] = None) -> str:
        return self.generator.generate_simple(prompt, self.settings, num_tokens=500)

# Usage
llm = ExLlamaV2LLM(model=model, tokenizer=tokenizer, generator=generator, settings=settings)
result = llm("What is quantum computing?")
```

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* vLLM Inference - High throughput serving
* [llama.cpp Server](/guides/language-models/llamacpp-server) - Cross-platform
* [Text Generation WebUI](/guides/language-models/text-generation-webui) - Web interface


# LocalAI

Self-hosted OpenAI-compatible API with LocalAI on Clore.ai

Run a self-hosted OpenAI-compatible API with LocalAI.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Server Requirements

| Parameter    | Minimum          | Recommended |
| ------------ | ---------------- | ----------- |
| RAM          | 8GB              | 16GB+       |
| VRAM         | 6GB              | 8GB+        |
| Network      | 200Mbps          | 500Mbps+    |
| Startup Time | **5-10 minutes** | -           |

{% hint style="warning" %}
**Important:** LocalAI takes 5-10 minutes to fully initialize on first startup. HTTP 502 during this time is normal - the service is downloading and loading models.
{% endhint %}

{% hint style="info" %}
LocalAI is lightweight. For running LLMs (7B+ models), choose servers with 16GB+ RAM and 8GB+ VRAM.
{% endhint %}

## What is LocalAI?

LocalAI provides:

* Drop-in OpenAI API replacement
* Support for multiple model formats
* Text, image, audio, and embedding generation
* No GPU required (but faster with GPU)

## Supported Models

| Type       | Formats     | Examples            |
| ---------- | ----------- | ------------------- |
| LLM        | GGUF, GGML  | Llama, Mistral, Phi |
| Embeddings | GGUF        | all-MiniLM, BGE     |
| Images     | Diffusers   | SD 1.5, SDXL        |
| Audio      | Whisper     | Speech-to-text      |
| TTS        | Piper, Bark | Text-to-speech      |

## Quick Deploy

**Docker Image:**

```
localai/localai:master-aio-gpu-nvidia-cuda-12
```

**Ports:**

```
22/tcp
8080/http
```

**No command needed** - server starts automatically.

### Verify It's Working

After deployment, find your `http_pub` URL in **My Orders** and test:

```bash
# Check if service is ready
curl https://your-http-pub.clorecloud.net/readyz

# List available models
curl https://your-http-pub.clorecloud.net/v1/models

# Get version
curl https://your-http-pub.clorecloud.net/version
```

{% hint style="warning" %}
If you get HTTP 502, wait 5-10 minutes - LocalAI takes longer to initialize than other services.
{% endhint %}

## Pre-Built Models

LocalAI comes with several models available out of the box:

| Model Name                 | Type       | Description             |
| -------------------------- | ---------- | ----------------------- |
| `gpt-4`                    | Chat       | General-purpose LLM     |
| `gpt-4o`                   | Chat       | General-purpose LLM     |
| `gpt-4o-mini`              | Chat       | Smaller, faster LLM     |
| `whisper-1`                | STT        | Speech-to-text          |
| `tts-1`                    | TTS        | Text-to-speech          |
| `text-embedding-ada-002`   | Embeddings | 384-dimensional vectors |
| `jina-reranker-v1-base-en` | Reranking  | Document reranking      |

{% hint style="info" %}
These models work immediately after startup without additional configuration.
{% endhint %}

## Accessing Your Service

When deployed on CLORE.AI, access LocalAI via the `http_pub` URL:

```bash
# Chat completion
curl https://your-http-pub.clorecloud.net/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "gpt-4",
        "messages": [{"role": "user", "content": "Hello!"}]
    }'
```

{% hint style="info" %}
All `localhost:8080` examples below work when connected via SSH. For external access, replace with your `https://your-http-pub.clorecloud.net/` URL.
{% endhint %}

## Docker Deploy (Alternative)

```bash
docker run -d \
    --gpus all \
    -p 8080:8080 \
    -v /workspace/models:/models \
    -e THREADS=4 \
    -e CONTEXT_SIZE=4096 \
    localai/localai:master-aio-gpu-nvidia-cuda-12
```

## Download Models

### From Model Gallery

LocalAI has a built-in model gallery:

```bash
# List available models
curl http://localhost:8080/models/available

# Install from gallery
curl http://localhost:8080/models/apply -d '{"id": "mistral-7b-instruct"}'
```

### From Hugging Face

```bash
mkdir -p /workspace/models

# Llama 3.1 8B GGUF
wget https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
    -O /workspace/models/llama-3.1-8b.gguf

# Mistral 7B GGUF
wget https://huggingface.co/bartowski/Mistral-7B-Instruct-v0.3-GGUF/resolve/main/Mistral-7B-Instruct-v0.3-Q4_K_M.gguf \
    -O /workspace/models/mistral-7b.gguf
```

## Model Configuration

Create YAML config for each model:

**models/llama-3.1-8b.yaml:**

```yaml
name: llama-3.1-8b
backend: llama-cpp
parameters:
  model: llama-3.1-8b.gguf
  context_size: 4096
  threads: 8
  gpu_layers: 35
template:
  chat: |
    {{.Input}}
    ### Response:
  completion: |
    {{.Input}}
```

## API Usage

### Chat Completions (OpenAI Compatible)

```python
import openai

# For external access, use your http_pub URL:
client = openai.OpenAI(
    base_url="https://your-http-pub.clorecloud.net/v1",
    api_key="not-needed"
)

# Or via SSH tunnel:
# client = openai.OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="mistral-7b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
    temperature=0.7,
    max_tokens=500
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="mistral-7b",
    messages=[{"role": "user", "content": "Write a poem about AI"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
```

### Embeddings

```python
response = client.embeddings.create(
    model="all-minilm",
    input="The quick brown fox jumps over the lazy dog"
)

embedding = response.data[0].embedding
print(f"Embedding dimension: {len(embedding)}")
```

### Image Generation

```python
response = client.images.generate(
    model="stablediffusion",
    prompt="a beautiful sunset over mountains",
    size="512x512",
    n=1
)

image_url = response.data[0].url
```

## cURL Examples

### Chat

```bash
curl https://your-http-pub.clorecloud.net/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "mistral-7b",
        "messages": [{"role": "user", "content": "Hello!"}]
    }'
```

### Embeddings

```bash
curl https://your-http-pub.clorecloud.net/v1/embeddings \
    -H "Content-Type: application/json" \
    -d '{
        "model": "text-embedding-ada-002",
        "input": "Your text here"
    }'
```

Response:

```json
{
  "data": [{"embedding": [0.1, -0.2, ...], "index": 0}],
  "model": "text-embedding-ada-002",
  "usage": {"prompt_tokens": 4, "total_tokens": 4}
}
```

### Text-to-Speech (TTS)

```bash
curl https://your-http-pub.clorecloud.net/v1/audio/speech \
    -H "Content-Type: application/json" \
    -d '{
        "model": "tts-1",
        "input": "Hello, welcome to LocalAI!",
        "voice": "alloy"
    }' \
    --output speech.wav
```

Available voices: `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`

### Speech-to-Text (STT)

```bash
curl https://your-http-pub.clorecloud.net/v1/audio/transcriptions \
    -F "file=@audio.mp3" \
    -F "model=whisper-1"
```

Response:

```json
{"text": "Transcribed text here..."}
```

### Reranking

Rerank documents by relevance to a query:

```bash
curl https://your-http-pub.clorecloud.net/v1/rerank \
    -H "Content-Type: application/json" \
    -d '{
        "model": "jina-reranker-v1-base-en",
        "query": "What is machine learning?",
        "documents": [
            "Machine learning is a subset of AI",
            "The weather is nice today",
            "Deep learning uses neural networks"
        ],
        "top_n": 2
    }'
```

Response:

```json
{
  "results": [
    {"index": 0, "relevance_score": 0.95},
    {"index": 2, "relevance_score": 0.82}
  ]
}
```

## Complete API Reference

### Standard Endpoints (OpenAI Compatible)

| Endpoint                   | Method | Description           |
| -------------------------- | ------ | --------------------- |
| `/v1/models`               | GET    | List available models |
| `/v1/chat/completions`     | POST   | Chat completion       |
| `/v1/completions`          | POST   | Text completion       |
| `/v1/embeddings`           | POST   | Generate embeddings   |
| `/v1/audio/speech`         | POST   | Text-to-speech        |
| `/v1/audio/transcriptions` | POST   | Speech-to-text        |
| `/v1/images/generations`   | POST   | Image generation      |

### Additional Endpoints

| Endpoint            | Method | Description                |
| ------------------- | ------ | -------------------------- |
| `/readyz`           | GET    | Readiness check            |
| `/healthz`          | GET    | Health check               |
| `/version`          | GET    | Get LocalAI version        |
| `/v1/rerank`        | POST   | Document reranking         |
| `/models/available` | GET    | List gallery models        |
| `/models/apply`     | POST   | Install model from gallery |
| `/swagger/`         | GET    | Swagger UI documentation   |
| `/metrics`          | GET    | Prometheus metrics         |

#### Get Version

```bash
curl https://your-http-pub.clorecloud.net/version
```

Response:

```json
{"version": "v2.26.0"}
```

#### Swagger Documentation

Open in browser for interactive API documentation:

```
https://your-http-pub.clorecloud.net/swagger/
```

## GPU Acceleration

### CUDA Backend

```yaml
# In model config
parameters:
  gpu_layers: 35  # Number of layers on GPU
  f16: true       # Use FP16
```

### Full GPU Offload

```yaml
parameters:
  gpu_layers: 99  # All layers on GPU
  main_gpu: 0     # Primary GPU ID
```

## Multiple Models

LocalAI can serve multiple models simultaneously:

```
models/
├── llama-3.1-8b.yaml
├── llama-3.1-8b.gguf
├── mistral-7b.yaml
├── mistral-7b.gguf
├── whisper.yaml
└── whisper-base.bin
```

Access each via model name in API calls.

## Performance Tuning

### For Speed

```yaml
parameters:
  threads: 8
  gpu_layers: 99
  batch_size: 512
  use_mmap: true
  use_mlock: true
```

### For Memory

```yaml
parameters:
  gpu_layers: 20  # Partial offload
  context_size: 2048  # Smaller context
  batch_size: 256
```

## Benchmarks

| Model           | GPU      | Tokens/sec |
| --------------- | -------- | ---------- |
| Llama 3.1 8B Q4 | RTX 3090 | \~100      |
| Mistral 7B Q4   | RTX 3090 | \~110      |
| Llama 3.1 8B Q4 | RTX 4090 | \~140      |
| Mixtral 8x7B Q4 | A100     | \~60       |

*Benchmarks updated January 2026.*

## Troubleshooting

### HTTP 502 on http\_pub URL

LocalAI takes longer to start than other services. Wait **5-10 minutes** and retry:

```bash
# Check readiness
curl https://your-http-pub.clorecloud.net/readyz

# Check health
curl https://your-http-pub.clorecloud.net/healthz
```

### Model Not Loading

* Check file path in YAML
* Verify GGUF format compatibility
* Check available VRAM

### Slow Responses

* Increase `gpu_layers`
* Enable `use_mmap`
* Reduce `context_size`

### Out of Memory

* Reduce `gpu_layers`
* Use smaller quantization (Q4 instead of Q8)
* Reduce batch size

### Image Generation Issues

{% hint style="warning" %}
Stable Diffusion may have CUDA compatibility issues on some GPU configurations. If you encounter CUDA errors with image generation, consider using a dedicated Stable Diffusion image instead.
{% endhint %}

## Cost Estimate

Typical CLORE.AI marketplace rates:

| GPU      | VRAM | Price/day  | Good For       |
| -------- | ---- | ---------- | -------------- |
| RTX 3060 | 12GB | $0.15–0.30 | 7B models      |
| RTX 3090 | 24GB | $0.30–1.00 | 13B models     |
| RTX 4090 | 24GB | $0.50–2.00 | Fast inference |
| A100     | 40GB | $1.50–3.00 | Large models   |

*Prices in USD/day. Rates vary by provider — check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

## Next Steps

* [vLLM Inference](/guides/language-models/vllm) - Higher throughput
* [Ollama Guide](/guides/language-models/ollama) - Simpler setup
* [RAG with LangChain](/guides/advanced/api-integration) - Build applications


# Llama 3.3 70B

Run Meta's Llama 3.3 70B model on Clore.ai GPUs

{% hint style="info" %}
**Newer version available!** Meta released [**Llama 4**](/guides/language-models/llama4) in April 2025 with MoE architecture — Scout (17B active, fits on RTX 4090) delivers similar quality at a fraction of the VRAM. Consider upgrading.
{% endhint %}

Meta's latest and most efficient 70B model on CLORE.AI GPUs.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Why Llama 3.3?

* **Best 70B model** - Matches Llama 3.1 405B performance at fraction of cost
* **Multilingual** - Supports 8 languages natively
* **128K context** - Long document processing
* **Open weights** - Free for commercial use

## Model Overview

| Spec           | Value                          |
| -------------- | ------------------------------ |
| Parameters     | 70B                            |
| Context Length | 128K tokens                    |
| Training Data  | 15T+ tokens                    |
| Languages      | EN, DE, FR, IT, PT, HI, ES, TH |
| License        | Llama 3.3 Community License    |

### Performance vs Other Models

| Benchmark    | Llama 3.3 70B | Llama 3.1 405B | GPT-4o |
| ------------ | ------------- | -------------- | ------ |
| MMLU         | 86.0          | 87.3           | 88.7   |
| HumanEval    | 88.4          | 89.0           | 90.2   |
| MATH         | 77.0          | 73.8           | 76.6   |
| Multilingual | 91.1          | 91.6           | -      |

## GPU Requirements

| Setup        | VRAM  | Performance | Cost                      |
| ------------ | ----- | ----------- | ------------------------- |
| Q4 quantized | 40GB  | Good        | A100 40GB (\~$0.17/hr)    |
| Q8 quantized | 70GB  | Better      | A100 80GB (\~$0.25/hr)    |
| FP16 full    | 140GB | Best        | 2x A100 80GB (\~$0.50/hr) |

**Recommended:** A100 40GB with Q4 quantization for best price/performance.

## Quick Deploy on CLORE.AI

### Using Ollama (Easiest)

**Docker Image:**

```
ollama/ollama
```

**Ports:**

```
22/tcp
11434/http
```

**After deploy:**

```bash
ollama pull llama3.3
ollama run llama3.3
```

### Using vLLM (Production)

**Docker Image:**

```
vllm/vllm-openai:latest
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --tensor-parallel-size 1 \
    --max-model-len 32768 \
    --host 0.0.0.0
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Installation Methods

### Method 1: Ollama (Recommended for Testing)

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Pull Llama 3.3 (auto-downloads Q4 version)
ollama pull llama3.3

# Run interactively
ollama run llama3.3

# Or serve API
ollama serve
```

**API usage:**

```bash
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.3",
  "prompt": "Explain quantum computing in simple terms"
}'
```

### Method 2: vLLM (Production)

```bash
pip install vllm

# Single GPU (A100 40GB with AWQ quantization)
python -m vllm.entrypoints.openai.api_server \
    --model casperhansen/llama-3.3-70b-instruct-awq \
    --quantization awq \
    --max-model-len 16384 \
    --host 0.0.0.0

# Multi-GPU (2x A100 for full precision)
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --tensor-parallel-size 2 \
    --max-model-len 32768 \
    --host 0.0.0.0
```

**API usage (OpenAI-compatible):**

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")

response = client.chat.completions.create(
    model="meta-llama/Llama-3.3-70B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Write a Python function to calculate fibonacci numbers"}
    ],
    temperature=0.7,
    max_tokens=1024
)

print(response.choices[0].message.content)
```

### Method 3: Transformers + bitsandbytes

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)

model_id = "meta-llama/Llama-3.3-70B-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=bnb_config,
    device_map="auto"
)

# Generate
messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "Write a Python web scraper using BeautifulSoup"}
]

input_ids = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt"
).to("cuda")

outputs = model.generate(
    input_ids,
    max_new_tokens=512,
    temperature=0.7,
    do_sample=True
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

### Method 4: llama.cpp (CPU+GPU hybrid)

```bash
# Clone and build
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_CUDA=1

# Download GGUF model
wget https://huggingface.co/bartowski/Llama-3.3-70B-Instruct-GGUF/resolve/main/Llama-3.3-70B-Instruct-Q4_K_M.gguf

# Run server
./llama-server \
    -m Llama-3.3-70B-Instruct-Q4_K_M.gguf \
    -c 8192 \
    -ngl 80 \
    --host 0.0.0.0 \
    --port 8080
```

## Benchmarks

### Throughput (tokens/second)

| GPU          | Q4    | Q8    | FP16  |
| ------------ | ----- | ----- | ----- |
| A100 40GB    | 25-30 | -     | -     |
| A100 80GB    | 35-40 | 25-30 | -     |
| 2x A100 80GB | 50-60 | 40-45 | 30-35 |
| H100 80GB    | 60-70 | 45-50 | 35-40 |

### Time to First Token (TTFT)

| GPU          | Q4       | FP16     |
| ------------ | -------- | -------- |
| A100 40GB    | 0.8-1.2s | -        |
| A100 80GB    | 0.6-0.9s | -        |
| 2x A100 80GB | 0.4-0.6s | 0.8-1.0s |

### Context Length vs VRAM

| Context | Q4 VRAM | Q8 VRAM |
| ------- | ------- | ------- |
| 4K      | 38GB    | 72GB    |
| 8K      | 40GB    | 75GB    |
| 16K     | 44GB    | 80GB    |
| 32K     | 52GB    | 90GB    |
| 64K     | 68GB    | 110GB   |
| 128K    | 100GB   | 150GB   |

## Use Cases

### Code Generation

```python
messages = [
    {"role": "system", "content": "You are an expert programmer. Write clean, efficient, well-documented code."},
    {"role": "user", "content": "Create a REST API in FastAPI with user authentication using JWT tokens"}
]
```

### Document Analysis (Long Context)

```python
# Load long document
with open("large_document.txt") as f:
    document = f.read()

messages = [
    {"role": "system", "content": "You are a document analyst. Provide detailed, accurate analysis."},
    {"role": "user", "content": f"Analyze this document and provide a summary with key points:\n\n{document}"}
]
```

### Multilingual Tasks

```python
messages = [
    {"role": "system", "content": "You are a multilingual assistant."},
    {"role": "user", "content": "Translate this to German, French, and Spanish: 'The quick brown fox jumps over the lazy dog'"}
]
```

### Reasoning & Analysis

```python
messages = [
    {"role": "system", "content": "Think step by step. Show your reasoning."},
    {"role": "user", "content": "A train leaves Station A at 9:00 AM traveling at 60 mph. Another train leaves Station B (300 miles away) at 10:00 AM traveling toward Station A at 90 mph. When and where do they meet?"}
]
```

## Optimization Tips

### Memory Optimization

```python
# vLLM with memory optimization
python -m vllm.entrypoints.openai.api_server \
    --model casperhansen/llama-3.3-70b-instruct-awq \
    --quantization awq \
    --gpu-memory-utilization 0.95 \
    --max-model-len 8192
```

### Speed Optimization

```python
# Enable Flash Attention
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --tensor-parallel-size 2 \
    --enable-prefix-caching
```

### Batch Processing

```python
# Process multiple requests efficiently
responses = client.chat.completions.create(
    model="meta-llama/Llama-3.3-70B-Instruct",
    messages=messages,
    n=4,  # Generate 4 responses
    temperature=0.8
)
```

## Comparison with Other Models

| Feature   | Llama 3.3 70B | Llama 3.1 70B | Qwen 2.5 72B | Mixtral 8x22B |
| --------- | ------------- | ------------- | ------------ | ------------- |
| MMLU      | 86.0          | 83.6          | 85.3         | 77.8          |
| Coding    | 88.4          | 80.5          | 85.4         | 75.5          |
| Math      | 77.0          | 68.0          | 80.0         | 60.0          |
| Context   | 128K          | 128K          | 128K         | 64K           |
| Languages | 8             | 8             | 29           | 8             |
| License   | Open          | Open          | Open         | Open          |

**Verdict:** Llama 3.3 70B offers the best overall performance in its class, especially for coding and reasoning tasks.

## Troubleshooting

### Out of Memory

```bash
# Use AWQ quantization (most memory efficient)
--model casperhansen/llama-3.3-70b-instruct-awq --quantization awq

# Reduce context length
--max-model-len 8192

# Use tensor parallelism
--tensor-parallel-size 2
```

### Slow First Response

* First request loads model to GPU - wait 30-60 seconds
* Use `--enable-prefix-caching` for faster subsequent requests
* Pre-warm with dummy request

### Hugging Face Access

```bash
# Login to HF (required for gated model)
huggingface-cli login

# Or set environment variable
export HUGGING_FACE_HUB_TOKEN=hf_xxxxx
```

## Cost Estimate

| Setup       | GPU            | $/hour  | tokens/$ |
| ----------- | -------------- | ------- | -------- |
| Budget      | A100 40GB (Q4) | \~$0.17 | \~530K   |
| Balanced    | A100 80GB (Q4) | \~$0.25 | \~500K   |
| Performance | 2x A100 80GB   | \~$0.50 | \~360K   |
| Maximum     | H100 80GB      | \~$0.50 | \~500K   |

## Next Steps

* [vLLM Guide](/guides/language-models/vllm) - Production deployment
* [Ollama Guide](/guides/language-models/ollama) - Easy local setup
* [Multi-GPU Setup](/guides/advanced/multi-gpu-setup) - Scale to larger models
* [API Integration](/guides/advanced/api-integration) - Build applications


# Mistral & Mixtral

Run Mistral and Mixtral models on Clore.ai GPUs

{% hint style="info" %}
**Newer versions available!** Check out [**Mistral Small 3.1**](/guides/language-models/mistral-small) (24B, Apache 2.0, fits on RTX 4090) and [**Mistral Large 3**](/guides/language-models/mistral-large3) (675B MoE, frontier-class).
{% endhint %}

Run Mistral and Mixtral models for high-quality text generation.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## Model Overview

| Model               | Parameters           | VRAM  | Specialty         |
| ------------------- | -------------------- | ----- | ----------------- |
| Mistral-7B          | 7B                   | 8GB   | General purpose   |
| Mistral-7B-Instruct | 7B                   | 8GB   | Chat/instruction  |
| Mixtral-8x7B        | 46.7B (12.9B active) | 24GB  | MoE, best quality |
| Mixtral-8x22B       | 141B                 | 80GB+ | Largest MoE       |

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
pip install vllm && \
python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-Instruct-v0.2 \
    --port 8000
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Installation Options

### Using Ollama (Easiest)

```bash

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run Mistral
ollama run mistral

# Run Mixtral
ollama run mixtral
```

### Using vLLM

```bash
pip install vllm

# Start server
python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-Instruct-v0.2 \
    --dtype float16
```

### Using Transformers

```bash
pip install transformers accelerate
```

## Mistral-7B with Transformers

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "mistralai/Mistral-7B-Instruct-v0.2"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Explain quantum computing in simple terms"}
]

inputs = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt"
).to("cuda")

outputs = model.generate(
    inputs,
    max_new_tokens=500,
    do_sample=True,
    temperature=0.7,
    top_p=0.95
)

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

## Mixtral-8x7B

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "mistralai/Mixtral-8x7B-Instruct-v0.1"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Write a Python function to calculate fibonacci numbers"}
]

inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to("cuda")

outputs = model.generate(
    inputs,
    max_new_tokens=1000,
    do_sample=True,
    temperature=0.7
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## Quantized Models (Lower VRAM)

### 4-bit Quantization

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_quant_type="nf4"
)

model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mixtral-8x7B-Instruct-v0.1",
    quantization_config=quantization_config,
    device_map="auto"
)
```

### GGUF with llama.cpp

```bash

# Download GGUF model
wget https://huggingface.co/bartowski/Mistral-7B-Instruct-v0.3-GGUF/resolve/main/Mistral-7B-Instruct-v0.3-Q4_K_M.gguf

# Run with llama.cpp
./main -m Mistral-7B-Instruct-v0.3-Q4_K_M.gguf \
    -p "Explain machine learning" \
    -n 500
```

## vLLM Server (Production)

```bash
python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-Instruct-v0.2 \
    --dtype float16 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9
```

### OpenAI-Compatible API

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ],
    temperature=0.7,
    max_tokens=500
)

print(response.choices[0].message.content)
```

## Streaming

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

stream = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[{"role": "user", "content": "Write a story about a robot"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

## Function Calling

Mistral supports function calling:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get weather for a location",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string"},
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
                },
                "required": ["location"]
            }
        }
    }
]

response = client.chat.completions.create(
    model="mistralai/Mistral-7B-Instruct-v0.2",
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
    tools=tools
)

print(response.choices[0].message.tool_calls)
```

## Gradio Interface

```python
import gradio as gr
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "mistralai/Mistral-7B-Instruct-v0.2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

def chat(message, history, temperature, max_tokens):
    messages = []
    for h in history:
        messages.append({"role": "user", "content": h[0]})
        messages.append({"role": "assistant", "content": h[1]})
    messages.append({"role": "user", "content": message})

    inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to("cuda")

    outputs = model.generate(
        inputs,
        max_new_tokens=max_tokens,
        temperature=temperature,
        do_sample=True
    )

    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
    # Extract assistant response
    return response.split("[/INST]")[-1].strip()

demo = gr.ChatInterface(
    fn=chat,
    additional_inputs=[
        gr.Slider(0.1, 2.0, value=0.7, label="Temperature"),
        gr.Slider(100, 2000, value=500, step=100, label="Max Tokens")
    ],
    title="Mistral-7B Chat"
)

demo.launch(server_name="0.0.0.0", server_port=7860)
```

## Performance Comparison

### Throughput (tokens/sec)

| Model             | RTX 3060 | RTX 3090 | RTX 4090 | A100 40GB |
| ----------------- | -------- | -------- | -------- | --------- |
| Mistral-7B FP16   | 45       | 80       | 120      | 150       |
| Mistral-7B Q4     | 70       | 110      | 160      | 200       |
| Mixtral-8x7B FP16 | -        | -        | 30       | 60        |
| Mixtral-8x7B Q4   | -        | 25       | 50       | 80        |
| Mixtral-8x22B Q4  | -        | -        | -        | 25        |

### Time to First Token (TTFT)

| Model         | RTX 3090 | RTX 4090 | A100  |
| ------------- | -------- | -------- | ----- |
| Mistral-7B    | 80ms     | 50ms     | 35ms  |
| Mixtral-8x7B  | -        | 150ms    | 90ms  |
| Mixtral-8x22B | -        | -        | 200ms |

### Context Length vs VRAM (Mistral-7B)

| Context | FP16 | Q8   | Q4   |
| ------- | ---- | ---- | ---- |
| 4K      | 15GB | 9GB  | 5GB  |
| 8K      | 18GB | 11GB | 7GB  |
| 16K     | 24GB | 15GB | 9GB  |
| 32K     | 36GB | 22GB | 14GB |

## VRAM Requirements

| Model         | FP16  | 8-bit | 4-bit |
| ------------- | ----- | ----- | ----- |
| Mistral-7B    | 14GB  | 8GB   | 5GB   |
| Mixtral-8x7B  | 90GB  | 45GB  | 24GB  |
| Mixtral-8x22B | 180GB | 90GB  | 48GB  |

## Use Cases

### Code Generation

```python
prompt = """
Write a Python class for a REST API client with:
- Authentication handling
- Retry logic
- Error handling
"""
```

### Data Analysis

```python
prompt = """
Analyze this data and provide insights:
Sales Q1: $100K
Sales Q2: $150K
Sales Q3: $120K
Sales Q4: $200K
"""
```

### Creative Writing

```python
prompt = """
Write a short story about an AI that becomes self-aware,
in the style of Isaac Asimov.
"""
```

## Troubleshooting

### Out of Memory

* Use 4-bit quantization
* Use Mistral-7B instead of Mixtral
* Reduce max\_model\_len

### Slow Generation

* Use vLLM for production
* Enable flash attention
* Use tensor parallelism for multi-GPU

### Poor Output Quality

* Adjust temperature (0.1-0.9)
* Use instruct variant
* Better system prompts

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* [vLLM](/guides/language-models/vllm) - Production serving
* [Ollama](/guides/language-models/ollama) - Easy deployment
* [DeepSeek-V3](/guides/language-models/deepseek-v3) - Best reasoning model
* [Qwen2.5](/guides/language-models/qwen25) - Multilingual alternative


# DeepSeek Coder

Best-in-class code generation with DeepSeek Coder on Clore.ai

{% hint style="info" %}
**Newer versions available!** [**DeepSeek-R1**](/guides/language-models/deepseek-r1) (reasoning + coding) and [**DeepSeek-V3**](/guides/language-models/deepseek-v3) (general purpose) are significantly more capable. Also see [**Qwen2.5-Coder**](/guides/language-models/qwen25) for a strong coding alternative.
{% endhint %}

Best-in-class code generation with DeepSeek Coder models.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## What is DeepSeek Coder?

DeepSeek Coder offers:

* State-of-the-art code generation
* 338 programming languages
* Fill-in-the-middle support
* Repository-level understanding

## Model Variants

| Model               | Parameters | VRAM  | Context |
| ------------------- | ---------- | ----- | ------- |
| DeepSeek-Coder-1.3B | 1.3B       | 3GB   | 16K     |
| DeepSeek-Coder-6.7B | 6.7B       | 8GB   | 16K     |
| DeepSeek-Coder-33B  | 33B        | 40GB  | 16K     |
| DeepSeek-Coder-V2   | 16B/236B   | 20GB+ | 128K    |

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
pip install vllm && \
vllm serve deepseek-ai/deepseek-coder-6.7b-instruct --port 8000
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Using Ollama

```bash

# Run DeepSeek Coder
ollama run deepseek-coder

# Specific sizes
ollama run deepseek-coder:1.3b
ollama run deepseek-coder:6.7b
ollama run deepseek-coder:33b

# V2 (latest)
ollama run deepseek-coder-v2
```

## Installation

```bash
pip install transformers accelerate torch
```

## Code Generation

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "deepseek-ai/deepseek-coder-6.7b-instruct"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": """
Write a Python class for a REST API client with:
- Authentication support
- Retry logic with exponential backoff
- Request/response logging
"""}
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to("cuda")

outputs = model.generate(
    inputs,
    max_new_tokens=1024,
    temperature=0.2,
    do_sample=True
)

print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
```

## Fill-in-the-Middle (FIM)

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "deepseek-ai/deepseek-coder-6.7b-base"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

# Fill-in-the-middle format
prefix = """def calculate_statistics(data):
    \"\"\"Calculate mean, median, and std of a list.\"\"\"
    import statistics

    mean = statistics.mean(data)
"""

suffix = """
    return {
        'mean': mean,
        'median': median,
        'std': std
    }
"""

# FIM tokens
prompt = f"<｜fim▁begin｜>{prefix}<｜fim▁hole｜>{suffix}<｜fim▁end｜>"

inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=128)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## DeepSeek-Coder-V2

Latest and most powerful:

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Implement a thread-safe LRU cache in Python"}
]

inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to("cuda")
outputs = model.generate(inputs, max_new_tokens=1024, temperature=0.2)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
```

## vLLM Server

```bash
vllm serve deepseek-ai/deepseek-coder-6.7b-instruct \
    --port 8000 \
    --dtype bfloat16 \
    --max-model-len 16384 \
    --trust-remote-code
```

### API Usage

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

response = client.chat.completions.create(
    model="deepseek-ai/deepseek-coder-6.7b-instruct",
    messages=[
        {"role": "system", "content": "You are an expert programmer."},
        {"role": "user", "content": "Write a FastAPI websocket server"}
    ],
    temperature=0.2,
    max_tokens=1500
)

print(response.choices[0].message.content)
```

## Code Review

````python
code_to_review = """
def process_data(data):
    result = []
    for i in range(len(data)):
        if data[i] > 0:
            result.append(data[i] * 2)
    return result
"""

messages = [
    {"role": "user", "content": f"""
Review this code and suggest improvements:

```python
{code_to_review}
````

Focus on:

1. Performance
2. Readability
3. Best practices """} ]

````

## Bug Fixing

```python
buggy_code = """
def merge_sorted_lists(list1, list2):
    result = []
    i = j = 0
    while i < len(list1) and j < len(list2):
        if list1[i] < list2[j]:
            result.append(list1[i])
            i += 1
        else:
            result.append(list2[j])
    return result
"""

messages = [
    {"role": "user", "content": f"""
Find and fix the bug in this code:

```python
{buggy_code}
````

"""} ]

````

## Gradio Interface

```python
import gradio as gr
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "deepseek-ai/deepseek-coder-6.7b-instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)

def generate_code(prompt, temperature, max_tokens):
    messages = [{"role": "user", "content": prompt}]
    inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to("cuda")
    outputs = model.generate(inputs, max_new_tokens=max_tokens, temperature=temperature, do_sample=True)
    return tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)

demo = gr.Interface(
    fn=generate_code,
    inputs=[
        gr.Textbox(label="Prompt", lines=5, placeholder="Describe the code you need..."),
        gr.Slider(0.1, 1.0, value=0.2, label="Temperature"),
        gr.Slider(256, 2048, value=1024, step=128, label="Max Tokens")
    ],
    outputs=gr.Code(language="python", label="Generated Code"),
    title="DeepSeek Coder"
)

demo.launch(server_name="0.0.0.0", server_port=7860)
````

## Performance

| Model            | GPU      | Tokens/sec |
| ---------------- | -------- | ---------- |
| DeepSeek-1.3B    | RTX 3060 | \~120      |
| DeepSeek-6.7B    | RTX 3090 | \~70       |
| DeepSeek-6.7B    | RTX 4090 | \~100      |
| DeepSeek-33B     | A100     | \~40       |
| DeepSeek-V2-Lite | RTX 4090 | \~50       |

## Comparison

| Model              | HumanEval | Code Quality |
| ------------------ | --------- | ------------ |
| DeepSeek-Coder-33B | 79.3%     | Excellent    |
| CodeLlama-34B      | 53.7%     | Good         |
| GPT-3.5-Turbo      | 72.6%     | Good         |

## Troubleshooting

### Code completion not working

* Ensure correct prompt format with `<|fim_prefix|>`, `<|fim_suffix|>`, `<|fim_middle|>`
* Set appropriate `max_new_tokens` for code generation

### Model outputs garbage

* Check model is fully downloaded
* Verify CUDA is being used: `model.device`
* Try lower temperature (0.2-0.5 for code)

### Slow inference

* Use vLLM for 5-10x speedup
* Enable `torch.compile()` for transformers
* Use quantized model for large variants

### Import errors

* Install dependencies: `pip install transformers accelerate`
* Update PyTorch to 2.0+

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* [DeepSeek-V3](/guides/language-models/deepseek-v3) - Latest DeepSeek flagship model
* [CodeLlama](/guides/language-models/codellama) - Alternative code model
* [Qwen2.5-Coder](/guides/language-models/qwen25) - Alibaba's code model
* [vLLM](/guides/language-models/vllm) - Production deployment


# DeepSeek-V3

Run DeepSeek-V3 with exceptional reasoning on Clore.ai GPUs

Run DeepSeek-V3, the state-of-the-art open-source LLM with exceptional reasoning capabilities on CLORE.AI GPUs.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

{% hint style="info" %}
**Updated: DeepSeek-V3-0324 (March 2024)** — The latest revision of DeepSeek-V3 brings significant improvements in code generation, mathematical reasoning, and general problem-solving. See the [changelog](#whats-new-in-deepseek-v3-0324) section for details.
{% endhint %}

## Why DeepSeek-V3?

* **State-of-the-art** - Competes with GPT-4o and Claude 3.5 Sonnet
* **671B MoE** - 671B total params, 37B active per token (efficient inference)
* **Improved reasoning** - DeepSeek-V3-0324 is significantly better at math and code
* **Efficient** - MoE architecture reduces compute costs vs dense models
* **Open source** - Fully open weights under MIT license
* **Long context** - 128K token context window

## What's New in DeepSeek-V3-0324

DeepSeek-V3-0324 (March 2024 revision) introduces meaningful improvements across key domains:

### Code Generation

* **+8-12% on HumanEval** compared to original V3
* Better at multi-file codebases and complex refactoring tasks
* Improved understanding of modern frameworks (FastAPI, Pydantic v2, LangChain v0.3)
* More reliable at generating complete, runnable code without omissions

### Mathematical Reasoning

* **+5% on MATH-500** benchmark
* Better step-by-step proof construction
* Improved numerical accuracy for multi-step problems
* Enhanced ability to identify and correct mistakes mid-solution

### General Reasoning

* Stronger logical deduction and causal inference
* Better at multi-step planning tasks
* More consistent performance on edge cases and ambiguous prompts
* Improved instruction following on complex, multi-constraint requests

## Quick Deploy on CLORE.AI

**Docker Image:**

```
vllm/vllm-openai:latest
```

**Ports:**

```
22/tcp
8000/http
```

**Command (Multi-GPU Required):**

```bash
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V3-0324 \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 8 \
    --trust-remote-code
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

### Verify It's Working

```bash
# Check if service is ready
curl https://your-http-pub.clorecloud.net/health

# List available models
curl https://your-http-pub.clorecloud.net/v1/models

# Get version
curl https://your-http-pub.clorecloud.net/version
```

{% hint style="warning" %}
**Important:** DeepSeek-V3 requires **8x A100 80GB** GPUs and significant download time. HTTP 502 may persist for 15-30 minutes while the model downloads.
{% endhint %}

## Model Variants

| Model             | Parameters | Active | VRAM Required | HuggingFace                                                                                             |
| ----------------- | ---------- | ------ | ------------- | ------------------------------------------------------------------------------------------------------- |
| DeepSeek-V3-0324  | 671B       | 37B    | 8x80GB        | [deepseek-ai/DeepSeek-V3-0324](https://huggingface.co/deepseek-ai/DeepSeek-V3-0324)                     |
| DeepSeek-V3       | 671B       | 37B    | 8x80GB        | [deepseek-ai/DeepSeek-V3](https://huggingface.co/deepseek-ai/DeepSeek-V3)                               |
| DeepSeek-V3-Base  | 671B       | 37B    | 8x80GB        | [deepseek-ai/DeepSeek-V3-Base](https://huggingface.co/deepseek-ai/DeepSeek-V3-Base)                     |
| DeepSeek-V2.5     | 236B       | 21B    | 4x80GB        | [deepseek-ai/DeepSeek-V2.5](https://huggingface.co/deepseek-ai/DeepSeek-V2.5)                           |
| DeepSeek-V2-Lite  | 16B        | 2.4B   | 16GB          | [deepseek-ai/DeepSeek-V2-Lite](https://huggingface.co/deepseek-ai/DeepSeek-V2-Lite)                     |
| DeepSeek-Coder-V2 | 236B       | 21B    | 4x80GB        | [deepseek-ai/DeepSeek-Coder-V2-Instruct](https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Instruct) |

## Hardware Requirements

### Full Precision

| Model            | Minimum       | Recommended  |
| ---------------- | ------------- | ------------ |
| DeepSeek-V3-0324 | 8x A100 80GB  | 8x H100 80GB |
| DeepSeek-V2.5    | 4x A100 80GB  | 4x H100 80GB |
| DeepSeek-V2-Lite | RTX 4090 24GB | A100 40GB    |

### Quantized (AWQ/GPTQ)

| Model            | Quantization | VRAM   |
| ---------------- | ------------ | ------ |
| DeepSeek-V3-0324 | INT4         | 4x80GB |
| DeepSeek-V2.5    | INT4         | 2x80GB |
| DeepSeek-V2-Lite | INT4         | 8GB    |

## Installation

### Using vLLM (Recommended)

```bash
pip install vllm==0.7.3

# DeepSeek-V3-0324 (latest, 8 GPUs)
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V3-0324 \
    --tensor-parallel-size 8 \
    --trust-remote-code \
    --host 0.0.0.0 \
    --port 8000

# Original V3 (still available)
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V3 \
    --tensor-parallel-size 8 \
    --trust-remote-code \
    --host 0.0.0.0 \
    --port 8000
```

### Using Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepseek-ai/DeepSeek-V3-0324"

tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [{"role": "user", "content": "Explain quantum computing in simple terms."}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)

outputs = model.generate(inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

### Using Ollama

```bash
# Pull DeepSeek-V3 (requires significant resources)
ollama pull deepseek-v3

# Or lighter variant
ollama pull deepseek-coder-v2:16b

# Run
ollama run deepseek-v3
```

## API Usage

### OpenAI-Compatible API (vLLM)

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[
        {"role": "system", "content": "You are a helpful AI assistant."},
        {"role": "user", "content": "Write a Python function to find prime numbers."}
    ],
    temperature=0.7,
    max_tokens=1000
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[{"role": "user", "content": "Explain machine learning"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

### cURL

```bash
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "deepseek-ai/DeepSeek-V3-0324",
        "messages": [
            {"role": "user", "content": "What is the capital of France?"}
        ],
        "temperature": 0.7
    }'
```

## DeepSeek-V2-Lite (Single GPU)

For users with limited hardware:

```bash
# Using vLLM
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V2-Lite \
    --trust-remote-code \
    --host 0.0.0.0

# Using Ollama
ollama run deepseek-coder-v2:16b
```

```python
# Using Transformers on single GPU
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "deepseek-ai/DeepSeek-V2-Lite",
    torch_dtype=torch.float16,
    device_map="cuda",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V2-Lite", trust_remote_code=True)
```

## Code Generation

DeepSeek-V3-0324 is best-in-class for code:

```python
prompt = """Write a Python class for a binary search tree with:
- insert
- search
- delete
- in-order traversal
Include type hints and docstrings."""

response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.2  # Lower for code
)

print(response.choices[0].message.content)
```

Advanced code tasks where V3-0324 excels:

```python
# Multi-file refactoring
prompt = """I have a Flask application with all code in app.py (500 lines).
Refactor it to use the application factory pattern with blueprints for:
- auth (login, register, logout)
- api (REST endpoints)
- admin (dashboard)
Show the complete file structure and all files."""

response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.1,
    max_tokens=4000
)
```

## Math & Reasoning

```python
# Complex math problem
prompt = """Prove that for any integer n >= 1, the sum 1^2 + 2^2 + ... + n^2 = n(n+1)(2n+1)/6.
Use mathematical induction and show all steps clearly."""

response = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.1  # Very low for math
)

print(response.choices[0].message.content)
```

## Multi-GPU Configuration

### 8x GPU (Full Model — V3-0324)

```bash
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V3-0324 \
    --tensor-parallel-size 8 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.9 \
    --trust-remote-code
```

### 4x GPU (V2.5)

```bash
python -m vllm.entrypoints.openai.api_server \
    --model deepseek-ai/DeepSeek-V2.5 \
    --tensor-parallel-size 4 \
    --max-model-len 16384 \
    --trust-remote-code
```

## Performance

### Throughput (tokens/sec)

| Model                 | GPUs         | Context | Tokens/sec |
| --------------------- | ------------ | ------- | ---------- |
| DeepSeek-V3-0324      | 8x H100      | 32K     | \~85       |
| DeepSeek-V3-0324      | 8x A100 80GB | 32K     | \~52       |
| DeepSeek-V3-0324 INT4 | 4x A100 80GB | 16K     | \~38       |
| DeepSeek-V2.5         | 4x A100 80GB | 16K     | \~70       |
| DeepSeek-V2.5         | 2x A100 80GB | 8K      | \~45       |
| DeepSeek-V2-Lite      | RTX 4090     | 8K      | \~40       |
| DeepSeek-V2-Lite      | RTX 3090     | 4K      | \~25       |

### Time to First Token (TTFT)

| Model            | Configuration | TTFT     |
| ---------------- | ------------- | -------- |
| DeepSeek-V3-0324 | 8x H100       | \~750ms  |
| DeepSeek-V3-0324 | 8x A100       | \~1100ms |
| DeepSeek-V2.5    | 4x A100       | \~500ms  |
| DeepSeek-V2-Lite | RTX 4090      | \~150ms  |

### Memory Usage

| Model            | Precision | VRAM Required |
| ---------------- | --------- | ------------- |
| DeepSeek-V3-0324 | FP16      | 8x 80GB       |
| DeepSeek-V3-0324 | INT4      | 4x 80GB       |
| DeepSeek-V2.5    | FP16      | 4x 80GB       |
| DeepSeek-V2.5    | INT4      | 2x 80GB       |
| DeepSeek-V2-Lite | FP16      | 20GB          |
| DeepSeek-V2-Lite | INT4      | 10GB          |

## Benchmarks

### DeepSeek-V3-0324 vs Competition

| Benchmark         | V3-0324 | V3 (original) | GPT-4o | Claude 3.5 Sonnet |
| ----------------- | ------- | ------------- | ------ | ----------------- |
| MMLU              | 88.5%   | 87.1%         | 88.7%  | 88.3%             |
| HumanEval         | 90.2%   | 82.6%         | 90.2%  | 92.0%             |
| MATH-500          | 67.1%   | 61.6%         | 76.6%  | 71.1%             |
| GSM8K             | 92.1%   | 89.3%         | 95.8%  | 96.4%             |
| LiveCodeBench     | 72.4%   | 65.9%         | 71.3%  | 73.8%             |
| Codeforces Rating | 1850    | 1720          | 1780   | 1790              |

*Note: MATH-500 improvement from V3 → V3-0324 is +5.5 percentage points.*

## Docker Compose

```yaml
version: '3.8'

services:
  deepseek:
    image: vllm/vllm-openai:latest
    ports:
      - "8000:8000"
    volumes:
      - ~/.cache/huggingface:/root/.cache/huggingface
    environment:
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
    command: >
      --model deepseek-ai/DeepSeek-V2-Lite
      --host 0.0.0.0
      --port 8000
      --trust-remote-code
      --gpu-memory-utilization 0.9
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
```

## GPU Requirements Summary

| Use Case              | Recommended Setup  | Cost/Hour |
| --------------------- | ------------------ | --------- |
| Full DeepSeek-V3-0324 | 8x A100 80GB       | \~$2.00   |
| DeepSeek-V2.5         | 4x A100 80GB       | \~$1.00   |
| Development/Testing   | RTX 4090 (V2-Lite) | \~$0.10   |
| Production API        | 8x H100 80GB       | \~$3.00   |

## Cost Estimate

Typical CLORE.AI marketplace rates:

| GPU Configuration | Hourly Rate | Daily Rate |
| ----------------- | ----------- | ---------- |
| RTX 4090 24GB     | \~$0.10     | \~$2.30    |
| A100 40GB         | \~$0.17     | \~$4.00    |
| A100 80GB         | \~$0.25     | \~$6.00    |
| 4x A100 80GB      | \~$1.00     | \~$24.00   |
| 8x A100 80GB      | \~$2.00     | \~$48.00   |

*Prices vary by provider. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for development (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Use DeepSeek-V2-Lite for testing before scaling up

## Troubleshooting

### Out of Memory

```bash
# Reduce context length
--max-model-len 8192

# Or use quantization
--quantization awq

# For V2-Lite on 12GB GPU
--gpu-memory-utilization 0.85
--max-model-len 4096
```

### Model Download Slow

```bash
# Pre-download
huggingface-cli download deepseek-ai/DeepSeek-V3-0324

# Or use mirror
export HF_ENDPOINT=https://hf-mirror.com
```

### trust\_remote\_code Error

```bash
# Always include this flag for DeepSeek models
--trust-remote-code
```

### Multi-GPU Not Working

```bash
# Check NCCL
nvidia-smi topo -m

# Set NCCL variables
export NCCL_DEBUG=INFO
export NCCL_P2P_DISABLE=0
```

## DeepSeek vs Others

| Feature    | DeepSeek-V3-0324  | Llama 3.1 405B | Mixtral 8x22B     |
| ---------- | ----------------- | -------------- | ----------------- |
| Parameters | 671B (37B active) | 405B           | 176B (44B active) |
| Context    | 128K              | 128K           | 64K               |
| Code       | **Excellent**     | Great          | Good              |
| Math       | **Excellent**     | Good           | Good              |
| Min VRAM   | 8x80GB            | 8x80GB         | 2x80GB            |
| License    | MIT               | Llama 3.1      | Apache 2.0        |

**Use DeepSeek-V3 when:**

* Best reasoning performance needed
* Code generation is primary use
* Math/logic tasks are important
* Have multi-GPU setup available
* Want fully open-source weights (MIT license)

## Next Steps

* [vLLM](/guides/language-models/vllm) - Deployment server
* [DeepSeek-R1](/guides/language-models/deepseek-r1) - Reasoning-specialized variant
* [DeepSeek Coder](/guides/language-models/deepseek-coder) - Code-specific variant
* [Ollama](/guides/language-models/ollama) - Simpler deployment
* [Fine-tune LLM](/guides/training/finetune-llm) - Custom training


# DeepSeek-R1 Reasoning Model

Run DeepSeek-R1 open-source reasoning model on Clore.ai GPUs

{% hint style="success" %}
All examples run on GPU servers rented via the [CLORE.AI Marketplace](https://clore.ai/marketplace). RTX 4090 instances start at \~$0.50/day.
{% endhint %}

## Overview

DeepSeek-R1 is a 671B-parameter open-weight reasoning model released in January 2025 by DeepSeek under the **Apache 2.0** license. It is the first open model to match OpenAI o1 across math, coding, and scientific benchmarks — while exposing its entire chain-of-thought through explicit `<think>` tags.

The full model uses **Mixture-of-Experts (MoE)** with 37B active parameters per token, making inference tractable despite the headline parameter count. For most practitioners, the **distilled variants** (1.5B → 70B) are more practical: they inherit R1's reasoning patterns through knowledge distillation into Qwen-2.5 and Llama-3 base architectures and run on commodity GPUs.

## Key Features

* **Explicit chain-of-thought** — every response begins with a `<think>` block where the model reasons, backtracks, and self-corrects before producing a final answer
* **Reinforcement-learning trained** — reasoning ability emerges from RL reward signals rather than hand-authored chain-of-thought data
* **Six distilled variants** — 1.5B, 7B, 8B, 14B, 32B, 70B parameter models distilled from the full 671B into Qwen and Llama architectures
* **Apache 2.0 license** — fully commercial, no royalties, no usage restrictions
* **Wide framework support** — Ollama, vLLM, llama.cpp, SGLang, Transformers, TGI all work out of the box
* **AIME 2024 Pass\@1: 79.8%** — ties with OpenAI o1 on competition math
* **Codeforces 2029 Elo** — exceeds o1's 1891 on competitive programming

## Model Variants

| Variant                | Parameters        | Architecture | FP16 VRAM | Q4 VRAM  | Q4 Disk  |
| ---------------------- | ----------------- | ------------ | --------- | -------- | -------- |
| DeepSeek-R1 (full MoE) | 671B (37B active) | DeepSeek MoE | \~1.3 TB  | \~350 GB | \~340 GB |
| R1-Distill-Llama-70B   | 70B               | Llama 3      | 140 GB    | 40 GB    | 42 GB    |
| R1-Distill-Qwen-32B    | 32B               | Qwen 2.5     | 64 GB     | 22 GB    | 20 GB    |
| R1-Distill-Qwen-14B    | 14B               | Qwen 2.5     | 28 GB     | 10 GB    | 9 GB     |
| R1-Distill-Llama-8B    | 8B                | Llama 3      | 16 GB     | 6 GB     | 5.5 GB   |
| R1-Distill-Qwen-7B     | 7B                | Qwen 2.5     | 14 GB     | 5 GB     | 4.5 GB   |
| R1-Distill-Qwen-1.5B   | 1.5B              | Qwen 2.5     | 3 GB      | 2 GB     | 1.2 GB   |

### Choosing a Variant

| Use Case                              | Recommended Variant    | GPU on Clore                                                                                                                |
| ------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| Quick experiments, edge testing       | R1-Distill-Qwen-1.5B   | Any GPU                                                                                                                     |
| Budget deployment, fast inference     | R1-Distill-Qwen-7B     | RTX 3090 (\~$0.30–1/day)                                                                                                    |
| Single-GPU production sweet spot      | R1-Distill-Qwen-14B Q4 | [RTX 4090](https://clore.ai/rent-4090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=deepseek-r1) (\~$0.50–2/day)     |
| Best quality-per-dollar (recommended) | R1-Distill-Qwen-32B Q4 | [RTX 4090 24 GB](https://clore.ai/rent-4090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=deepseek-r1) or A100 40 GB |
| Maximum distilled quality             | R1-Distill-Llama-70B   | 2× A100 80 GB                                                                                                               |
| Research, full-fidelity reasoning     | DeepSeek-R1 671B       | 8× H100 cluster                                                                                                             |

### HuggingFace Repositories

| Variant           | Repository                                                                                                    |
| ----------------- | ------------------------------------------------------------------------------------------------------------- |
| Full R1           | [deepseek-ai/DeepSeek-R1](https://huggingface.co/deepseek-ai/DeepSeek-R1)                                     |
| Llama-70B distill | [deepseek-ai/DeepSeek-R1-Distill-Llama-70B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B) |
| Qwen-32B distill  | [deepseek-ai/DeepSeek-R1-Distill-Qwen-32B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B)   |
| Qwen-14B distill  | [deepseek-ai/DeepSeek-R1-Distill-Qwen-14B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B)   |
| Llama-8B distill  | [deepseek-ai/DeepSeek-R1-Distill-Llama-8B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B)   |
| Qwen-7B distill   | [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B)     |
| Qwen-1.5B distill | [deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B) |

## Requirements

| Component  | Minimum (7B Q4) | Recommended (32B Q4) |
| ---------- | --------------- | -------------------- |
| GPU VRAM   | 6 GB            | 24 GB                |
| System RAM | 16 GB           | 32 GB                |
| Disk       | 10 GB           | 30 GB                |
| CUDA       | 12.1+           | 12.4+                |
| Docker     | 24.0+           | 25.0+                |

## Ollama Quick Start

Ollama handles quantization, downloading, and serving automatically — the fastest path to a running DeepSeek-R1.

### Install and run

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Pull and run (pick your variant):
ollama run deepseek-r1:1.5b    # Edge / testing
ollama run deepseek-r1:7b      # Budget — fits any 8 GB+ GPU
ollama run deepseek-r1:14b     # RTX 4090 sweet spot
ollama run deepseek-r1:32b     # A100 / 2× RTX 4090
ollama run deepseek-r1:70b     # 2× A100 80 GB
```

### Example interactive session

```
>>> Prove that √2 is irrational.

<think>
I'll use proof by contradiction.
Assume √2 is rational, so √2 = p/q where p,q are integers with gcd(p,q) = 1.
Then 2 = p²/q², which gives p² = 2q².
This means p² is even, so p must be even. Write p = 2k.
Substituting: (2k)² = 2q² → 4k² = 2q² → q² = 2k².
So q² is also even, meaning q is even.
But if both p and q are even, gcd(p,q) ≥ 2, contradicting our assumption.
</think>

**Proof that √2 is irrational (by contradiction):**

Assume √2 = p/q in lowest terms (gcd(p,q) = 1). Squaring gives p² = 2q².
Since p² is even, p is even — write p = 2k. Then 4k² = 2q², so q² = 2k²,
meaning q is also even. This contradicts gcd(p,q) = 1. ∎
```

### Use the OpenAI-compatible API

```bash
# Start Ollama as a server (if not already running)
ollama serve &

# Query via curl
curl -s http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-r1:32b",
    "messages": [{"role": "user", "content": "Factor x^4 - 1 completely over the integers."}],
    "temperature": 0.6
  }' | python3 -m json.tool
```

### Python client (via OpenAI SDK)

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

response = client.chat.completions.create(
    model="deepseek-r1:32b",
    messages=[
        {"role": "user", "content": "Write a Python function to find the longest palindromic substring."}
    ],
    temperature=0.6,
    max_tokens=4096,
)
print(response.choices[0].message.content)
```

## vLLM Production Setup

vLLM delivers the highest throughput for multi-user serving with continuous batching, PagedAttention, and prefix caching.

### Single GPU — 7B / 14B

```bash
pip install vllm

# 7B on any 16 GB+ GPU
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-7B \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 16384

# 14B on RTX 4090 (24 GB)
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-14B \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.92
```

### Multi-GPU — 32B (recommended)

```bash
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
    --host 0.0.0.0 --port 8000 \
    --tensor-parallel-size 2 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching
```

> **Tip:** The 32B Q4 GPTQ or AWQ checkpoint fits on a single RTX 4090 (24 GB):
>
> ```bash
> vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
>     --quantization awq --host 0.0.0.0 --port 8000 \
>     --max-model-len 16384
> ```

### Multi-GPU — 70B

```bash
vllm serve deepseek-ai/DeepSeek-R1-Distill-Llama-70B \
    --host 0.0.0.0 --port 8000 \
    --tensor-parallel-size 4 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.90
```

### Query the vLLM endpoint

```bash
curl -s http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
    "messages": [{"role": "user", "content": "Solve: find all primes p such that p^2 + 2 is also prime."}],
    "temperature": 0.6,
    "max_tokens": 4096
  }'
```

## Transformers / Python (with `<think>` Tag Parsing)

Use HuggingFace Transformers when you need fine-grained control over generation or want to integrate R1 into a Python pipeline.

### Basic generation

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch, re

MODEL = "deepseek-ai/DeepSeek-R1-Distill-Qwen-7B"

tokenizer = AutoTokenizer.from_pretrained(MODEL, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    MODEL,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

prompt = "What is the sum of the first 100 positive integers?"
messages = [{"role": "user", "content": prompt}]
input_text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=2048,
        temperature=0.6,
        do_sample=True,
    )

full_response = tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
print(full_response)
```

### Parsing `<think>` tags

```python
def parse_r1_response(text: str) -> dict:
    """Split a DeepSeek-R1 response into thinking and answer parts."""
    think_match = re.search(r"<think>(.*?)</think>", text, re.DOTALL)
    thinking = think_match.group(1).strip() if think_match else ""
    answer = re.sub(r"<think>.*?</think>", "", text, flags=re.DOTALL).strip()
    return {
        "thinking": thinking,
        "answer": answer,
        "thinking_tokens": len(thinking.split()),
    }

result = parse_r1_response(full_response)
print(f"Model reasoned for {result['thinking_tokens']} words")
print(f"Answer: {result['answer']}")
```

### Streaming with `<think>` state tracking

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

stream = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
    messages=[{"role": "user", "content": "Derive the quadratic formula from ax² + bx + c = 0."}],
    stream=True,
    max_tokens=4096,
    temperature=0.6,
)

in_think = False
for chunk in stream:
    token = chunk.choices[0].delta.content or ""
    if "<think>" in token:
        in_think = True
        print("[Reasoning] ", end="", flush=True)
        continue
    if "</think>" in token:
        in_think = False
        print("\n[Answer] ", end="", flush=True)
        continue
    if not in_think:
        print(token, end="", flush=True)
print()
```

## Docker Deployment on Clore.ai

### Ollama Docker (simplest)

**Docker image:** `ollama/ollama` **Ports:** `22/tcp, 11434/http`

```bash
# On the Clore instance
docker run -d --gpus all \
    -v ollama_data:/root/.ollama \
    -p 11434:11434 \
    --name deepseek-r1 \
    ollama/ollama

# Pull and serve the model
docker exec deepseek-r1 ollama pull deepseek-r1:32b
```

### vLLM Docker (production)

**Docker image:** `vllm/vllm-openai:latest` **Ports:** `22/tcp, 8000/http`

```yaml
# docker-compose.yml
version: "3.8"
services:
  deepseek-r1:
    image: vllm/vllm-openai:latest
    ports:
      - "8000:8000"
    volumes:
      - hf_cache:/root/.cache/huggingface
    environment:
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
    command: >
      --model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
      --host 0.0.0.0 --port 8000
      --tensor-parallel-size 2
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --enable-prefix-caching
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 5
      start_period: 300s
volumes:
  hf_cache:
```

Deploy on Clore.ai:

1. Open [clore.ai/marketplace](https://clore.ai/marketplace)
2. Filter by **2× GPU, 48 GB+ VRAM total** (e.g. 2× RTX 4090 or A100 80 GB)
3. Set the Docker image to `vllm/vllm-openai:latest`
4. Map port **8000** as HTTP
5. Paste the command from the compose file above into the startup command
6. Connect via the HTTP endpoint once the health check passes

## Tips for Clore.ai Deployments

### Choosing the right GPU

| Budget       | GPU              | Daily Cost   | Best Variant                       |
| ------------ | ---------------- | ------------ | ---------------------------------- |
| Minimal      | RTX 3090 (24 GB) | $0.30 – 1.00 | R1-Distill-Qwen-7B or 14B Q4       |
| Standard     | RTX 4090 (24 GB) | $0.50 – 2.00 | R1-Distill-Qwen-14B FP16 or 32B Q4 |
| Production   | A100 80 GB       | $3 – 8       | R1-Distill-Qwen-32B FP16           |
| High quality | 2× A100 80 GB    | $6 – 16      | R1-Distill-Llama-70B FP16          |

### Performance tuning

* **Temperature 0.6** is the recommended default for reasoning tasks — DeepSeek's own papers use this value
* **Set `max_tokens` generously** — reasoning models produce long `<think>` blocks; 4096+ for non-trivial problems
* **Enable prefix caching** (`--enable-prefix-caching` in vLLM) when using a shared system prompt
* **Limit concurrency** (`--max-num-seqs 16`) for reasoning workloads — each request uses more compute than a standard chat
* **Use Q4 quantization** to fit 32B on a single 24 GB GPU with minimal quality loss (the distill already compresses R1's knowledge)

### Context length considerations

Reasoning models consume more context than standard chat models because of the `<think>` block:

| Task Complexity              | Typical Thinking Length | Total Context Needed |
| ---------------------------- | ----------------------- | -------------------- |
| Simple arithmetic            | \~100 tokens            | \~300 tokens         |
| Code generation              | \~500–1000 tokens       | \~2000 tokens        |
| Competition math (AIME)      | \~2000–4000 tokens      | \~5000 tokens        |
| Multi-step research analysis | \~4000–8000 tokens      | \~10000 tokens       |

## Troubleshooting

### Out of memory (OOM)

```bash
# Reduce context length
--max-model-len 8192    # instead of 32768

# Limit concurrent sequences
--max-num-seqs 8

# Use quantization
--quantization awq      # or gptq
```

### Model produces no `<think>` block

Some system prompts suppress thinking. Avoid instructions like "be concise" or "don't explain your reasoning." Use a minimal system prompt or none at all:

```python
# Good — preserves reasoning
messages = [{"role": "user", "content": "..."}]

# Bad — may suppress thinking
messages = [
    {"role": "system", "content": "Be extremely brief. No explanations."},
    {"role": "user", "content": "..."}
]
```

### Repetitive or looping `<think>` output

Lower the temperature to reduce randomness in the reasoning chain:

```python
temperature = 0.0   # Deterministic — best for math/code
temperature = 0.3   # Slight variation — good for analysis
```

### Slow first token (high TTFT)

This is expected — the model generates `<think>` tokens before the visible answer. For latency-sensitive applications where reasoning is not needed, use [DeepSeek-V3](/guides/language-models/deepseek-v3) instead.

### Download stalls on Clore instance

HuggingFace downloads can be slow on some providers. Pre-cache the model into a persistent volume:

```bash
# Download once into a volume
huggingface-cli download deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
    --local-dir /data/models/deepseek-r1-32b

# Point vLLM at local path
vllm serve /data/models/deepseek-r1-32b --host 0.0.0.0 --port 8000
```

## Further Reading

* [DeepSeek-R1 Paper](https://arxiv.org/abs/2501.12948) — *Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*
* [DeepSeek-R1 GitHub](https://github.com/deepseek-ai/DeepSeek-R1) — Official repository with model cards
* [DeepSeek-V3 Guide](/guides/language-models/deepseek-v3) — Non-reasoning general-purpose model from the same lab
* [vLLM Guide](/guides/language-models/vllm) — Comprehensive production serving setup
* [Ollama Guide](/guides/language-models/ollama) — Simple local deployment for any model
* [Open WebUI Guide](/guides/language-models/open-webui) — Chat UI with native `<think>` tag rendering
* [Qwen 2.5 Guide](/guides/language-models/qwen25) — The base architecture used by most R1 distills


# Qwen2.5

Run Alibaba's Qwen2.5 multilingual LLMs on Clore.ai GPUs

Run Alibaba's Qwen2.5 family of models - powerful multilingual LLMs with excellent code and math capabilities on CLORE.AI GPUs.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Why Qwen2.5?

* **Versatile sizes** - 0.5B to 72B parameters
* **Multilingual** - 29 languages including Chinese
* **Long context** - Up to 128K tokens
* **Specialized variants** - Coder, Math editions
* **Open source** - Apache 2.0 license

## Quick Deploy on CLORE.AI

**Docker Image:**

```
vllm/vllm-openai:latest
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-7B-Instruct \
    --host 0.0.0.0 \
    --port 8000
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

### Verify It's Working

```bash
# Check if service is ready
curl https://your-http-pub.clorecloud.net/health

# List available models
curl https://your-http-pub.clorecloud.net/v1/models
```

{% hint style="warning" %}
If you get HTTP 502, wait 5-15 minutes - the model is still downloading from HuggingFace.
{% endhint %}

## Qwen3 Reasoning Mode

{% hint style="info" %}
**New in Qwen3:** Some Qwen3 models support a reasoning mode that shows the model's thought process in `<think>` tags before the final answer.
{% endhint %}

When using Qwen3 models via vLLM, responses may include reasoning:

```json
{
  "content": "<think>\nLet me think about this step by step...\n</think>\n\nThe answer is..."
}
```

To use Qwen3 with reasoning:

```bash
vllm serve Qwen/Qwen3-0.6B --host 0.0.0.0 --port 8000
```

## Model Variants

### Base Models

| Model                | Parameters | VRAM (FP16) | Context | Notes               |
| -------------------- | ---------- | ----------- | ------- | ------------------- |
| Qwen2.5-0.5B         | 0.5B       | 2GB         | 32K     | Edge/testing        |
| Qwen2.5-1.5B         | 1.5B       | 4GB         | 32K     | Very light          |
| Qwen2.5-3B           | 3B         | 8GB         | 32K     | Budget              |
| Qwen2.5-7B           | 7B         | 16GB        | 128K    | Balanced            |
| Qwen2.5-14B          | 14B        | 32GB        | 128K    | High quality        |
| Qwen2.5-32B          | 32B        | 70GB        | 128K    | Very high quality   |
| Qwen2.5-72B          | 72B        | 150GB       | 128K    | **Best quality**    |
| Qwen2.5-72B-Instruct | 72B        | 150GB       | 128K    | Chat/instruct tuned |

### Specialized Variants

| Model                      | Focus       | Best For               | VRAM (FP16) |
| -------------------------- | ----------- | ---------------------- | ----------- |
| Qwen2.5-Coder-7B-Instruct  | Code        | Programming, debugging | 16GB        |
| Qwen2.5-Coder-14B-Instruct | Code        | Complex code tasks     | 32GB        |
| Qwen2.5-Coder-32B-Instruct | Code        | **Best code model**    | 70GB        |
| Qwen2.5-Math-7B-Instruct   | Mathematics | Calculations, proofs   | 16GB        |
| Qwen2.5-Math-72B-Instruct  | Mathematics | Research-grade math    | 150GB       |
| Qwen2.5-Instruct           | Chat        | General assistant      | varies      |

## Hardware Requirements

| Model     | Minimum GPU   | Recommended  | VRAM (Q4) |
| --------- | ------------- | ------------ | --------- |
| 0.5B-3B   | RTX 3060 12GB | RTX 3080     | 2-6GB     |
| 7B        | RTX 3090 24GB | RTX 4090     | 6GB       |
| 14B       | A100 40GB     | A100 80GB    | 12GB      |
| 32B       | A100 80GB     | 2x A100 40GB | 22GB      |
| 72B       | 2x A100 80GB  | 4x A100 80GB | 48GB      |
| Coder-32B | A100 80GB     | 2x A100 40GB | 22GB      |

## Installation

### Using vLLM (Recommended)

```bash
pip install vllm==0.7.3

python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-7B-Instruct \
    --host 0.0.0.0 \
    --port 8000
```

### Using Ollama

```bash
# Standard models
ollama pull qwen2.5:7b
ollama pull qwen2.5:14b
ollama pull qwen2.5:32b
ollama pull qwen2.5:72b       # New: largest Qwen2.5

# Specialized
ollama pull qwen2.5-coder:7b
ollama pull qwen2.5-coder:32b  # New: best code model

# Run chat
ollama run qwen2.5:7b
```

### Using Transformers

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "Qwen/Qwen2.5-7B-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto"
)

messages = [{"role": "user", "content": "Hello!"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## API Usage

### OpenAI-Compatible API

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain machine learning in simple terms."}
    ],
    temperature=0.7,
    max_tokens=500
)

print(response.choices[0].message.content)
```

### Streaming

```python
stream = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "Write a poem about AI"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

### cURL

```bash
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "Qwen/Qwen2.5-7B-Instruct",
        "messages": [
            {"role": "user", "content": "What is Python?"}
        ]
    }'
```

## Qwen2.5-72B-Instruct

The flagship Qwen2.5 model — the largest and most capable in the family. It competes with GPT-4 on many benchmarks and is fully open-source under Apache 2.0.

### Running via vLLM (Multi-GPU)

```bash
# 4x A100 80GB setup
vllm serve Qwen/Qwen2.5-72B-Instruct \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 4 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.9

# AWQ quantized — runs on 2x A100 80GB
vllm serve Qwen/Qwen2.5-72B-Instruct-AWQ \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 2 \
    --quantization awq \
    --max-model-len 32768
```

### Running via Ollama

```bash
# Pull 72B model (requires 48GB+ VRAM for Q4)
ollama pull qwen2.5:72b

# Run interactive session
ollama run qwen2.5:72b

# API access
curl http://localhost:11434/api/chat -d '{
  "model": "qwen2.5:72b",
  "messages": [{"role": "user", "content": "Analyze this complex scenario..."}],
  "stream": false
}'
```

### Python Example

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

# The 72B model excels at complex analytical tasks
response = client.chat.completions.create(
    model="Qwen/Qwen2.5-72B-Instruct",
    messages=[
        {
            "role": "system",
            "content": "You are an expert analyst. Provide detailed, nuanced responses."
        },
        {
            "role": "user",
            "content": """Compare the architectural differences between transformer and 
            state space models (SSMs) for sequence modeling. Include efficiency tradeoffs."""
        }
    ],
    temperature=0.7,
    max_tokens=2000
)

print(response.choices[0].message.content)
```

## Qwen2.5-Coder-32B-Instruct

The best open-source code model available. Qwen2.5-Coder-32B-Instruct matches or exceeds GPT-4o on many coding benchmarks, supporting 40+ programming languages.

### Running via vLLM

```bash
# Single A100 80GB
vllm serve Qwen/Qwen2.5-Coder-32B-Instruct \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.9

# Dual RTX 4090 (24GB each = 48GB total, using Q4 quantization)
vllm serve Qwen/Qwen2.5-Coder-32B-Instruct-AWQ \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 2 \
    --quantization awq
```

### Running via Ollama

```bash
# Pull Coder-32B (requires ~22GB VRAM for Q4)
ollama pull qwen2.5-coder:32b

# Run
ollama run qwen2.5-coder:32b

# Test with a coding prompt
ollama run qwen2.5-coder:32b "Write a Python async web scraper using aiohttp"
```

### Code Generation Examples

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

# Full-stack code generation
response = client.chat.completions.create(
    model="Qwen/Qwen2.5-Coder-32B-Instruct",
    messages=[
        {
            "role": "system",
            "content": "You are an expert software engineer. Write clean, production-ready code with proper error handling and documentation."
        },
        {
            "role": "user",
            "content": """Write a Python FastAPI service that:
1. Accepts POST /summarize with JSON body {"text": "...", "max_length": 150}
2. Uses a local Ollama instance to summarize the text
3. Returns {"summary": "...", "original_length": N, "summary_length": N}
4. Includes proper error handling, input validation with Pydantic, and async support"""
        }
    ],
    temperature=0.1,  # Low temperature for code
    max_tokens=3000
)

print(response.choices[0].message.content)
```

````python
# Code review and debugging
code_to_review = """
def find_duplicates(lst):
    seen = []
    duplicates = []
    for item in lst:
        if item in seen:
            duplicates.append(item)
        seen.append(item)
    return duplicates
"""

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-Coder-32B-Instruct",
    messages=[
        {
            "role": "user",
            "content": f"Review this Python code for performance issues and suggest improvements:\n\n```python\n{code_to_review}\n```"
        }
    ],
    temperature=0.3
)

print(response.choices[0].message.content)
````

## Qwen2.5-Coder

Optimized for code generation:

```bash
# Using vLLM
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-Coder-7B-Instruct \
    --host 0.0.0.0

# Using Ollama
ollama run qwen2.5-coder:7b
```

```python
prompt = """Write a Python function that:
1. Takes a list of numbers
2. Returns the median value
3. Handles empty lists gracefully
Include type hints and docstrings."""

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-Coder-7B-Instruct",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.2
)

print(response.choices[0].message.content)
```

## Qwen2.5-Math

Specialized for mathematical reasoning:

```bash
# Using vLLM
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-Math-7B-Instruct \
    --host 0.0.0.0
```

```python
prompt = """Solve step by step:
Find all values of x where: x^3 - 6x^2 + 11x - 6 = 0"""

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-Math-7B-Instruct",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.1
)

print(response.choices[0].message.content)
```

## Multilingual Support

Qwen2.5 supports 29 languages:

```python
# Chinese
response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "用中文解释什么是人工智能"}]
)

# Japanese
response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "人工知能について日本語で説明してください"}]
)

# Korean
response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "인공지능에 대해 한국어로 설명해주세요"}]
)
```

## Long Context (128K)

```python
# Read a long document
with open("long_document.txt", "r") as f:
    document = f.read()

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[
        {"role": "user", "content": f"Summarize this document:\n\n{document}"}
    ],
    max_tokens=2000
)
```

## Quantization

### GGUF with Ollama

```bash
# 4-bit quantized
ollama pull qwen2.5:7b-instruct-q4_K_M
ollama pull qwen2.5:72b-instruct-q4_K_M   # 72B in 4-bit (~48GB)

# 8-bit quantized
ollama pull qwen2.5:7b-instruct-q8_0

# Coder variants
ollama pull qwen2.5-coder:32b-instruct-q4_K_M
```

### AWQ with vLLM

```bash
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-72B-Instruct-AWQ \
    --quantization awq \
    --tensor-parallel-size 2
```

### GGUF with llama.cpp

```bash
# Download GGUF
wget https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-GGUF/resolve/main/qwen2.5-7b-instruct-q4_k_m.gguf

# Run server
./llama-server -m qwen2.5-7b-instruct-q4_k_m.gguf \
    --host 0.0.0.0 \
    --port 8080 \
    -ngl 35
```

## Multi-GPU Setup

### Tensor Parallelism

```bash
# 72B on 4 GPUs
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-72B-Instruct \
    --tensor-parallel-size 4 \
    --max-model-len 32768

# 32B on 2 GPUs
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-32B-Instruct \
    --tensor-parallel-size 2

# Coder-32B on 2 GPUs
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-Coder-32B-Instruct \
    --tensor-parallel-size 2 \
    --max-model-len 16384
```

## Performance

### Throughput (tokens/sec)

| Model             | RTX 3090 | RTX 4090 | A100 40GB | A100 80GB |
| ----------------- | -------- | -------- | --------- | --------- |
| Qwen2.5-0.5B      | 250      | 320      | 380       | 400       |
| Qwen2.5-3B        | 150      | 200      | 250       | 280       |
| Qwen2.5-7B        | 75       | 100      | 130       | 150       |
| Qwen2.5-7B Q4     | 110      | 140      | 180       | 200       |
| Qwen2.5-14B       | -        | 55       | 70        | 85        |
| Qwen2.5-32B       | -        | -        | 35        | 50        |
| Qwen2.5-72B       | -        | -        | 20 (2x)   | 40 (2x)   |
| Qwen2.5-72B Q4    | -        | -        | -         | 55 (2x)   |
| Qwen2.5-Coder-32B | -        | -        | 32        | 48        |

### Time to First Token (TTFT)

| Model | RTX 4090 | A100 40GB  | A100 80GB  |
| ----- | -------- | ---------- | ---------- |
| 7B    | 60ms     | 40ms       | 35ms       |
| 14B   | 120ms    | 80ms       | 60ms       |
| 32B   | -        | 200ms      | 140ms      |
| 72B   | -        | 400ms (2x) | 280ms (2x) |

### Context Length vs VRAM (7B)

| Context | FP16 | Q8   | Q4   |
| ------- | ---- | ---- | ---- |
| 8K      | 16GB | 10GB | 6GB  |
| 32K     | 24GB | 16GB | 10GB |
| 64K     | 40GB | 26GB | 16GB |
| 128K    | 72GB | 48GB | 28GB |

## Benchmarks

| Model             | MMLU  | HumanEval | GSM8K | MATH  | LiveCodeBench |
| ----------------- | ----- | --------- | ----- | ----- | ------------- |
| Qwen2.5-7B        | 74.2% | 75.6%     | 85.4% | 55.2% | 42.1%         |
| Qwen2.5-14B       | 79.7% | 81.1%     | 89.5% | 65.8% | 51.3%         |
| Qwen2.5-32B       | 83.3% | 84.2%     | 91.2% | 72.1% | 60.7%         |
| Qwen2.5-72B       | 86.1% | 86.2%     | 93.2% | 79.5% | 67.4%         |
| Qwen2.5-Coder-7B  | 72.8% | 88.4%     | 86.1% | 58.4% | 64.2%         |
| Qwen2.5-Coder-32B | 83.1% | **92.7%** | 92.3% | 76.8% | **78.5%**     |

## Docker Compose

```yaml
version: '3.8'

services:
  qwen:
    image: vllm/vllm-openai:latest
    ports:
      - "8000:8000"
    volumes:
      - ~/.cache/huggingface:/root/.cache/huggingface
    command: >
      --model Qwen/Qwen2.5-7B-Instruct
      --host 0.0.0.0
      --port 8000
      --gpu-memory-utilization 0.9
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
```

## Cost Estimate

Typical CLORE.AI marketplace rates:

| GPU           | Hourly Rate | Best For              |
| ------------- | ----------- | --------------------- |
| RTX 3090 24GB | \~$0.06     | 7B models             |
| RTX 4090 24GB | \~$0.10     | 7B-14B models         |
| A100 40GB     | \~$0.17     | 14B-32B models        |
| A100 80GB     | \~$0.25     | 32B models, Coder-32B |
| 2x A100 80GB  | \~$0.50     | 72B models            |
| 4x A100 80GB  | \~$1.00     | 72B max context       |

*Prices vary by provider. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads
* Pay with **CLORE** tokens
* Start with smaller models (7B) for testing

## Troubleshooting

### Out of Memory

```bash
# Reduce context
--max-model-len 8192

# Enable memory optimization
--gpu-memory-utilization 0.85

# Use quantized model
ollama pull qwen2.5:7b-instruct-q4_K_M
```

### Slow Generation

```bash
# Enable flash attention
pip install flash-attn

# Use vLLM for better throughput
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-7B-Instruct \
    --enable-prefix-caching
```

### Chinese Characters Display

```python
# Ensure UTF-8 encoding
import sys
sys.stdout.reconfigure(encoding='utf-8')
```

### Model Not Found

```bash
# Check model name
huggingface-cli search Qwen/Qwen2.5

# Common names:
# Qwen/Qwen2.5-7B-Instruct
# Qwen/Qwen2.5-72B-Instruct       ← New
# Qwen/Qwen2.5-Coder-7B-Instruct
# Qwen/Qwen2.5-Coder-32B-Instruct ← New
# Qwen/Qwen2.5-Math-7B-Instruct
```

## Qwen2.5 vs Others

| Feature      | Qwen2.5-7B | Qwen2.5-72B | Llama 3.1 70B | GPT-4o      |
| ------------ | ---------- | ----------- | ------------- | ----------- |
| Context      | 128K       | 128K        | 128K          | 128K        |
| Multilingual | Excellent  | Excellent   | Good          | Excellent   |
| Code         | Excellent  | Excellent   | Good          | Excellent   |
| Math         | Excellent  | Excellent   | Good          | Excellent   |
| Chinese      | Excellent  | Excellent   | Poor          | Good        |
| License      | Apache 2.0 | Apache 2.0  | Llama 3.1     | Proprietary |
| Cost         | Free       | Free        | Free          | Paid API    |

**Use Qwen2.5 when:**

* Chinese language support needed
* Math/code tasks are priority
* Long context is required
* Want Apache 2.0 license
* Need best open-source code model (Coder-32B)

## Next Steps

* [vLLM](/guides/language-models/vllm) - Production deployment
* [Ollama](/guides/language-models/ollama) - Easy local setup
* [DeepSeek-V3](/guides/language-models/deepseek-v3) - Larger reasoning model
* [DeepSeek-R1](/guides/language-models/deepseek-r1) - Open-source reasoning model
* [Fine-tune LLM](/guides/training/finetune-llm) - Custom training


# CodeLlama

Generate, complete, and explain code with CodeLlama on Clore.ai

{% hint style="info" %}
**Newer alternatives!** For coding tasks, consider [**Qwen2.5-Coder**](/guides/language-models/qwen25) (32B, state-of-the-art code gen) or [**DeepSeek-R1**](/guides/language-models/deepseek-r1) (reasoning + coding). CodeLlama is still useful for lightweight deployments.
{% endhint %}

Generate, complete, and explain code with Meta's CodeLlama.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## Model Variants

| Model         | Size | VRAM  | Best For        |
| ------------- | ---- | ----- | --------------- |
| CodeLlama-7B  | 7B   | 8GB   | Fast completion |
| CodeLlama-13B | 13B  | 16GB  | Balanced        |
| CodeLlama-34B | 34B  | 40GB  | Best quality    |
| CodeLlama-70B | 70B  | 80GB+ | Maximum quality |

### Variants

* **Base**: Code completion
* **Instruct**: Follow instructions
* **Python**: Python-specialized

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
pip install vllm && \
python -m vllm.entrypoints.openai.api_server \
    --model codellama/CodeLlama-7b-Instruct-hf \
    --port 8000
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Installation

### Using Ollama

```bash

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run CodeLlama
ollama run codellama

# Run Python variant
ollama run codellama:python
```

### Using Transformers

```bash
pip install transformers accelerate
```

## Code Completion

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "codellama/CodeLlama-7b-hf"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# Code completion
code = """
def fibonacci(n):
    '''Calculate the nth fibonacci number'''
"""

inputs = tokenizer(code, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=200,
    temperature=0.2,
    do_sample=True
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## Instruct Model

For following coding instructions:

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "codellama/CodeLlama-7b-Instruct-hf"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

prompt = """[INST] Write a Python function that:
1. Takes a list of numbers
2. Removes duplicates
3. Sorts in descending order
4. Returns top 5 elements
[/INST]"""

inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=500,
    temperature=0.2
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## Fill-in-the-Middle (FIM)

```python

# CodeLlama supports FIM for code insertion
prefix = """def calculate_area(shape, dimensions):
    if shape == "circle":
        radius = dimensions[0]
"""

suffix = """
    elif shape == "rectangle":
        length, width = dimensions
        return length * width
    return None
"""

# Use special tokens for FIM
prompt = f"<PRE> {prefix} <SUF>{suffix} <MID>"

inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=100)
```

## Python-Specialized Model

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "codellama/CodeLlama-7b-Python-hf"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# Python-specific completion
code = """
import pandas as pd
import numpy as np

def analyze_sales_data(df):
    '''Analyze sales data and return key metrics'''
"""

inputs = tokenizer(code, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=300)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## vLLM Server

```bash
python -m vllm.entrypoints.openai.api_server \
    --model codellama/CodeLlama-13b-Instruct-hf \
    --dtype float16 \
    --max-model-len 8192
```

### API Usage

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

response = client.chat.completions.create(
    model="codellama/CodeLlama-13b-Instruct-hf",
    messages=[
        {"role": "user", "content": "Write a FastAPI endpoint for user authentication"}
    ],
    temperature=0.2,
    max_tokens=1000
)

print(response.choices[0].message.content)
```

## Code Explanation

```python
code_to_explain = """
def quicksort(arr):
    if len(arr) <= 1:
        return arr
    pivot = arr[len(arr) // 2]
    left = [x for x in arr if x < pivot]
    middle = [x for x in arr if x == pivot]
    right = [x for x in arr if x > pivot]
    return quicksort(left) + middle + quicksort(right)
"""

prompt = f"[INST] Explain this code step by step:\n\n{code_to_explain}\n[/INST]"
```

## Bug Fixing

```python
buggy_code = """
def reverse_string(s):
    result = ""
    for i in range(len(s)):
        result += s[i]
    return result
"""

prompt = f"""[INST] Find and fix the bug in this code. The function should reverse a string:

{buggy_code}
[/INST]"""
```

## Code Translation

```python
python_code = """
def factorial(n):
    if n <= 1:
        return 1
    return n * factorial(n - 1)
"""

prompt = f"""[INST] Convert this Python code to JavaScript:

{python_code}
[/INST]"""
```

## Gradio Interface

```python
import gradio as gr
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "codellama/CodeLlama-7b-Instruct-hf"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

def generate_code(instruction, temperature, max_tokens):
    prompt = f"[INST] {instruction} [/INST]"
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

    outputs = model.generate(
        **inputs,
        max_new_tokens=max_tokens,
        temperature=temperature,
        do_sample=True
    )

    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
    return response.split("[/INST]")[-1].strip()

demo = gr.Interface(
    fn=generate_code,
    inputs=[
        gr.Textbox(label="Instruction", lines=5, placeholder="Write a Python function that..."),
        gr.Slider(0.1, 1.0, value=0.2, label="Temperature"),
        gr.Slider(100, 2000, value=500, step=100, label="Max Tokens")
    ],
    outputs=gr.Code(language="python", label="Generated Code"),
    title="CodeLlama Code Generator"
)

demo.launch(server_name="0.0.0.0", server_port=7860)
```

## Batch Processing

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "codellama/CodeLlama-7b-Instruct-hf"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

tasks = [
    "Write a function to validate email addresses",
    "Create a class for managing a shopping cart",
    "Write a function to parse JSON from a URL",
    "Create a decorator for timing function execution",
    "Write a function to generate random passwords"
]

for task in tasks:
    prompt = f"[INST] {task} [/INST]"
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

    outputs = model.generate(
        **inputs,
        max_new_tokens=500,
        temperature=0.2
    )

    result = tokenizer.decode(outputs[0], skip_special_tokens=True)
    print(f"\n=== {task} ===")
    print(result.split("[/INST]")[-1].strip())
```

## Use with Continue (VSCode)

Configure Continue extension:

```json
{
  "models": [
    {
      "title": "CodeLlama",
      "provider": "ollama",
      "model": "codellama:7b-instruct"
    }
  ],
  "tabAutocompleteModel": {
    "title": "CodeLlama",
    "provider": "ollama",
    "model": "codellama:7b-code"
  }
}
```

## Performance

| Model         | GPU      | Tokens/sec |
| ------------- | -------- | ---------- |
| CodeLlama-7B  | RTX 3090 | \~90       |
| CodeLlama-7B  | RTX 4090 | \~130      |
| CodeLlama-13B | RTX 4090 | \~70       |
| CodeLlama-34B | A100     | \~50       |

## Troubleshooting

### Poor Code Quality

* Lower temperature (0.1-0.3)
* Use Instruct variant
* Larger model if possible

### Incomplete Output

* Increase max\_new\_tokens
* Check context length

### Slow Generation

* Use vLLM
* Quantize model
* Use smaller variant

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* Open Interpreter - Execute code
* vLLM Inference - Production serving
* Mistral/Mixtral - Alternative models


# Gemma 2

Run Google's Gemma 2 models efficiently on Clore.ai GPUs

{% hint style="info" %}
**Newer version available!** Google released [**Gemma 3**](/guides/language-models/gemma3) in March 2025 — the 27B model beats Llama 3.1 405B and adds native multimodal support. Consider upgrading.
{% endhint %}

Run Google's Gemma 2 models for efficient inference.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## What is Gemma 2?

Gemma 2 from Google offers:

* Models from 2B to 27B parameters
* Excellent performance per size
* Strong instruction following
* Efficient architecture

## Model Variants

| Model       | Parameters | VRAM | Context |
| ----------- | ---------- | ---- | ------- |
| Gemma-2-2B  | 2B         | 3GB  | 8K      |
| Gemma-2-9B  | 9B         | 12GB | 8K      |
| Gemma-2-27B | 27B        | 32GB | 8K      |

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
pip install vllm && \
vllm serve google/gemma-2-9b-it --port 8000
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Using Ollama

```bash

# Run Gemma 2
ollama run gemma2

# Specific sizes
ollama run gemma2:2b
ollama run gemma2:9b
ollama run gemma2:27b
```

## Installation

```bash
pip install transformers accelerate torch
```

## Basic Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "google/gemma-2-9b-it"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Explain how neural networks learn."}
]

inputs = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt",
    add_generation_prompt=True
).to("cuda")

outputs = model.generate(
    inputs,
    max_new_tokens=512,
    temperature=0.7,
    do_sample=True
)

response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
```

## Gemma 2 2B (Lightweight)

For edge/mobile deployment:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "google/gemma-2-2b-it"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# Fast inference for simple tasks
messages = [{"role": "user", "content": "Summarize in one sentence: AI is transforming industries."}]
```

## Gemma 2 27B (Best Quality)

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

model_id = "google/gemma-2-27b-it"

# Use 4-bit to fit in 24GB VRAM
quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16
)

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=quantization_config,
    device_map="auto"
)
```

## vLLM Server

```bash
vllm serve google/gemma-2-9b-it \
    --port 8000 \
    --dtype bfloat16 \
    --max-model-len 8192
```

### OpenAI-Compatible API

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

response = client.chat.completions.create(
    model="google/gemma-2-9b-it",
    messages=[
        {"role": "user", "content": "Write a haiku about programming"}
    ],
    temperature=0.8
)

print(response.choices[0].message.content)
```

## Streaming

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")

stream = client.chat.completions.create(
    model="google/gemma-2-9b-it",
    messages=[{"role": "user", "content": "Tell me a short story"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

## Gradio Interface

```python
import gradio as gr
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "google/gemma-2-9b-it"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

def chat(message, history, temperature):
    messages = []
    for h in history:
        messages.append({"role": "user", "content": h[0]})
        messages.append({"role": "assistant", "content": h[1]})
    messages.append({"role": "user", "content": message})

    inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to("cuda")
    outputs = model.generate(inputs, max_new_tokens=512, temperature=temperature, do_sample=True)

    return tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)

demo = gr.ChatInterface(
    fn=chat,
    additional_inputs=[gr.Slider(0.1, 1.5, value=0.7, label="Temperature")],
    title="Gemma 2 Chat"
)

demo.launch(server_name="0.0.0.0", server_port=7860)
```

## Batch Processing

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "google/gemma-2-9b-it"
tokenizer = AutoTokenizer.from_pretrained(model_id, padding_side="left")
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

prompts = [
    "Explain gravity in one sentence.",
    "What is photosynthesis?",
    "Define machine learning.",
    "What is the speed of light?"
]

messages_batch = [[{"role": "user", "content": p}] for p in prompts]

inputs = tokenizer.apply_chat_template(
    messages_batch,
    return_tensors="pt",
    padding=True,
    add_generation_prompt=True
).to("cuda")

outputs = model.generate(inputs, max_new_tokens=128, pad_token_id=tokenizer.pad_token_id)

for i, output in enumerate(outputs):
    response = tokenizer.decode(output, skip_special_tokens=True)
    print(f"Q: {prompts[i]}")
    print(f"A: {response.split('<start_of_turn>model')[-1].strip()}\n")
```

## Performance

| Model               | GPU      | Tokens/sec |
| ------------------- | -------- | ---------- |
| Gemma-2-2B          | RTX 3060 | \~100      |
| Gemma-2-9B          | RTX 3090 | \~60       |
| Gemma-2-9B          | RTX 4090 | \~85       |
| Gemma-2-27B         | A100     | \~45       |
| Gemma-2-27B (4-bit) | RTX 4090 | \~30       |

## Comparison

| Model        | MMLU  | Quality | Speed |
| ------------ | ----- | ------- | ----- |
| Gemma-2-9B   | 71.3% | Great   | Fast  |
| Llama-3.1-8B | 69.4% | Good    | Fast  |
| Mistral-7B   | 62.5% | Good    | Fast  |

## Troubleshooting

{% hint style="danger" %}
**CUDA out of memory**
{% endhint %}

for 27B - Use 4-bit quantization with BitsAndBytesConfig - Reduce \`max\_new\_tokens\` - Clear GPU cache: \`torch.cuda.empty\_cache()\`

### Slow generation

* Use vLLM for production deployment
* Enable Flash Attention
* Try 9B model for faster inference

### Output quality issues

* Use instruction-tuned version (`-it` suffix)
* Adjust temperature (0.7-0.9 recommended)
* Add system prompt for context

### Tokenizer warnings

* Update transformers to latest version
* Use `padding_side="left"` for batch inference

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Next Steps

* Llama 3.2 - Meta's model
* Qwen2.5 - Alibaba's model
* vLLM Inference - Production serving


# Phi-4

Run Microsoft's Phi-4 small language model on Clore.ai GPUs

Run Microsoft's Phi-4 - a small but powerful language model.

{% hint style="success" %}
All examples can be run on GPU servers rented through [CLORE.AI Marketplace](https://clore.ai/marketplace).
{% endhint %}

## Renting on CLORE.AI

1. Visit [CLORE.AI Marketplace](https://clore.ai/marketplace)
2. Filter by GPU type, VRAM, and price
3. Choose **On-Demand** (fixed rate) or **Spot** (bid price)
4. Configure your order:
   * Select Docker image
   * Set ports (TCP for SSH, HTTP for web UIs)
   * Add environment variables if needed
   * Enter startup command
5. Select payment: **CLORE**, **BTC**, or **USDT/USDC**
6. Create order and wait for deployment

### Access Your Server

* Find connection details in **My Orders**
* Web interfaces: Use the HTTP port URL
* SSH: `ssh -p <port> root@<proxy-address>`

## What is Phi-4?

Phi-4 from Microsoft offers:

* 14B parameters with excellent performance
* Beats larger models on benchmarks
* Strong reasoning and math
* Efficient inference

## Model Variants

| Model          | Parameters        | VRAM | Specialty          |
| -------------- | ----------------- | ---- | ------------------ |
| Phi-4          | 14B               | 16GB | General            |
| Phi-3.5-mini   | 3.8B              | 4GB  | Lightweight        |
| Phi-3.5-MoE    | 42B (6.6B active) | 16GB | Mixture of Experts |
| Phi-3.5-vision | 4.2B              | 6GB  | Vision             |

## Quick Deploy

**Docker Image:**

```
pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
```

**Ports:**

```
22/tcp
8000/http
```

**Command:**

```bash
pip install transformers accelerate torch && \
python phi4_server.py
```

## Accessing Your Service

After deployment, find your `http_pub` URL in **My Orders**:

1. Go to **My Orders** page
2. Click on your order
3. Find the `http_pub` URL (e.g., `abc123.clorecloud.net`)

Use `https://YOUR_HTTP_PUB_URL` instead of `localhost` in examples below.

## Using Ollama

```bash

# Run Phi-4
ollama run phi4

# Phi-3.5 mini (faster)
ollama run phi3.5

# Phi-3.5 vision
ollama run phi3.5-vision
```

## Installation

```bash
pip install transformers accelerate torch
```

## Basic Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "microsoft/Phi-4"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "system", "content": "You are a helpful AI assistant."},
    {"role": "user", "content": "Explain the difference between TCP and UDP."}
]

inputs = tokenizer.apply_chat_template(
    messages,
    return_tensors="pt",
    add_generation_prompt=True
).to("cuda")

outputs = model.generate(
    inputs,
    max_new_tokens=512,
    temperature=0.7,
    do_sample=True
)

response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
```

## Phi-3.5-Vision

For image understanding:

```python
from transformers import AutoModelForCausalLM, AutoProcessor
from PIL import Image
import torch

model_id = "microsoft/Phi-3.5-vision-instruct"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

image = Image.open("diagram.png")

messages = [
    {"role": "user", "content": "<|image_1|>\nDescribe this diagram in detail."}
]

prompt = processor.tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

inputs = processor(prompt, [image], return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.7
)

response = processor.decode(outputs[0], skip_special_tokens=True)
print(response)
```

## Math and Reasoning

```python
messages = [
    {"role": "user", "content": """
Solve step by step:
A farmer has chickens and rabbits.
Total heads: 35
Total legs: 94
How many of each animal?
"""}
]

# Phi-4 excels at step-by-step reasoning
```

## Code Generation

```python
messages = [
    {"role": "user", "content": """
Write a Python implementation of binary search tree with:
- Insert
- Search
- Delete
- In-order traversal
Include type hints and docstrings.
"""}
]
```

## Quantized Inference

```python
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16
)

model = AutoModelForCausalLM.from_pretrained(
    "microsoft/Phi-4",
    quantization_config=quantization_config,
    device_map="auto",
    trust_remote_code=True
)
```

## Gradio Interface

```python
import gradio as gr
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "microsoft/Phi-4"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)

def chat(message, history, system_prompt, temperature):
    messages = [{"role": "system", "content": system_prompt}]
    for h in history:
        messages.append({"role": "user", "content": h[0]})
        messages.append({"role": "assistant", "content": h[1]})
    messages.append({"role": "user", "content": message})

    inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to("cuda")
    outputs = model.generate(inputs, max_new_tokens=512, temperature=temperature, do_sample=True)

    return tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)

demo = gr.ChatInterface(
    fn=chat,
    additional_inputs=[
        gr.Textbox(value="You are a helpful assistant.", label="System"),
        gr.Slider(0.1, 1.5, value=0.7, label="Temperature")
    ],
    title="Phi-4 Chat"
)

demo.launch(server_name="0.0.0.0", server_port=7860)
```

## Performance

| Model         | GPU      | Tokens/sec |
| ------------- | -------- | ---------- |
| Phi-3.5-mini  | RTX 3060 | \~100      |
| Phi-3.5-mini  | RTX 4090 | \~150      |
| Phi-4         | RTX 4090 | \~60       |
| Phi-4         | A100     | \~90       |
| Phi-4 (4-bit) | RTX 3090 | \~40       |

## Benchmarks

| Model         | MMLU  | HumanEval | GSM8K |
| ------------- | ----- | --------- | ----- |
| Phi-4         | 84.8% | 82.6%     | 94.6% |
| GPT-4-Turbo   | 86.4% | 85.4%     | 94.2% |
| Llama-3.1-70B | 83.6% | 80.5%     | 92.1% |

*Phi-4 matches or beats much larger models*

## Troubleshooting

### "trust\_remote\_code" error

* Add `trust_remote_code=True` to `from_pretrained()`
* This is required for Phi models

### Repetitive outputs

* Lower temperature (0.3-0.6)
* Add repetition\_penalty=1.1
* Use proper chat template

### Memory issues

* Phi-4 is efficient but still needs \~8GB for 14B
* Use 4-bit quantization if needed
* Reduce context length

### Wrong output format

* Use `apply_chat_template()` for proper formatting
* Check you're using instruct version, not base

## Cost Estimate

Typical CLORE.AI marketplace rates (as of 2024):

| GPU       | Hourly Rate | Daily Rate | 4-Hour Session |
| --------- | ----------- | ---------- | -------------- |
| RTX 3060  | \~$0.03     | \~$0.70    | \~$0.12        |
| RTX 3090  | \~$0.06     | \~$1.50    | \~$0.25        |
| RTX 4090  | \~$0.10     | \~$2.30    | \~$0.40        |
| A100 40GB | \~$0.17     | \~$4.00    | \~$0.70        |
| A100 80GB | \~$0.25     | \~$6.00    | \~$1.00        |

*Prices vary by provider and demand. Check* [*CLORE.AI Marketplace*](https://clore.ai/marketplace) *for current rates.*

**Save money:**

* Use **Spot** market for flexible workloads (often 30-50% cheaper)
* Pay with **CLORE** tokens
* Compare prices across different providers

## Use Cases

* Math tutoring
* Code assistance
* Document analysis (vision)
* Efficient edge deployment
* Cost-effective inference

## Next Steps

* Qwen2.5 - Alternative model
* Gemma 2 - Google's model
* Llama 3.2 - Meta's model


# Llama 4 (Scout & Maverick)

Run Meta Llama 4 Scout & Maverick MoE models on Clore.ai GPUs

Meta's Llama 4, released April 2025, marks a fundamental shift to **Mixture of Experts (MoE)** architecture. Instead of activating all parameters for every token, Llama 4 routes each token to specialized "expert" sub-networks — delivering frontier performance at a fraction of the compute cost. Two open-weight models are available: **Scout** (ideal for single-GPU) and **Maverick** (multi-GPU powerhouse).

## Key Features

* **MoE Architecture**: Only 17B parameters active per token (out of 109B/400B total)
* **Massive Context Windows**: Scout supports 10M tokens, Maverick supports 1M tokens
* **Natively Multimodal**: Understands both text and images out of the box
* **Two Models**: Scout (16 experts, single-GPU friendly) and Maverick (128 experts, multi-GPU)
* **Competitive Performance**: Scout matches Gemma 3 27B; Maverick competes with GPT-4o class models
* **Open Weights**: Llama Community License (free for most commercial uses)

## Model Variants

| Model        | Total Params | Active Params | Experts | Context | Min VRAM (Q4) | Min VRAM (FP16) |
| ------------ | ------------ | ------------- | ------- | ------- | ------------- | --------------- |
| **Scout**    | 109B         | 17B           | 16      | 10M     | 12GB          | 80GB            |
| **Maverick** | 400B         | 17B           | 128     | 1M      | 48GB (multi)  | 320GB (multi)   |

## Requirements

| Component | Scout (Q4)  | Scout (FP16) | Maverick (Q4) |
| --------- | ----------- | ------------ | ------------- |
| GPU       | 1× RTX 4090 | 1× H100      | 4× RTX 4090   |
| VRAM      | 24GB        | 80GB         | 4×24GB        |
| RAM       | 32GB        | 64GB         | 128GB         |
| Disk      | 50GB        | 120GB        | 250GB         |
| CUDA      | 11.8+       | 12.0+        | 12.0+         |

**Recommended Clore.ai GPU**: RTX 4090 24GB (\~$0.5–2/day) for Scout — best value

## Quick Start with Ollama

The fastest way to get Llama 4 running:

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run Scout (quantized, ~12GB VRAM)
ollama run llama4-scout

# For longer context (uses more VRAM)
ollama run llama4-scout --ctx-size 32768
```

### Ollama as API Server

```bash
# Start server in background
ollama serve &

# Pull model
ollama pull llama4-scout

# Query via OpenAI-compatible API
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama4-scout",
    "messages": [{"role": "user", "content": "Explain MoE architecture in 3 sentences"}]
  }'
```

## vLLM Setup (Production)

For production workloads with higher throughput:

```bash
# Install vLLM
pip install vllm

# Serve Scout on single GPU (quantized)
vllm serve meta-llama/Llama-4-Scout-17B-16E-Instruct \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

# Serve Scout on 2 GPUs (longer context)
vllm serve meta-llama/Llama-4-Scout-17B-16E-Instruct \
  --tensor-parallel-size 2 \
  --max-model-len 128000 \
  --gpu-memory-utilization 0.90

# Serve Maverick on 4 GPUs
vllm serve meta-llama/Llama-4-Maverick-17B-128E-Instruct \
  --tensor-parallel-size 4 \
  --max-model-len 65536
```

### Query vLLM Server

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")

response = client.chat.completions.create(
    model="meta-llama/Llama-4-Scout-17B-16E-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Write a Python function to calculate Fibonacci numbers"}
    ],
    temperature=0.7,
    max_tokens=1024
)
print(response.choices[0].message.content)
```

## HuggingFace Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "meta-llama/Llama-4-Scout-17B-16E-Instruct"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    load_in_4bit=True  # 4-bit quantization for 24GB GPUs
)

messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "Write a REST API with FastAPI that manages a todo list"}
]

input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(input_ids, max_new_tokens=2048, temperature=0.7, do_sample=True)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
```

## Docker Quick Start

```bash
# Using vLLM Docker image
docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-4-Scout-17B-16E-Instruct \
  --max-model-len 32768
```

## Why MoE Matters on Clore.ai

Traditional dense models (like Llama 3.3 70B) need massive VRAM because all 70B parameters are active. Llama 4 Scout has 109B total but only activates 17B per token — meaning:

* **Same quality as 70B+ dense models** at a fraction of the VRAM cost
* **Fits on a single RTX 4090** in quantized mode
* **10M token context** — process entire codebases, long documents, books
* **Cheaper to rent** — $0.5–2/day instead of $6–12/day for 70B models

## Tips for Clore.ai Users

* **Start with Scout Q4**: Best bang for buck on RTX 4090 — $0.5–2/day, covers 95% of use cases
* **Use `--max-model-len` wisely**: Don't set context higher than you need — it reserves VRAM. Start at 8192, increase as needed
* **Tensor Parallel for Maverick**: Rent 4× RTX 4090 machines for Maverick; use `--tensor-parallel-size 4`
* **HuggingFace Login Required**: `huggingface-cli login` — you need to accept the Llama license on HF first
* **Ollama for Quick Tests, vLLM for Production**: Ollama is faster to set up; vLLM gives higher throughput for API serving
* **Monitor GPU Memory**: `watch nvidia-smi` — MoE models can spike VRAM on long sequences

## Troubleshooting

| Issue                    | Solution                                                                  |
| ------------------------ | ------------------------------------------------------------------------- |
| `OutOfMemoryError`       | Reduce `--max-model-len`, use Q4 quantization, or upgrade GPU             |
| Model download fails     | Run `huggingface-cli login` and accept Llama 4 license at hf.co           |
| Slow generation          | Ensure GPU is being used (`nvidia-smi`); check `--gpu-memory-utilization` |
| vLLM crashes on start    | Reduce context length; ensure CUDA 11.8+ installed                        |
| Ollama shows wrong model | Run `ollama list` to verify; `ollama rm` + `ollama pull` to re-download   |

## Further Reading

* [Meta Llama 4 Blog Post](https://llama.meta.com/)
* [HuggingFace Model Card](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct)
* [vLLM Documentation](https://docs.vllm.ai/)
* [Ollama Model Library](https://ollama.com/library/llama4-scout)


# Gemma 3

Run Google Gemma 3 multimodal models on Clore.ai — beats Llama-405B at 15x smaller

Gemma 3, released March 2025 by Google DeepMind, is built on the same technology as Gemini 2.0. Its standout achievement: **the 27B model beats Llama 3.1 405B** on LMArena benchmarks — a model 15 times its size. It's natively multimodal (text + images + video), supports 128K context, and runs on a single RTX 4090 with quantization.

## Key Features

* **Punches way above its weight**: 27B beats 405B-class models on major benchmarks
* **Natively multimodal**: Text, image, and video understanding built-in
* **128K context window**: Process long documents, codebases, conversations
* **Four sizes**: 1B, 4B, 12B, 27B — something for every GPU budget
* **QAT versions**: Quantization-Aware Training variants let 27B run on consumer GPUs
* **Wide framework support**: Ollama, vLLM, Transformers, Keras, JAX, PyTorch

## Model Variants

| Model           | Parameters | VRAM (Q4) | VRAM (FP16) | Best For                    |
| --------------- | ---------- | --------- | ----------- | --------------------------- |
| Gemma 3 1B      | 1B         | 1.5GB     | 3GB         | Edge, mobile, testing       |
| Gemma 3 4B      | 4B         | 4GB       | 9GB         | Budget GPUs, fast tasks     |
| Gemma 3 12B     | 12B        | 10GB      | 25GB        | Balanced quality/speed      |
| Gemma 3 27B     | 27B        | 18GB      | 54GB        | Best quality, production    |
| Gemma 3 27B QAT | 27B        | 14GB      | —           | Optimized for consumer GPUs |

## Requirements

| Component | Gemma 3 4B | Gemma 3 27B (Q4) | Gemma 3 27B (FP16) |
| --------- | ---------- | ---------------- | ------------------ |
| GPU       | RTX 3060   | RTX 4090         | 2× RTX 4090 / A100 |
| VRAM      | 6GB        | 24GB             | 48GB+              |
| RAM       | 16GB       | 32GB             | 64GB               |
| Disk      | 10GB       | 25GB             | 55GB               |
| CUDA      | 11.8+      | 11.8+            | 12.0+              |

**Recommended Clore.ai GPU**: RTX 4090 24GB (\~$0.5–2/day) for 27B quantized — the sweet spot

## Quick Start with Ollama

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run different sizes
ollama run gemma3:1b     # Tiny — 1.5GB VRAM
ollama run gemma3:4b     # Small — 4GB VRAM
ollama run gemma3:12b    # Medium — 10GB VRAM
ollama run gemma3:27b    # Large — 18-20GB VRAM (quantized)

# QAT version (optimized quantization)
ollama run gemma3:27b-qat
```

### Ollama API Server

```bash
ollama serve &

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma3:27b",
    "messages": [{"role": "user", "content": "Compare REST vs GraphQL for a new API"}]
  }'
```

### Vision with Ollama

```bash
# Analyze an image
ollama run gemma3:27b "Describe this image in detail" --images ./photo.jpg
```

## vLLM Setup (Production)

```bash
pip install vllm

# Serve 27B model
vllm serve google/gemma-3-27b-it \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90

# Serve with longer context on 2 GPUs
vllm serve google/gemma-3-27b-it \
  --tensor-parallel-size 2 \
  --max-model-len 65536

# Serve 4B for budget setups
vllm serve google/gemma-3-4b-it \
  --max-model-len 32768
```

## HuggingFace Transformers

### Text Generation

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "google/gemma-3-27b-it"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    load_in_4bit=True  # Fits on 24GB GPU
)

messages = [
    {"role": "user", "content": "Write a Python class for a binary search tree with insert, search, and delete methods"}
]

input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(input_ids, max_new_tokens=2048, temperature=0.7, do_sample=True)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
```

### Vision (Image Understanding)

```python
import torch
from transformers import AutoProcessor, Gemma3ForConditionalGeneration
from PIL import Image

model_name = "google/gemma-3-27b-it"
processor = AutoProcessor.from_pretrained(model_name)
model = Gemma3ForConditionalGeneration.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

# Load image
image = Image.open("screenshot.png")

messages = [
    {"role": "user", "content": [
        {"type": "image", "image": image},
        {"type": "text", "text": "What does this screenshot show? List all UI elements."}
    ]}
]

inputs = processor.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(output[0], skip_special_tokens=True))
```

## Docker Quick Start

```bash
docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model google/gemma-3-27b-it \
  --max-model-len 8192
```

## Benchmark Highlights

| Benchmark         | Gemma 3 27B    | Llama 3.1 70B | Llama 3.1 405B |
| ----------------- | -------------- | ------------- | -------------- |
| LMArena ELO       | 1354           | 1298          | 1337           |
| MMLU              | 75.6           | 79.3          | 85.2           |
| HumanEval         | 72.0           | 72.6          | 80.5           |
| VRAM (Q4)         | 18GB           | 40GB          | 200GB+         |
| **Cost on Clore** | **$0.5–2/day** | **$3–6/day**  | **$12–24/day** |

The 27B delivers 405B-class conversational quality at 1/10th the VRAM cost.

## Tips for Clore.ai Users

* **27B QAT is the sweet spot**: Quantization-Aware Training means less quality loss than post-training quantization — run it on a single RTX 4090
* **Vision is free**: No extra setup needed — Gemma 3 understands images natively. Great for document parsing, screenshot analysis, chart reading
* **Start with short context**: Use `--max-model-len 8192` initially; increase only when needed to save VRAM
* **4B for budget runs**: If you're on RTX 3060/3070 ($0.15–0.3/day), the 4B model still outperforms last-gen 27B models
* **Google auth not required**: Unlike some models, Gemma 3 downloads without gating (just accept license on HuggingFace)

## Troubleshooting

| Issue                        | Solution                                                                   |
| ---------------------------- | -------------------------------------------------------------------------- |
| `OutOfMemoryError` on 27B    | Use QAT version or reduce `--max-model-len` to 4096                        |
| Vision not working in Ollama | Update Ollama to latest: `curl -fsSL https://ollama.com/install.sh \| sh`  |
| Slow generation speed        | Check you're using bfloat16, not float32. Use `--dtype bfloat16`           |
| Model outputs garbage        | Ensure you're using the `-it` (instruct-tuned) variant, not the base model |
| Download 403 error           | Accept the Gemma license at <https://huggingface.co/google/gemma-3-27b-it> |

## Further Reading

* [Gemma 3 Technical Report](https://ai.google.dev/gemma)
* [HuggingFace Model Card](https://huggingface.co/google/gemma-3-27b-it)
* [Ollama Library](https://ollama.com/library/gemma3)
* [Google AI Studio](https://aistudio.google.com/) — try Gemma 3 online before renting a GPU


# Gemma 4 (26B MoE, 4B active)

Deploy Gemma 4 (26B MoE, 4B active) by Google on Clore.ai — the open-weight model released April 2026 that climbed to

{% hint style="info" %}
**Status (April 2026):** Gemma 4 was released on **April 2, 2026** by Google as the next generation of the Gemma open-weight family. Two variants ship: a **31B dense** model (`google/gemma-4-31b-it`) and a **26B MoE with \~4B active parameters** (`google/gemma-4-26b-it`). Both are published under the standard **Gemma terms of use** at [huggingface.co/google/gemma-4-26b-it](https://huggingface.co/google/gemma-4-26b-it) and [huggingface.co/google/gemma-4-31b-it](https://huggingface.co/google/gemma-4-31b-it).
{% endhint %}

Gemma 4 is Google's first MoE entry in the Gemma line and the first Gemma release that climbed into the top of the LMSYS Arena (vendor reports **#3 overall at release**, edging out several closed models on factuality and instruction-following). The headline number is the MoE variant: **26B total parameters, \~4B active per token**, which gives you near-frontier instruction-following at the inference cost of a small dense model.

For Clore.ai users the practical takeaway is simple — the 26B MoE runs comfortably on a single **RTX 4090 (24GB)** with FP8 or 4-bit quantization (\~10 tok/s) and hits production-grade throughput on a single **H100 80GB** (\~40+ tok/s), putting Gemma-quality instruction-following within reach at roughly $0.5–2/day on the marketplace. The 31B dense variant is the more capable but more expensive sibling, needing 2× RTX 4090 or 1× H100 to serve.

## Key Features

* **MoE architecture (26B variant)** — 26B total parameters, \~4B activated per token; pay 4B-class inference cost for 26B-class quality
* **Dense fallback (31B variant)** — for teams that prefer the predictability and tooling maturity of dense inference
* **128K context window** — long-document Q\&A, RAG over mid-sized codebases, multi-turn agent loops
* **Strong instruction-following** — Gemma 4 is explicitly tuned for tool use, structured output, and faithful constraint following
* **Multilingual** — full multilingual coverage out of Gemma 3 carried forward, plus an expanded non-English benchmark suite
* **Open weights, Gemma terms** — free for most commercial use; review the [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy) before shipping
* **First-class tooling** — supported out of the box in vLLM, SGLang, Ollama, and Hugging Face Transformers

## Choose Your Variant

| Variant                                  | Total Params | Active         | Context | Recommended Quant | Recommended Clore GPU                                                                                                   |
| ---------------------------------------- | ------------ | -------------- | ------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------- |
| **Gemma 4 26B MoE** (`gemma-4-26b-it`)   | 26B          | \~4B per token | 128K    | FP8 or 4-bit GPTQ | 1× [RTX 4090](https://clore.ai/rent-4090.html?utm_source=docs\&utm_medium=guide\&utm_campaign=gemma4) (24GB, quantized) |
| **Gemma 4 31B Dense** (`gemma-4-31b-it`) | 31B          | 31B (all)      | 128K    | FP8 or BF16       | 1× [H100](https://clore.ai/rent-h100.html?utm_source=docs\&utm_medium=guide\&utm_campaign=gemma4) (80GB, BF16)          |

{% hint style="success" %}
**Practical pick:** For 90% of single-GPU deployments, go with **Gemma 4 26B MoE on FP8**. You get the headline Arena quality at \~10–15 tok/s on a 4090 and \~40+ tok/s on an H100, without the latency cost of dense 31B inference.
{% endhint %}

***

## Server Requirements

| Component  | 26B MoE (4-bit, 4090) | 26B MoE (FP8, H100) | 31B Dense (BF16, H100) |
| ---------- | --------------------- | ------------------- | ---------------------- |
| GPU VRAM   | 24GB                  | 80GB                | 80GB                   |
| System RAM | 32GB                  | 64GB                | 64GB                   |
| Disk       | 60GB NVMe             | 80GB NVMe           | 90GB NVMe              |
| Network    | 100 Mbps for HF pull  | 1 Gbps preferred    | 1 Gbps preferred       |
| CUDA       | 12.1+                 | 12.4+               | 12.4+                  |
| Driver     | 550+                  | 555+                | 555+                   |

Plan for an extra \~20% VRAM headroom on top of the static weight footprint to cover KV cache at long contexts. Setting `--gpu-memory-utilization 0.90` in vLLM is a good default.

***

## Quick Deploy on CLORE.AI

The fastest path: rent a single GPU, pull the standard `vllm/vllm-openai` image, and serve the model with an OpenAI-compatible API. Below is the docker-compose layout used by the rest of these guides — adjust the model name and tensor-parallel size based on the variant you picked above.

### Option A — Gemma 4 26B MoE on a single GPU (vLLM, FP8)

```yaml
version: "3.8"
services:
  vllm:
    image: vllm/vllm-openai:latest
    ports:
      - "8000:8000"
    environment:
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
    volumes:
      - hf_cache:/root/.cache/huggingface
    command: >
      --model google/gemma-4-26b-it
      --quantization fp8
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --served-model-name gemma-4-26b
      --trust-remote-code
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    shm_size: "8gb"

volumes:
  hf_cache:
```

```bash
# Bring it up
HF_TOKEN=hf_xxx docker compose up -d

# Tail logs while the weights download
docker compose logs -f vllm
```

{% hint style="info" %}
**License gating:** Gemma models on Hugging Face require accepting Google's terms once per account. Visit the model page in a browser, click "Acknowledge license", then export `HF_TOKEN` so the container can pull the weights.
{% endhint %}

### Option B — Gemma 4 31B Dense on H100 (vLLM, BF16)

```bash
docker run --gpus all -p 8000:8000 \
  -e HUGGING_FACE_HUB_TOKEN=$HF_TOKEN \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model google/gemma-4-31b-it \
  --dtype bfloat16 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --served-model-name gemma-4-31b
```

### Option C — Gemma 4 31B Dense on 2× RTX 4090 (FP8, tensor-parallel)

```bash
docker run --gpus all -p 8000:8000 \
  -e HUGGING_FACE_HUB_TOKEN=$HF_TOKEN \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model google/gemma-4-31b-it \
  --quantization fp8 \
  --tensor-parallel-size 2 \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90 \
  --served-model-name gemma-4-31b
```

### Option D — Quick local testing with Ollama

For laptop-class experimentation, Ollama wraps the GGUF community builds. Expect quants to land a few days after the official release.

```bash
# Once a community GGUF is published
ollama pull gemma4:26b-moe-q4_k_m
ollama run gemma4:26b-moe-q4_k_m

# OpenAI-compatible API on :11434
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma4:26b-moe-q4_k_m",
    "messages": [{"role":"user","content":"Summarize the MoE routing approach in two sentences."}]
  }'
```

See the [Ollama guide](/guides/language-models/ollama) for general setup, model management, and persistence tips.

***

## Usage Examples

The vLLM container exposes an OpenAI-compatible API on `:8000`. Anything that speaks the OpenAI chat-completions schema works directly.

### Curl chat completion

```bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "messages": [
      {"role": "system", "content": "You are a careful technical writer."},
      {"role": "user", "content": "Explain MoE routing in three sentences without using analogies."}
    ],
    "max_tokens": 512,
    "temperature": 0.7
  }'
```

### Python (OpenAI client)

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

resp = client.chat.completions.create(
    model="gemma-4-26b",
    messages=[
        {"role": "system", "content": "You answer in plain text, no markdown."},
        {"role": "user", "content": "Give me a 5-bullet code review checklist for a Go HTTP handler."},
    ],
    temperature=0.7,
    max_tokens=1024,
)
print(resp.choices[0].message.content)
```

### Streaming responses

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

stream = client.chat.completions.create(
    model="gemma-4-26b",
    messages=[{"role": "user", "content": "Write a haiku about distributed inference."}],
    stream=True,
    max_tokens=128,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
print()
```

### Hugging Face Transformers (offline use)

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "google/gemma-4-26b-it"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    load_in_4bit=True,  # Fits the MoE on a single 24GB card
)

messages = [
    {"role": "user", "content": "Refactor this Python function for readability:\n\ndef f(x): return [i for i in x if i%2==0 and i>10]"},
]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(input_ids, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
```

***

## Performance Tips

* **Use FP8 on Hopper.** On H100 the FP8 checkpoint is roughly half the memory of BF16 with no measurable quality loss for instruction-following tasks. Pass `--quantization fp8` to vLLM.
* **Use 4-bit GPTQ on Ada (RTX 4090).** For the MoE variant on a single 4090, a community GPTQ 4-bit build is the practical sweet spot — expect \~10–15 tok/s. Ollama's Q4\_K\_M GGUF builds give similar quality with simpler ops.
* **Tensor parallelism for 31B Dense.** Across 2× RTX 4090, pass `--tensor-parallel-size 2`. Pin the context to what you actually need (`--max-model-len 16384`) — every doubling of context roughly doubles the KV cache footprint.
* **Expert parallelism for the MoE.** On multi-GPU setups for the 26B MoE, vLLM's `--enable-expert-parallel` can give a meaningful throughput bump at higher batch sizes. It's overkill for single-GPU.
* **Chunked prefill for long contexts.** When pushing past 32K, add `--enable-chunked-prefill` to vLLM. This keeps prefill latency manageable and prevents stalls on the decode path.
* **Pre-pull weights.** For ephemeral Clore rentals, mount a persistent volume at `/root/.cache/huggingface` so subsequent runs skip the 50–60GB download.
* **Pick the right serving backend.** vLLM is the safe default. SGLang often wins on Hopper for high-concurrency workloads; see the [vLLM guide](/guides/language-models/vllm) for the broader comparison.

***

## Benchmarks

{% hint style="warning" %}
**Vendor-published numbers — independent verification pending.** The figures below come from Google's April 2, 2026 launch materials. Independent reproductions on private evals are still rolling in. Treat the Arena ranking and factuality scores as directional, not absolute.
{% endhint %}

| Benchmark                       | Gemma 4 26B MoE                               | Gemma 4 31B Dense                        | Reference       |
| ------------------------------- | --------------------------------------------- | ---------------------------------------- | --------------- |
| LMSYS Arena (overall)           | #3 at release                                 | \~#5 at release                          | vendor-reported |
| Instruction-following (IFEval)  | vendor reports strong gains over Gemma 3      | vendor reports strong gains over Gemma 3 | vendor-reported |
| Factuality (SimpleQA / similar) | beats several closed models per Google        | comparable                               | vendor-reported |
| Multilingual (Global-MMLU)      | vendor reports parity with much larger models | best Gemma score to date                 | vendor-reported |

Gemma 4's positioning argument is "more useful per active parameter," not "raw HumanEval king." If you need pure code generation, compare against [GLM-5.1](/guides/language-models/glm-5-1) (frontier coding) or [Qwen3.5](/guides/language-models/qwen35) (best 35B-class dense). If you need long-horizon agentic loops, GLM-5.1 is still the sharper tool.

***

## Troubleshooting

| Issue                                          | Solution                                                                                                                                                                                  |
| ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `OutOfMemoryError` loading the 26B MoE on 24GB | Switch to FP8 (`--quantization fp8`) or 4-bit (`load_in_4bit=True` in Transformers). Drop `--max-model-len` to 16384 to shrink the KV cache.                                              |
| `OutOfMemoryError` loading 31B Dense on H100   | BF16 at 32K context is right at the edge on 80GB. Lower `--max-model-len` to 16384 or move to FP8.                                                                                        |
| Hugging Face download fails with 403           | You have not accepted the Gemma license on the model page. Open the URL in a browser, acknowledge the terms, then re-pull with a token that has `read` scope.                             |
| Very slow first token                          | Cold weight load (\~30–60s on first request) plus prefill on long inputs. Run a dummy warm-up request after the server starts. Add `--enable-chunked-prefill` for long-context workloads. |
| Garbled output / repetition loops              | Check the chat template — `tokenizer.apply_chat_template` is required; do not concatenate `system`+`user` strings manually. Set `temperature=0.7` and `top_p=0.95` for general use.       |
| Tool / JSON output unreliable                  | Use vLLM's `--guided-decoding-backend` or pass a JSON schema via `response_format`. The model follows constraints well but unstructured prompts will still drift.                         |
| `unsupported quantization` error in vLLM       | Update to a vLLM version released after April 2026 (`pip install -U vllm --pre`). The Gemma 4 architecture needs the latest config parsers.                                               |

***

## FAQ

**Gemma 4 vs Llama 4?** Different shapes for different jobs. [Llama 4 Scout](/guides/language-models/llama4) is 109B/17B-active with a headline 10M context — great when you need to dump huge inputs at the model. Gemma 4 26B MoE is much smaller in total params (26B vs 109B), activates fewer params per token (4B vs 17B), and is tuned harder for instruction-following and factuality. For tight VRAM budgets and quality-per-parameter, Gemma 4 wins. For absurd context length, Llama 4 Scout wins.

**How much VRAM for Gemma 4 26B MoE?**

* 4-bit GGUF / GPTQ: fits in **24GB** (single RTX 4090), \~10–15 tok/s.
* FP8: comfortable on **40GB**, fast on **80GB** (H100) at \~40+ tok/s.
* BF16 full: \~55GB of weights plus KV cache — plan for an **80GB** card.

**Can I use Gemma 4 commercially?** Yes, under the standard Gemma terms of use. Review the [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy) before deploying — there are restrictions around specific use cases (deception, generating CSAM, illegal activity), and you must pass downstream license notices to your users. It is not an Apache 2.0 / MIT model — it is open-weight under a usage policy. If you need a fully unrestricted license, [Qwen3.5](/guides/language-models/qwen35) (Apache 2.0) or [GLM-5.1](/guides/language-models/glm-5-1) (MIT) are alternatives.

**Gemma 4 vs DeepSeek-V4?** [DeepSeek-V4](/guides/language-models/deepseek-v4) is a different weight class — \~1T params, multimodal, 1M context. Use DeepSeek-V4 when you need raw capability and have a serious GPU rack. Use Gemma 4 26B MoE when you want strong instruction-following on a **single GPU** and care about \~$1–2/day rentals on Clore. Gemma 4 is the "best model that fits on a 4090" candidate; DeepSeek-V4 is the "I will pay for 8× H200" candidate.

**Does Gemma 4 support vision / multimodal inputs?** Gemma 4's headline release is text-only instruction-tuned (`*-it`). Google has historically followed text releases with PaliGemma vision variants — track [huggingface.co/google](https://huggingface.co/google) for updates. For an image-capable open model today, look at [Kimi K2.5](/guides/language-models/kimi-k2) or [Llama 4 Scout](/guides/language-models/llama4).

***

## Related Guides

* [vLLM](/guides/language-models/vllm) — production serving backend used in this guide
* [Ollama](/guides/language-models/ollama) — quickest path to local testing with GGUF builds
* [Llama 4](/guides/language-models/llama4) — Meta's MoE alternative with 10M context
* [GLM-5.1](/guides/language-models/glm-5-1) — frontier-class coding MoE (744B/40B-active) when Gemma's size class is not enough
* [Qwen3.5](/guides/language-models/qwen35) — Apache-2.0 35B dense, the other strong single-GPU option
* [Gemma 3](/guides/language-models/gemma3) — the predecessor generation, useful baseline for migration

### Links

* [Gemma 4 26B MoE on Hugging Face](https://huggingface.co/google/gemma-4-26b-it)
* [Gemma 4 31B Dense on Hugging Face](https://huggingface.co/google/gemma-4-31b-it)
* [Gemma terms of use](https://ai.google.dev/gemma/terms)
* [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy)
* [vLLM docs](https://docs.vllm.ai)
* [SGLang repo](https://github.com/sgl-project/sglang)


# Mistral Small 3.1

Deploy Mistral Small 3.1 (24B) on Clore.ai — the ideal single-GPU production model

Mistral Small 3.1, released March 2025 by Mistral AI, is a **24-billion parameter dense model** that punches way above its weight. With a 128K context window, native vision capabilities, best-in-class function calling, and an **Apache 2.0 license**, it's arguably the best model you can run on a single RTX 4090. It outperforms GPT-4o Mini and Claude 3.5 Haiku on most benchmarks while fitting comfortably on consumer hardware when quantized.

## Key Features

* **24B dense parameters** — no MoE complexity, straightforward deployment
* **128K context window** — RULER 128K score of 81.2%, beats GPT-4o Mini (65.8%)
* **Native vision** — analyze images, charts, documents, and screenshots
* **Apache 2.0 license** — fully open for commercial and personal use
* **Elite function calling** — native tool use with JSON output, ideal for agentic workflows
* **Multilingual** — 25+ languages including CJK, Arabic, Hindi, and European languages

## Requirements

| Component | Quantized (Q4)   | Full Precision (BF16)  |
| --------- | ---------------- | ---------------------- |
| GPU       | 1× RTX 4090 24GB | 2× RTX 4090 or 1× H100 |
| VRAM      | \~16GB           | \~55GB                 |
| RAM       | 32GB             | 64GB                   |
| Disk      | 20GB             | 50GB                   |
| CUDA      | 11.8+            | 12.0+                  |

**Clore.ai recommendation**: RTX 4090 (\~$0.5–2/day) for quantized inference — best price/performance ratio

## Quick Start with Ollama

The fastest way to get Mistral Small 3.1 running:

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run Mistral Small 3.1 (auto-downloads ~14GB Q4 quantization)
ollama run mistral-small3.1

# Or specify a specific quantization
ollama run mistral-small3.1:24b-instruct-2503-q4_K_M
```

### Ollama as OpenAI-Compatible API

```bash
# Start Ollama server
ollama serve &

# Pull the model
ollama pull mistral-small3.1

# Query via API
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-small3.1",
    "messages": [
      {"role": "system", "content": "You are a helpful coding assistant."},
      {"role": "user", "content": "Write a Python decorator for rate limiting"}
    ],
    "temperature": 0.15
  }'
```

### Ollama with Vision

```bash
# Send an image for analysis
curl http://localhost:11434/api/chat -d '{
  "model": "mistral-small3.1",
  "messages": [{
    "role": "user",
    "content": "What does this image show?",
    "images": ["/path/to/image.jpg"]
  }]
}'
```

## vLLM Setup (Production)

For production workloads with high throughput and concurrent requests:

```bash
# Install vLLM (v0.8.1+ required)
pip install -U vllm

# Verify mistral_common is installed (should be automatic)
python -c "import mistral_common; print(mistral_common.__version__)"
```

### Serve on Single GPU (Text Only)

```bash
vllm serve mistralai/Mistral-Small-3.1-24B-Instruct-2503 \
  --tokenizer-mode mistral \
  --config-format mistral \
  --load-format mistral \
  --tool-call-parser mistral \
  --enable-auto-tool-choice \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90
```

### Serve with Vision (2 GPUs Recommended)

```bash
vllm serve mistralai/Mistral-Small-3.1-24B-Instruct-2503 \
  --tokenizer-mode mistral \
  --config-format mistral \
  --load-format mistral \
  --tool-call-parser mistral \
  --enable-auto-tool-choice \
  --limit-mm-per-prompt 'image=10' \
  --tensor-parallel-size 2 \
  --max-model-len 65536
```

### Query the Server

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="mistralai/Mistral-Small-3.1-24B-Instruct-2503",
    messages=[
        {"role": "system", "content": "You are a helpful assistant. Today is 2026-02-20."},
        {"role": "user", "content": "Write a complete REST API in FastAPI with CRUD operations for a blog"}
    ],
    temperature=0.15,
    max_tokens=4096
)
print(response.choices[0].message.content)
```

## HuggingFace Transformers

For direct Python integration and experimentation:

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "mistralai/Mistral-Small-3.1-24B-Instruct-2503"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    load_in_4bit=True  # 4-bit quantization — fits on 24GB GPU
)

messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "Implement a binary search tree in Python with insert, delete, and search methods"}
]

input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)

output = model.generate(
    input_ids,
    max_new_tokens=2048,
    temperature=0.15,
    do_sample=True
)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
```

## Function Calling Example

Mistral Small 3.1 is one of the best small models for tool use:

```python
import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_stock_price",
            "description": "Get the current stock price for a given ticker symbol",
            "parameters": {
                "type": "object",
                "required": ["ticker"],
                "properties": {
                    "ticker": {"type": "string", "description": "Stock ticker symbol (e.g., AAPL)"}
                }
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "calculate_portfolio_value",
            "description": "Calculate total portfolio value given holdings",
            "parameters": {
                "type": "object",
                "required": ["holdings"],
                "properties": {
                    "holdings": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "properties": {
                                "ticker": {"type": "string"},
                                "shares": {"type": "number"}
                            }
                        }
                    }
                }
            }
        }
    }
]

response = client.chat.completions.create(
    model="mistralai/Mistral-Small-3.1-24B-Instruct-2503",
    messages=[{"role": "user", "content": "What's the current price of AAPL and MSFT?"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.15
)

for tool_call in response.choices[0].message.tool_calls:
    print(f"Call: {tool_call.function.name}({tool_call.function.arguments})")
```

## Docker Quick Start

```bash
# Single GPU deployment
docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model mistralai/Mistral-Small-3.1-24B-Instruct-2503 \
  --tokenizer-mode mistral \
  --config-format mistral \
  --load-format mistral \
  --tool-call-parser mistral \
  --enable-auto-tool-choice \
  --max-model-len 32768

# With vision support (2 GPUs)
docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model mistralai/Mistral-Small-3.1-24B-Instruct-2503 \
  --tokenizer-mode mistral \
  --config-format mistral \
  --load-format mistral \
  --tool-call-parser mistral \
  --enable-auto-tool-choice \
  --limit-mm-per-prompt 'image=10' \
  --tensor-parallel-size 2
```

## Tips for Clore.ai Users

* **RTX 4090 is the sweet spot**: At $0.5–2/day, a single RTX 4090 runs Mistral Small 3.1 quantized with room to spare. Best cost/performance ratio on Clore.ai for a general-purpose LLM.
* **Use low temperature**: Mistral AI recommends `temperature=0.15` for most tasks. Higher temps cause inconsistent output with this model.
* **RTX 3090 works too**: At $0.3–1/day, RTX 3090 (24GB) runs Q4 quantized with Ollama just fine. Slightly slower than 4090 but half the price.
* **Ollama for quick setups, vLLM for production**: Ollama gives you a working model in 60 seconds. For concurrent API requests and higher throughput, switch to vLLM.
* **Function calling makes it special**: Many 24B models can chat — few can reliably call tools. Mistral Small 3.1's function calling is on par with GPT-4o Mini. Build agents, API backends, and automation pipelines with confidence.

## Troubleshooting

| Issue                          | Solution                                                                                       |
| ------------------------------ | ---------------------------------------------------------------------------------------------- |
| `OutOfMemoryError` on RTX 4090 | Use quantized model via Ollama or `load_in_4bit=True` in Transformers. Full BF16 needs \~55GB. |
| Ollama model not found         | Use `ollama run mistral-small3.1` (official library name).                                     |
| vLLM tokenizer errors          | Always pass `--tokenizer-mode mistral --config-format mistral --load-format mistral`.          |
| Poor output quality            | Set `temperature=0.15`. Add a system prompt. Mistral Small is sensitive to temperature.        |
| Vision not working on 1 GPU    | Vision features need more VRAM. Use `--tensor-parallel-size 2` or reduce `--max-model-len`.    |
| Function calls return empty    | Add `--tool-call-parser mistral --enable-auto-tool-choice` to vLLM serve.                      |

## Further Reading

* [Mistral Small 3.1 on HuggingFace](https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503)
* [Mistral AI Blog Post](https://mistral.ai/news/mistral-small-3-1/)
* [Ollama Model Page](https://ollama.com/library/mistral-small3.1)
* [vLLM Documentation](https://docs.vllm.ai/)
* [Mistral Common Library](https://github.com/mistralai/mistral-common)
* [Mistral AI Platform](https://console.mistral.ai/)


# Qwen3.5

Run Alibaba Qwen3.5 on Clore.ai — the freshest frontier model (Feb 2026)

Qwen3.5, released February 16, 2026, is Alibaba's latest flagship model and one of the hottest open-source releases of 2026. The **397B MoE flagship** beat Claude 4.5 Opus on the HMMT math benchmark, while the smaller **35B dense model** fits on a single RTX 4090. All models come with agentic capabilities (tool use, function calling, autonomous task execution) and multimodal understanding out of the box.

## Key Features

* **Three sizes**: 9B (dense), 35B (dense), 397B (MoE) — something for every GPU
* **Beat Claude 4.5 Opus** on HMMT math benchmark
* **Natively multimodal**: Text + image understanding
* **Agentic capabilities**: Tool use, function calling, autonomous workflows
* **128K context window**: Handle large documents and codebases
* **Apache 2.0 license**: Full commercial use, no restrictions

## Model Variants

| Model        | Params | Type  | VRAM (Q4) | VRAM (FP16) | Strength        |
| ------------ | ------ | ----- | --------- | ----------- | --------------- |
| Qwen3.5-9B   | 9B     | Dense | 6GB       | 18GB        | Fast, efficient |
| Qwen3.5-35B  | 35B    | Dense | 22GB      | 70GB        | Best single-GPU |
| Qwen3.5-397B | 397B   | MoE   | \~100GB   | 400GB+      | Frontier-class  |

## Requirements

| Component | 9B (Q4)       | 35B (Q4)      | 397B (multi-GPU) |
| --------- | ------------- | ------------- | ---------------- |
| GPU       | RTX 3080 10GB | RTX 4090 24GB | 4× H100 80GB     |
| VRAM      | 8GB           | 22GB          | 320GB+           |
| RAM       | 16GB          | 32GB          | 128GB            |
| Disk      | 15GB          | 30GB          | 250GB            |

**Recommended Clore.ai GPU**: RTX 4090 24GB (\~$0.5–2/day) for 35B — best quality per dollar

## Quick Start with Ollama

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# 9B — runs on anything (8GB VRAM)
ollama run qwen3.5:9b

# 35B quantized — needs RTX 4090 (24GB)
ollama run qwen3.5:35b

# As API server
ollama serve &
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5:35b",
    "messages": [{"role": "user", "content": "Solve this: if f(x) = x^3 - 3x + 1, find all real roots"}]
  }'
```

## vLLM Setup (Production)

```bash
pip install vllm

# 35B on single GPU
vllm serve Qwen/Qwen3.5-35B-Instruct \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

# 9B with long context
vllm serve Qwen/Qwen3.5-9B-Instruct \
  --max-model-len 65536

# 397B on multi-GPU cluster
vllm serve Qwen/Qwen3.5-397B-A45B-Instruct \
  --tensor-parallel-size 8 \
  --max-model-len 32768
```

## HuggingFace Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3.5-35B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    load_in_4bit=True  # Fits 35B on 24GB
)

messages = [
    {"role": "system", "content": "You are a helpful math tutor."},
    {"role": "user", "content": "Prove that the square root of 2 is irrational."}
]

input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(input_ids, max_new_tokens=2048, temperature=0.7, do_sample=True)
print(tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True))
```

## Agentic / Tool Use Example

```python
import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

tools = [{
    "type": "function",
    "function": {
        "name": "get_gpu_price",
        "description": "Get current rental price for a GPU model on Clore.ai",
        "parameters": {
            "type": "object",
            "properties": {
                "gpu_model": {"type": "string", "description": "GPU model name, e.g. RTX 4090"}
            },
            "required": ["gpu_model"]
        }
    }
}]

response = client.chat.completions.create(
    model="qwen3.5:35b",
    messages=[{"role": "user", "content": "What's the cheapest GPU I can rent for running a 7B model?"}],
    tools=tools,
    tool_choice="auto"
)

# Qwen3.5 will call get_gpu_price with appropriate parameters
print(response.choices[0].message)
```

## Why Qwen3.5 on Clore.ai?

The 35B model is arguably the **best model you can run on a single RTX 4090**:

* Beats Llama 4 Scout on math and reasoning
* Beats Gemma 3 27B on agentic tasks
* Tool use / function calling works out of the box
* Apache 2.0 = no license headaches

At $0.5–2/day for an RTX 4090, you get frontier-class AI for the cost of a coffee.

## Tips for Clore.ai Users

* **35B is the sweet spot**: Fits on RTX 4090 Q4, outperforms most 70B models
* **9B for budget**: Even RTX 3060 ($0.15/day) runs the 9B model well
* **Use Ollama for quick start**: One command to serve; OpenAI-compatible API included
* **Agentic workflows**: Qwen3.5 excels at tool use — combine with function calling for automation
* **Fresh model = less cached**: First download takes time (\~20GB for 35B). Pre-pull before your workload starts

## Troubleshooting

| Issue                  | Solution                                                        |
| ---------------------- | --------------------------------------------------------------- |
| 35B OOM on 24GB        | Use `load_in_4bit=True` or reduce `--max-model-len`             |
| Ollama model not found | Update Ollama: `curl -fsSL https://ollama.com/install.sh \| sh` |
| Slow on first request  | Model loading takes 30-60s; subsequent requests are fast        |
| Tool calls not working | Ensure you pass `tools` parameter; use instruct variant only    |

## Further Reading

* [Qwen Blog](https://qwenlm.github.io/)
* [HuggingFace Models](https://huggingface.co/Qwen)
* [Ollama Library](https://ollama.com/library/qwen3.5)


# Qwen3.5-Omni (Multimodal)

Alibaba's **Qwen3.5-Omni** is a unified end-to-end multimodal model released on March 30, 2026 under the Apache 2.0 license. It can understand and reason across text, audio, images, and video simultaneously — and generate both text and speech as output. Running it on a rented Clore.ai GPU gives you a production-grade multimodal assistant at a fraction of cloud API costs.

***

## What Is Qwen3.5-Omni?

Qwen3.5-Omni is an **end-to-end multimodal model** built on a sparse Mixture-of-Experts architecture. The HuggingFace release (`Qwen3.5-Omni-7B`) uses Alibaba's naming convention where "7B" refers to the active-parameter configuration per inference step; the full checkpoint includes all expert weights. That sparsity is what makes it deployable on a single RTX 4090 (24 GB) using INT4 quantization — a model that would otherwise require far more VRAM at full precision.

### Key Capabilities

| Modality | Input                            | Output               |
| -------- | -------------------------------- | -------------------- |
| Text     | ✅                                | ✅                    |
| Audio    | ✅ (transcription, understanding) | ✅ (speech synthesis) |
| Image    | ✅ (understanding, OCR, analysis) | —                    |
| Video    | ✅ (scene understanding, QA)      | —                    |

Unlike previous multimodal models that bolt together separate encoders, Qwen3.5-Omni processes all modalities in a single unified forward pass. It can simultaneously transcribe spoken audio, analyze a video frame, and respond with both text and a synthesized voice — in one inference call.

### Architecture Highlights

* **Gated Delta Networks (GDN)** for efficient sequence modeling with subquadratic complexity on long audio/video streams
* **Sparse Mixture-of-Experts** — 30B total params, \~3B active per token; comparable quality to 7–14B dense models but faster at scale
* **Unified tokenizer** covering text, audio frames, image patches, and video frame sequences
* **Built-in TTS decoder** — generates speech waveforms natively rather than through a separate pipeline

Released March 30, 2026 · License: **Apache 2.0** · [HuggingFace](https://huggingface.co/Qwen/Qwen3.5-Omni-7B)

***

## Qwen3.5-Omni vs. Related Models

| Model               | Params              | Modalities In             | Speech Out | License     | VRAM (INT4) |
| ------------------- | ------------------- | ------------------------- | ---------- | ----------- | ----------- |
| **Qwen3.5-Omni**    | 30B MoE (3B active) | Text, Audio, Image, Video | ✅          | Apache 2.0  | \~15 GB     |
| Qwen3.5 (text-only) | 32B                 | Text only                 | ❌          | Apache 2.0  | \~18 GB     |
| Qwen2.5-VL          | 72B                 | Text, Image, Video        | ❌          | Apache 2.0  | \~40 GB     |
| Gemini 2.0 Flash    | —                   | Text, Audio, Image, Video | ✅          | Proprietary | API only    |

Compared to **Qwen3.5 (text-only)**, the Omni variant adds audio/video understanding and speech generation while actually requiring *less* VRAM at INT4 thanks to the MoE architecture. Compared to **Qwen2.5-VL**, it adds audio I/O but requires far less hardware.

***

## Hardware Requirements

| Precision      | VRAM Required | Recommended GPU          |
| -------------- | ------------- | ------------------------ |
| BF16 (full)    | 64–80 GB      | A100 80GB, H100          |
| BF16 multi-GPU | 2× 40 GB      | 2× A40 / 2× A6000        |
| INT4 / GGUF    | \~15 GB       | RTX 4090 (24 GB) ✅       |
| INT8           | \~30 GB       | A6000 48GB, RTX 6000 Ada |

For most self-hosted use cases, **INT4 on an RTX 4090** is the sweet spot: full multimodal capability at $0.50–0.80/day on Clore.ai.

***

## Quick Start on Clore.ai

### Step 1: Rent a GPU

Go to [clore.ai/marketplace](https://clore.ai/marketplace) and rent:

* **INT4 / Single-GPU**: RTX 4090 (24 GB) — from **\~$0.50/day**
* **BF16 / Full Precision**: A100 80GB or H100 — from **\~$2.50/day**

Use the **vllm/vllm-openai** Docker image or the standard CUDA image.

### Step 2: Deploy with vLLM (Recommended)

vLLM v0.17.0+ is required for Qwen3.5-Omni support.

```bash
# Pull and run the vLLM OpenAI-compatible server
docker run --gpus all --rm -it \
  -p 8000:8000 \
  -v /workspace/models:/root/.cache/huggingface \
  vllm/vllm-openai:v0.17.0 \
  --model Qwen/Qwen3.5-Omni-7B \
  --quantization awq_marlin \
  --max-model-len 32768 \
  --trust-remote-code
```

> **Note:** The `awq_marlin` flag requires a pre-quantized AWQ model. Download `Qwen/Qwen3.5-Omni-7B-AWQ` instead of the base model, or omit `--quantization` for BF16 on A100/H100.

Once the server is running, it exposes an OpenAI-compatible API at `http://localhost:8000/v1`.

### Step 3: Deploy with Ollama (Simpler Setup)

For quick experimentation without Docker complexity:

```bash
# Install Ollama
curl -fsSL https://ollama.ai/install.sh | sh

# Pull Qwen3.5-Omni (quantized)
# Note: check https://ollama.com/library for availability — tag may vary
ollama pull qwen3.5-omni

# Start the server
ollama serve
```

Ollama handles quantization automatically and provides a simple `/api/generate` endpoint.

***

## Example API Calls

### Multimodal Input: Image + Text

```python
import openai
import base64

client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="none")

# Load an image
with open("screenshot.png", "rb") as f:
    image_b64 = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="Qwen/Qwen3.5-Omni-7B",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{image_b64}"}
                },
                {
                    "type": "text",
                    "text": "Describe what you see in this image and identify any text."
                }
            ]
        }
    ],
    max_tokens=512
)
print(response.choices[0].message.content)
```

### Audio Transcription + Understanding

```python
import openai
import base64

client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="none")

with open("meeting_recording.wav", "rb") as f:
    audio_b64 = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="Qwen/Qwen3.5-Omni-7B",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "audio_url",
                    "audio_url": {"url": f"data:audio/wav;base64,{audio_b64}"}
                },
                {
                    "type": "text",
                    "text": "Transcribe this audio and summarize the key points."
                }
            ]
        }
    ]
)
print(response.choices[0].message.content)
```

### Video Understanding

```python
# Video frames can be passed as a sequence of image URLs
# or as a video_url when using the Qwen3.5-Omni native API
response = client.chat.completions.create(
    model="Qwen/Qwen3.5-Omni-7B",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "video_url",
                    "video_url": {"url": "https://example.com/product-demo.mp4"}
                },
                {
                    "type": "text",
                    "text": "What is happening in this video? Describe each scene."
                }
            ]
        }
    ]
)
```

***

## Multi-GPU Setup for BF16

If you rent a multi-GPU machine on Clore.ai (e.g., 2× A40 or 2× A6000), use tensor parallelism:

```bash
docker run --gpus all --rm -it \
  -p 8000:8000 \
  -v /workspace/models:/root/.cache/huggingface \
  vllm/vllm-openai:v0.17.0 \
  --model Qwen/Qwen3.5-Omni-7B \
  --tensor-parallel-size 2 \
  --dtype bfloat16 \
  --max-model-len 65536 \
  --trust-remote-code
```

This splits the model across both GPUs for maximum throughput and quality.

***

## Use Cases

### 1. Customer Service Automation

Qwen3.5-Omni can listen to customer voice calls, transcribe them in real-time, understand the issue, and generate both a text summary and a spoken response. All in one model, no stitching together separate ASR + LLM + TTS pipelines.

### 2. Video Content Understanding

Upload product demo videos, lecture recordings, or surveillance footage and get detailed text descriptions, timestamped summaries, or Q\&A. The model handles up to 32K tokens of context, covering multi-minute videos.

### 3. Real-Time Voice Agents

Build conversational voice assistants that understand context across audio turns. Qwen3.5-Omni maintains conversational memory and can interleave its text reasoning with speech generation — ideal for phone-based customer support bots.

### 4. Document + Screenshot Analysis

OCR, layout understanding, chart interpretation — pass in screenshots of dashboards, PDFs, or handwritten notes and get structured text output or detailed analysis.

### 5. Multilingual Audio Processing

The model supports 29 languages for both text and speech, making it suitable for international customer support, multilingual transcription pipelines, and cross-lingual video analysis.

***

## Cost Estimate on Clore.ai

| GPU          | Precision            | VRAM    | Price/Day | Best For                             |
| ------------ | -------------------- | ------- | --------- | ------------------------------------ |
| RTX 4090     | INT4                 | 24 GB   | \~$0.50   | Dev, testing, small-scale production |
| RTX 6000 Ada | INT8                 | 48 GB   | \~$1.20   | Better quality, moderate throughput  |
| A100 80GB    | BF16                 | 80 GB   | \~$2.50   | Full quality, high throughput        |
| 2× A40       | BF16 tensor parallel | 2×48 GB | \~$2.00   | Full quality, cost-efficient         |

Running Qwen3.5-Omni at INT4 on an RTX 4090 costs less per day than a single OpenAI API call for a complex multimodal task at scale.

***

## Tips & Troubleshooting

**"CUDA out of memory" on RTX 4090**

* Add `--gpu-memory-utilization 0.90` to the vLLM command
* Reduce `--max-model-len` to 16384 if processing short inputs

**Audio input not working**

* Ensure vLLM version is exactly `v0.17.0` or newer — earlier versions lack Omni audio support
* WAV files must be 16kHz mono for best results; use `ffmpeg -ar 16000 -ac 1` to convert

**Slow first inference**

* vLLM compiles CUDA kernels on first run; warmup takes 2–5 minutes. Subsequent calls are fast.

**Ollama not recognizing video input**

* Ollama currently supports image+text and audio only; for video understanding use the vLLM deployment.

***

## Summary

Qwen3.5-Omni brings true end-to-end multimodal AI — text, audio, image, and video in, text and speech out — to a single open-source model that runs on consumer hardware. At INT4, it fits in a 24 GB RTX 4090 and costs under a dollar a day on Clore.ai. With Apache 2.0 licensing and OpenAI-compatible API via vLLM, it drops directly into existing pipelines.

**→** [**Rent an RTX 4090 on Clore.ai**](https://clore.ai/marketplace) and deploy Qwen3.5-Omni today.


# GLM-5

Deploy GLM-5 (744B MoE) by Zhipu AI on Clore.ai — API access and self-hosting with vLLM

GLM-5, released February 2026 by Zhipu AI (Z.AI), is a **744-billion parameter Mixture-of-Experts** language model that activates only 40B parameters per token. It achieves best-in-class open-source performance on reasoning, coding, and agentic tasks — scoring 77.8% on SWE-bench Verified and rivaling frontier models like Claude Opus 4.5 and GPT-5.2. The model is available under the **MIT license** on HuggingFace.

## Key Features

* **744B total / 40B active** — 256-expert MoE with highly efficient routing
* **Frontier coding performance** — 77.8% SWE-bench Verified, 73.3% SWE-bench Multilingual
* **Deep reasoning** — 92.7% on AIME 2026, 96.9% on HMMT Nov 2025, built-in thinking mode
* **Agentic capabilities** — native tool calling, function execution, and long-horizon task planning
* **200K+ context window** — handles massive codebases and long documents
* **MIT license** — fully open weights, commercial use permitted

## Requirements

Self-hosting GLM-5 is a serious undertaking — the FP8 checkpoint requires **\~860GB VRAM**.

| Component | Minimum (FP8) | Recommended   |
| --------- | ------------- | ------------- |
| GPU       | 8× H100 80GB  | 8× H200 141GB |
| VRAM      | 640GB         | 1,128GB       |
| RAM       | 256GB         | 512GB         |
| Disk      | 1.5TB NVMe    | 2TB NVMe      |
| CUDA      | 12.0+         | 12.4+         |

**Clore.ai recommendation**: For most users, **access GLM-5 via API** (Z.AI, OpenRouter). Self-hosting only makes sense if you can rent 8× H100/H200 (\~$24–48/day on Clore.ai).

## API Access (Recommended for Most Users)

The most practical way to use GLM-5 from a Clore.ai machine or anywhere:

### Via Z.AI Platform

```python
from openai import OpenAI

client = OpenAI(
    api_key="your-zai-api-key",
    base_url="https://api.z.ai/v1"
)

response = client.chat.completions.create(
    model="glm-5",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Write a Python async web scraper using aiohttp and BeautifulSoup"}
    ],
    temperature=1.0,
    max_tokens=4096
)
print(response.choices[0].message.content)
```

### Via OpenRouter

```python
from openai import OpenAI

client = OpenAI(
    api_key="your-openrouter-key",
    base_url="https://openrouter.ai/api/v1"
)

response = client.chat.completions.create(
    model="zai-org/glm-5",
    messages=[
        {"role": "user", "content": "Explain the MoE architecture used in GLM-5"}
    ],
    max_tokens=2048
)
print(response.choices[0].message.content)
```

## vLLM Setup (Self-Hosting)

For those with access to high-end multi-GPU machines on Clore.ai:

```bash
# Install vLLM (nightly required for GLM-5 support)
pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightly

# Install latest transformers (required)
pip install git+https://github.com/huggingface/transformers.git
```

### Serve FP8 on 8× H200 GPUs

```bash
vllm serve zai-org/GLM-5-FP8 \
  --tensor-parallel-size 8 \
  --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 1 \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --enable-auto-tool-choice \
  --served-model-name glm-5-fp8 \
  --gpu-memory-utilization 0.85
```

### Query the Server

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

# With thinking mode (default)
response = client.chat.completions.create(
    model="glm-5-fp8",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Solve: find all primes p where p^2 + 2 is also prime"}
    ],
    temperature=1.0,
    max_tokens=4096
)
print(response.choices[0].message.content)

# Without thinking mode (faster, shorter responses)
response = client.chat.completions.create(
    model="glm-5-fp8",
    messages=[
        {"role": "user", "content": "Write a quicksort in Rust"}
    ],
    temperature=1.0,
    max_tokens=4096,
    extra_body={
        "chat_template_kwargs": {"enable_thinking": False}
    }
)
print(response.choices[0].message.content)
```

## SGLang Alternative

SGLang also supports GLM-5 and may offer better performance on some hardware:

```bash
# Using Docker (Hopper GPUs)
docker pull lmsysorg/sglang:glm5-hopper

# Launch server
python3 -m sglang.launch_server \
  --model-path zai-org/GLM-5-FP8 \
  --tp-size 8 \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --mem-fraction-static 0.85 \
  --served-model-name glm-5-fp8
```

## Docker Quick Start

```bash
# vLLM Docker image with GLM-5 support
docker run --gpus all -p 8000:8000 \
  --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:glm5 zai-org/GLM-5-FP8 \
  --tensor-parallel-size 8 \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --enable-auto-tool-choice \
  --served-model-name glm5 \
  --trust-remote-code
```

## Tool Calling Example

GLM-5 has native tool-calling support — ideal for building agentic applications:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "required": ["city"],
            "properties": {
                "city": {"type": "string", "description": "City name"}
            }
        }
    }
}]

response = client.chat.completions.create(
    model="glm-5-fp8",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto"
)
print(response.choices[0].message.tool_calls)
```

## Tips for Clore.ai Users

* **API first, self-host second**: GLM-5 requires 8× H200 (\~$24–48/day on Clore.ai). For occasional use, the Z.AI API or OpenRouter is far more cost-effective. Self-host only if you need sustained throughput or data privacy.
* **Consider GLM-4.7 instead**: If 8× H200 is too much, the predecessor GLM-4.7 (355B, 32B active) runs on 4× H200 or 4× H100 (\~$12–24/day) and still delivers excellent performance.
* **Use FP8 weights**: Always use `zai-org/GLM-5-FP8` — same quality as BF16 but nearly half the memory footprint. The BF16 version requires 16× GPUs.
* **Monitor VRAM usage**: `watch nvidia-smi` — long context queries can spike memory. Set `--gpu-memory-utilization 0.85` to leave headroom.
* **Thinking mode tradeoff**: Thinking mode produces better results for complex tasks but uses more tokens and time. Disable it for simple queries with `enable_thinking: false`.

## Troubleshooting

| Issue                         | Solution                                                                       |
| ----------------------------- | ------------------------------------------------------------------------------ |
| `OutOfMemoryError` on startup | Ensure you have 8× H200 (141GB each). FP8 needs \~860GB total VRAM.            |
| Slow downloads (\~800GB)      | Use `huggingface-cli download zai-org/GLM-5-FP8` with `--local-dir` to resume. |
| vLLM version mismatch         | GLM-5 requires vLLM nightly. Install via `pip install -U vllm --pre`.          |
| Tool calls not working        | Add `--tool-call-parser glm47 --enable-auto-tool-choice` to serve command.     |
| DeepGEMM errors               | Install DeepGEMM for FP8: use the `install_deepgemm.sh` script from vLLM repo. |
| Thinking mode output empty    | Set `temperature=1.0` — thinking mode requires non-zero temperature.           |

## Further Reading

* [GLM-5 on HuggingFace](https://huggingface.co/zai-org/GLM-5)
* [GLM-5 FP8 Checkpoint](https://huggingface.co/zai-org/GLM-5-FP8)
* [Z.AI Platform](https://chat.z.ai)
* [Z.AI API Docs](https://docs.z.ai/guides/llm/glm-5)
* [vLLM GLM-5 Recipe](https://docs.vllm.ai/projects/recipes/en/latest/GLM/GLM5.html)
* [GLM-5 Technical Blog](https://z.ai/blog/glm-5)
* [Slime RL Infrastructure](https://github.com/THUDM/slime)


# GLM-4.7-Flash

Deploy GLM-4.7-Flash (30B MoE) by Zhipu AI on Clore.ai — efficient language model with 59.2% SWE-bench performance

> GLM-4.7-Flash is a **30-billion parameter Mixture-of-Experts** language model by Zhipu AI that activates only 3B parameters per token. It delivers exceptional performance on coding and reasoning tasks, achieving 59.2% on SWE-bench while requiring only 10-12GB VRAM for FP16 inference. Released under the **MIT license**, it's an ideal choice for developers seeking frontier model quality at affordable single-GPU costs.

## At a Glance

* **Model Size**: 30B total / 3B active parameters (MoE)
* **License**: MIT (fully commercial)
* **Context**: 128K tokens
* **Performance**: 59.2% SWE-bench, 75.4% HumanEval
* **VRAM**: \~10-12GB FP16, \~6GB INT8
* **Speed**: \~45-60 tok/s on RTX 4090

## Why GLM-4.7-Flash?

**Efficient Performance**: GLM-4.7-Flash punches above its weight class. Despite using only 3B active parameters, it outperforms many 70B+ dense models on coding benchmarks. The MoE architecture provides 30B model quality at 7B model inference cost.

**Single-GPU Friendly**: Unlike massive models requiring multi-GPU setups, GLM-4.7-Flash runs comfortably on a single RTX 4090 or A100 40GB. This makes it perfect for development, fine-tuning, and cost-effective production deployments.

**Coding Specialist**: With 59.2% SWE-bench performance, GLM-4.7-Flash excels at software engineering tasks — code generation, debugging, refactoring, and technical documentation. It understands 20+ programming languages with deep context awareness.

**MIT Licensed**: No usage restrictions. Deploy commercially, fine-tune, or modify without licensing concerns. The complete weights and training recipes are freely available.

## GPU Recommendations

| GPU          | VRAM | Performance | Daily Cost\* |
| ------------ | ---- | ----------- | ------------ |
| **RTX 4090** | 24GB | \~50 tok/s  | \~$2.10      |
| **RTX 3090** | 24GB | \~35 tok/s  | \~$1.10      |
| A100 40GB    | 40GB | \~80 tok/s  | \~$3.50      |
| A100 80GB    | 80GB | \~90 tok/s  | \~$4.00      |
| H100         | 80GB | \~120 tok/s | \~$6.00      |

**Best Value**: RTX 4090 offers the sweet spot of performance and cost for GLM-4.7-Flash.

\*Estimated Clore.ai marketplace prices

## Deploy with vLLM

### Install vLLM

```bash
pip install vllm>=0.6.0
# or latest
pip install git+https://github.com/vllm-project/vllm.git
```

### Single GPU Setup

```bash
vllm serve THUDM/glm-4-flash \
  --model THUDM/glm-4-flash \
  --tensor-parallel-size 1 \
  --dtype float16 \
  --max-model-len 32768 \
  --served-model-name glm-4.7-flash \
  --trust-remote-code
```

### Query the Server

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1", 
    api_key="EMPTY"
)

response = client.chat.completions.create(
    model="glm-4.7-flash",
    messages=[
        {"role": "system", "content": "You are an expert Python developer."},
        {"role": "user", "content": "Write a FastAPI app with async SQLAlchemy and JWT auth"}
    ],
    max_tokens=2048,
    temperature=0.7
)

print(response.choices[0].message.content)
```

## Deploy with SGLang

SGLang often provides better throughput for MoE models:

```bash
pip install "sglang[all]>=0.3.0"

# Launch server
python -m sglang.launch_server \
  --model-path THUDM/glm-4-flash \
  --port 30000 \
  --host 0.0.0.0 \
  --dtype float16 \
  --tp-size 1 \
  --context-length 32768
```

## Deploy with Ollama

Simple setup for local development:

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Pull model (will download ~18GB)
ollama pull glm4:7b-chat

# Run interactively
ollama run glm4:7b-chat

# API mode
ollama serve
```

Then query via REST API:

```python
import requests

response = requests.post('http://localhost:11434/api/generate',
    json={
        'model': 'glm4:7b-chat',
        'prompt': 'Explain the MoE architecture in GLM-4.7-Flash',
        'stream': False
    }
)

print(response.json()['response'])
```

## Docker Template

```dockerfile
FROM nvidia/cuda:12.1-devel-ubuntu22.04

# Install Python 3.10
RUN apt-get update && apt-get install -y python3.10 python3-pip curl

# Install vLLM
RUN pip install vllm>=0.6.0 transformers

# Pre-download model (optional)
# RUN python3 -c "from transformers import AutoModel; AutoModel.from_pretrained('THUDM/glm-4-flash', trust_remote_code=True)"

EXPOSE 8000

CMD ["vllm", "serve", "THUDM/glm-4-flash", \
     "--host", "0.0.0.0", \
     "--port", "8000", \
     "--tensor-parallel-size", "1", \
     "--dtype", "float16", \
     "--trust-remote-code"]
```

Build and run:

```bash
docker build -t glm-4.7-flash .
docker run --gpus all -p 8000:8000 glm-4.7-flash
```

## Code Generation Example

GLM-4.7-Flash excels at complex code generation:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="glm-4.7-flash",
    messages=[
        {"role": "user", 
         "content": """Create a Python class for a rate limiter with:
- Token bucket algorithm
- Async/await support  
- Redis backend
- Decorator for function rate limiting
- Proper error handling"""}
    ],
    max_tokens=2048,
    temperature=0.3
)

print(response.choices[0].message.content)
```

## Tips for Clore.ai Users

* **Memory Optimization**: Use `--dtype float16` to reduce VRAM usage. For 16GB GPUs, add `--max-model-len 16384` to limit context.
* **Batch Processing**: Increase `--max-num-seqs` for higher throughput when serving multiple requests.
* **Quantization**: For RTX 3060/4060 (12GB), use AWQ or GPTQ quantized versions for \~6GB VRAM usage.
* **Preemption**: GLM-4.7-Flash handles interruptions gracefully — good for preemptible Clore.ai instances.
* **Context Length**: Default 128K context may be overkill. Set `--max-model-len 32768` for most applications.

## Troubleshooting

| Issue              | Solution                                                    |
| ------------------ | ----------------------------------------------------------- |
| `OutOfMemoryError` | Reduce `--max-model-len` or use `--dtype float16`           |
| Slow model loading | Pre-cache with `huggingface-cli download THUDM/glm-4-flash` |
| Import errors      | Update transformers: `pip install transformers>=4.40.0`     |
| Poor performance   | Enable Flash Attention: `pip install flash-attn`            |
| Connection refused | Check firewall: `ufw allow 8000`                            |

## Alternative Models

If GLM-4.7-Flash doesn't fit your needs:

* **Qwen2.5-Coder-7B**: Better pure coding, smaller footprint
* **CodeQwen1.5-7B**: Chinese + English coding specialist
* **GLM-4-9B**: Larger sibling with better reasoning
* **DeepSeek-V3**: 671B MoE for ultimate performance (multi-GPU)

## Resources

* [GLM-4-Flash on Hugging Face](https://huggingface.co/THUDM/glm-4-flash)
* [GLM-4 Technical Report](https://arxiv.org/abs/2406.12793)
* [vLLM Documentation](https://docs.vllm.ai/)
* [SGLang GitHub](https://github.com/sgl-project/sglang)
* [Zhipu AI Platform](https://open.bigmodel.cn/)


# Kimi K2.5

Deploy Kimi K2.5 (1T MoE multimodal) by Moonshot AI on Clore.ai GPUs

Kimi K2.5, released January 27, 2026 by Moonshot AI, is a **1-trillion parameter Mixture-of-Experts multimodal model** with 32B active parameters per token. Built through continual pretraining on \~15 trillion mixed visual and text tokens atop the Kimi-K2-Base, it natively understands text, images, and video. K2.5 introduces **Agent Swarm** technology — coordinating up to 100 specialized AI agents simultaneously — and achieves frontier-level performance on coding (76.8% SWE-bench Verified), vision, and agentic tasks. Available under an **open-weight license** on HuggingFace.

## Key Features

* **1T total / 32B active** — 384-expert MoE architecture with MLA attention and SwiGLU
* **Native multimodal** — pre-trained on vision–language tokens; understands images, video, and text
* **Agent Swarm** — decomposes complex tasks into parallel sub-tasks via dynamically spawned agents
* **256K context window** — process entire codebases, long documents, and video transcripts
* **Hybrid reasoning** — supports both instant mode (fast) and thinking mode (deep reasoning)
* **Strong coding** — 76.8% SWE-bench Verified, 73.0% SWE-bench Multilingual

## Requirements

Kimi K2.5 is a massive model — the FP8 checkpoint is \~630GB. Self-hosting requires serious hardware.

| Component | Quantized (GGUF Q2)     | FP8 Full      |
| --------- | ----------------------- | ------------- |
| GPU       | 1× RTX 4090 + 256GB RAM | 8× H200 141GB |
| VRAM      | 24GB + CPU offload      | 1,128GB       |
| RAM       | 256GB+                  | 256GB         |
| Disk      | 400GB SSD               | 700GB NVMe    |
| CUDA      | 12.0+                   | 12.0+         |

**Clore.ai recommendation**: For full-precision serving, rent 8× H200 (\~$24–48/day). For quantized local inference, a single H100 80GB or even RTX 4090 + heavy CPU offloading works at reduced speed.

## Quick Start with llama.cpp (Quantized)

The most accessible way to run K2.5 locally — using Unsloth's GGUF quantizations:

```bash
# Clone and build llama.cpp
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release -j

# Download quantized model (Q2_K_XL — 375GB, good quality/size balance)
huggingface-cli download unsloth/Kimi-K2.5-GGUF \
  Kimi-K2.5-UD-Q2_K_XL-00001-of-00005.gguf \
  Kimi-K2.5-UD-Q2_K_XL-00002-of-00005.gguf \
  Kimi-K2.5-UD-Q2_K_XL-00003-of-00005.gguf \
  Kimi-K2.5-UD-Q2_K_XL-00004-of-00005.gguf \
  Kimi-K2.5-UD-Q2_K_XL-00005-of-00005.gguf \
  --local-dir ./models

# Run inference (adjust --n-gpu-layers for your VRAM)
./build/bin/llama-server \
  -m ./models/Kimi-K2.5-UD-Q2_K_XL-00001-of-00005.gguf \
  --n-gpu-layers 10 \
  --threads 32 \
  --ctx-size 16384 \
  --host 0.0.0.0 --port 8080
```

> **Note**: Vision is not yet supported in GGUF/llama.cpp for K2.5. For multimodal features, use vLLM.

## vLLM Setup (Production — Full Model)

For production serving with full multimodal support:

```bash
# Install vLLM nightly (K2.5 requires latest)
pip install -U vllm --pre \
  --extra-index-url https://wheels.vllm.ai/nightly/cu129 \
  --extra-index-url https://download.pytorch.org/whl/cu129 \
  --index-strategy unsafe-best-match
```

### Serve on 8× H200 GPUs

```bash
vllm serve moonshotai/Kimi-K2.5 \
  -tp 8 \
  --mm-encoder-tp-mode data \
  --tool-call-parser kimi_k2 \
  --reasoning-parser kimi_k2 \
  --trust-remote-code \
  --gpu-memory-utilization 0.90
```

### Query with Text

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="moonshotai/Kimi-K2.5",
    messages=[
        {"role": "system", "content": "You are Kimi, an AI assistant created by Moonshot AI."},
        {"role": "user", "content": "Write a FastAPI service with WebSocket support for real-time chat"}
    ],
    temperature=0.6,
    max_tokens=4096
)
print(response.choices[0].message.content)
```

### Query with Image (Multimodal)

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY", timeout=3600)

response = client.chat.completions.create(
    model="moonshotai/Kimi-K2.5",
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {"url": "https://example.com/diagram.png"}
            },
            {
                "type": "text",
                "text": "Describe this diagram in detail and extract all text."
            }
        ]
    }],
    max_tokens=2048
)
print(response.choices[0].message.content)
```

## API Access (No GPU Required)

If self-hosting is overkill, use Moonshot's official API:

```python
from openai import OpenAI

# Moonshot Platform — OpenAI-compatible API
client = OpenAI(
    api_key="your-moonshot-api-key",
    base_url="https://api.moonshot.ai/v1"
)

response = client.chat.completions.create(
    model="kimi-k2.5",
    messages=[
        {"role": "user", "content": "Explain the Agent Swarm architecture in Kimi K2.5"}
    ],
    temperature=0.6,
    max_tokens=2048
)
print(response.choices[0].message.content)
```

## Tool Calling

K2.5 excels at agentic tool use:

```python
import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

tools = [{
    "type": "function",
    "function": {
        "name": "search_code",
        "description": "Search a codebase for relevant files and functions",
        "parameters": {
            "type": "object",
            "required": ["query"],
            "properties": {
                "query": {"type": "string", "description": "Search query"}
            }
        }
    }
}]

response = client.chat.completions.create(
    model="moonshotai/Kimi-K2.5",
    messages=[{"role": "user", "content": "Find all authentication-related code in the project"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.6
)

for tool_call in response.choices[0].message.tool_calls:
    print(f"Function: {tool_call.function.name}")
    print(f"Args: {json.loads(tool_call.function.arguments)}")
```

## Docker Quick Start

```bash
# Using vLLM Docker with 8 GPUs
docker run --gpus all -p 8000:8000 \
  --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model moonshotai/Kimi-K2.5 \
  --tensor-parallel-size 8 \
  --mm-encoder-tp-mode data \
  --tool-call-parser kimi_k2 \
  --reasoning-parser kimi_k2 \
  --trust-remote-code
```

## Tips for Clore.ai Users

* **API vs self-hosting tradeoff**: Full K2.5 needs 8× H200 at \~$24–48/day. Moonshot's API is free-tier or pay-per-token — use API for exploration, self-host for sustained production loads.
* **Quantized on single GPU**: The Unsloth GGUF Q2\_K\_XL (\~375GB) can run on an RTX 4090 ($0.5–2/day) with 256GB RAM via CPU offloading — expect \~5–10 tok/s. Good enough for personal use and development.
* **Text-only K2 for budget setups**: If you don't need vision, `moonshotai/Kimi-K2-Instruct` is the text-only predecessor — same 1T MoE but lighter to deploy (no vision encoder overhead).
* **Set temperature correctly**: Use `temperature=0.6` for instant mode, `temperature=1.0` for thinking mode. Wrong temperature causes repetition or incoherence.
* **Expert Parallelism for throughput**: On multi-node setups, use `--enable-expert-parallel` in vLLM for higher throughput. Check vLLM docs for EP configuration.

## Troubleshooting

| Issue                              | Solution                                                                           |
| ---------------------------------- | ---------------------------------------------------------------------------------- |
| `OutOfMemoryError` with full model | Need 8× H200 (1128GB total). Use FP8 weights, set `--gpu-memory-utilization 0.90`. |
| GGUF inference very slow           | Ensure enough RAM for the quant size. Q2\_K\_XL needs \~375GB RAM+VRAM combined.   |
| Vision not working in llama.cpp    | Vision support for K2.5 GGUF is not available yet — use vLLM for multimodal.       |
| Repetitive output                  | Set `temperature=0.6` (instant) or `1.0` (thinking). Add `min_p=0.01`.             |
| Model download takes forever       | \~630GB FP8 checkpoint. Use `huggingface-cli download` with `--resume-download`.   |
| Tool calls not parsed              | Add `--tool-call-parser kimi_k2 --enable-auto-tool-choice` to vLLM serve command.  |

## Further Reading

* [Kimi K2.5 on HuggingFace](https://huggingface.co/moonshotai/Kimi-K2.5)
* [Kimi K2.5 Tech Blog](https://www.kimi.com/blog/kimi-k2-5.html)
* [Kimi K2.5 Paper](https://arxiv.org/abs/2602.02276)
* [vLLM K2.5 Recipe](https://docs.vllm.ai/projects/recipes/en/latest/moonshotai/Kimi-K2.5.html)
* [Unsloth GGUF Quantizations](https://huggingface.co/unsloth/Kimi-K2.5-GGUF)
* [Moonshot API Platform](https://platform.moonshot.ai)
* [Kimi K2 GitHub](https://github.com/MoonshotAI/Kimi-K2)




---

[Next Page](/llms-full.txt/1)

