> For the complete documentation index, see [llms.txt](https://docs.clore.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.clore.ai/guides/guides_v2-es/devops-con-gpu/onnx-runtime.md).

# ONNX Runtime GPU

> **Inferencia de ML multiplataforma y acelerada por hardware — despliega cualquier modelo desde cualquier framework**

ONNX Runtime (ORT) es el motor de inferencia de código abierto de Microsoft para modelos ONNX (Open Neural Network Exchange). Proporciona inferencia acelerada por hardware en CPUs, GPUs y aceleradores especializados a través de una API unificada. Tanto si tu modelo se entrenó en PyTorch, TensorFlow, Scikit-learn o XGBoost — si puedes exportarlo al formato ONNX, ORT puede ejecutarlo más rápido.

**GitHub:** [microsoft/onnxruntime](https://github.com/microsoft/onnxruntime) — 14K+ ⭐

***

## ¿Por qué ONNX Runtime?

| Función                     | ONNX Runtime      | TorchScript    | TensorFlow Serving |
| --------------------------- | ----------------- | -------------- | ------------------ |
| Agnóstico al framework      | ✅                 | ❌ Solo PyTorch | ❌ Solo TF          |
| Aceleración GPU             | ✅ CUDA/TensorRT   | ✅              | ✅                  |
| Cuantización INT8/FP16      | ✅                 | Parcial        | Parcial            |
| Despliegue móvil/periférico | ✅                 | Limitado       | Limitado           |
| Fusión de operadores        | ✅                 | Parcial        | ✅                  |
| Integración fácil           | ✅ Python/C++/Java | Python         | Python/gRPC        |

{% hint style="success" %}
**Beneficio clave:** ONNX Runtime con el proveedor de ejecución CUDA normalmente ofrece **una aceleración de 1.5–3x** frente a la inferencia nativa de PyTorch para modelos de visión por computadora y NLP.
{% endhint %}

***

## Proveedores de ejecución compatibles

ONNX Runtime admite múltiples backends de hardware (proveedores de ejecución):

| Proveedor                   | Hardware      | Caso de uso               |
| --------------------------- | ------------- | ------------------------- |
| `CUDAExecutionProvider`     | GPUs NVIDIA   | Inferencia general en GPU |
| `TensorrtExecutionProvider` | GPUs NVIDIA   | Máximo rendimiento        |
| `CPUExecutionProvider`      | CPU           | Alternativa / edge        |
| `ROCMExecutionProvider`     | GPUs AMD      | Hardware AMD              |
| `CoreMLExecutionProvider`   | Apple Silicon | macOS/iOS                 |
| `OpenVINOExecutionProvider` | Intel         | CPUs/GPUs Intel           |

***

## Requisitos previos

* Cuenta de Clore.ai con alquiler de una GPU
* Conocimientos básicos de Python
* Un modelo entrenado (PyTorch, TensorFlow o ONNX preexportado)

***

## Paso 1 — Alquila una GPU en Clore.ai

1. Ve a [clore.ai](https://clore.ai) → **Marketplace**
2. Cualquier GPU NVIDIA funciona — desde una RTX 3070 para modelos pequeños hasta una A100 para grandes transformers
3. **Para modelos transformer:** Se recomienda RTX 4090 o A100
4. **Para visión por computadora:** Una RTX 3090 o RTX 4090 es suficiente

***

## Paso 2 — Despliega tu contenedor

ONNX Runtime no tiene un contenedor oficial precompilado, pero la base de NVIDIA CUDA es ideal:

**Imagen de Docker:**

```
nvcr.io/nvidia/cuda:12.8.1-cudnn-devel-ubuntu22.04
```

**Puertos:**

```
22
```

**Variables de entorno:**

```
NVIDIA_VISIBLE_DEVICES=all
NVIDIA_DRIVER_CAPABILITIES=compute,utility
```

{% hint style="info" %}
Como alternativa, usa `pytorch/pytorch:2.11.0-cuda12.8-cudnn9-runtime` que incluye CUDA y un entorno de Python listo para la instalación de ORT.
{% endhint %}

***

## Paso 3 — Instala ONNX Runtime con soporte GPU

```bash
ssh root@<ip-del-servidor> -p <puerto-ssh>

# Actualiza los paquetes
apt-get update && apt-get install -y \\
    python3-pip \\
    python3-dev \
    wget \\
    git \\
    libgomp1

# Instala ONNX Runtime con soporte CUDA
pip install onnxruntime-gpu

# Instala los paquetes de apoyo
pip install \
    onnx \
    numpy \
    Pillow \
    transformers \
    torch \\
    torchvision \
    fastapi \\
    uvicorn

# Verificar la instalación
python3 << 'EOF'
import onnxruntime as ort
print(f"Versión de ORT: {ort.__version__}")
print(f"Proveedores disponibles: {ort.get_available_providers()}")
# Debería incluir: CUDAExecutionProvider, TensorrtExecutionProvider, CPUExecutionProvider
EOF
```

***

## Paso 4 — Exporta tu modelo a ONNX

### Exportación de modelo PyTorch

```python
import torch
import torch.nn as nn
import onnx

# Ejemplo: exportar ResNet50
model = torch.hub.load('pytorch/vision:v0.10.0', 'resnet50', pretrained=True)
model.eval()

# Crear entrada ficticia (lote=1, imagen RGB 224x224)
dummy_input = torch.randn(1, 3, 224, 224)

# Exportar a ONNX
torch.onnx.export(
    model,
    dummy_input,
    "resnet50.onnx",
    export_params=True,
    opset_version=17,              # Usa el último opset estable
    do_constant_folding=True,      # Optimiza operaciones constantes
    input_names=["input"],
    output_names=["output"],
    dynamic_axes={
        "input": {0: "batch_size"},    # Lote dinámico
        "output": {0: "batch_size"}
    }
)
print("¡Modelo exportado con éxito!")

# Verificar el modelo exportado
onnx_model = onnx.load("resnet50.onnx")
onnx.checker.check_model(onnx_model)
print("¡El modelo ONNX es válido!")
```

### Exportación de HuggingFace Transformers

```bash
# Instala optimum para la exportación a ONNX de HuggingFace
pip install optimum[exporters]

# Exporta BERT para clasificación de texto
optimum-cli export onnx \
    --model bert-base-uncased \
    --task text-classification \
    ./bert_onnx/

# Exportar con optimización
optimum-cli export onnx \
    --model microsoft/phi-2 \
    --task text-generation \
    --optimize O2 \
    ./phi2_onnx/
```

### Exportar con optimización de ORT

```python
from optimum.onnxruntime import ORTModelForSequenceClassification
from optimum.onnxruntime.configuration import OptimizationConfig, ORTConfig
from optimum.onnxruntime import ORTOptimizer

# Cargar y optimizar
model = ORTModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased-finetuned-sst-2-english",
    export=True
)

optimizer = ORTOptimizer.from_pretrained(model)
optimization_config = OptimizationConfig(
    optimization_level=2,
    optimize_for_gpu=True,
    fp16=True
)

optimizer.optimize(
    save_dir="./distilbert_optimized",
    optimization_config=optimization_config
)
```

***

## Paso 5 — Ejecuta inferencia con ONNX Runtime

### Inferencia básica en GPU

```python
import onnxruntime as ort
import numpy as np
from PIL import Image
import torchvision.transforms as transforms

# Configura la sesión con proveedores de ejecución GPU
# Los proveedores se prueban en orden — primero CUDA, luego fallback a CPU
providers = [
    ("CUDAExecutionProvider", {
        "device_id": 0,
        "arena_extend_strategy": "kNextPowerOfTwo",
        "gpu_mem_limit": 4 * 1024 * 1024 * 1024,  # límite de 4 GB
        "cudnn_conv_algo_search": "EXHAUSTIVE",
        "do_copy_in_default_stream": True,
    }),
    "CPUExecutionProvider"
]

# Opciones de sesión para rendimiento
opts = ort.SessionOptions()
opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
opts.intra_op_num_threads = 8
opts.execution_mode = ort.ExecutionMode.ORT_PARALLEL

# Cargar modelo
session = ort.InferenceSession(
    "resnet50.onnx",
    sess_options=opts,
    providers=providers
)

print(f"Ejecutando en: {session.get_providers()}")

# Preparar la entrada
transform = transforms.Compose([
    transforms.Resize(256),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

img = Image.open("test_image.jpg").convert("RGB")
img_tensor = transform(img).unsqueeze(0).numpy()

# Ejecutar inferencia
outputs = session.run(None, {"input": img_tensor})
probabilities = outputs[0][0]
top5_idx = probabilities.argsort()[-5:][::-1]
print("Las 5 mejores predicciones:", top5_idx, probabilities[top5_idx])
```

### Inferencia por lotes para mayor rendimiento

```python
import onnxruntime as ort
import numpy as np
import time

session = ort.InferenceSession(
    "resnet50.onnx",
    providers=["CUDAExecutionProvider"]
)

# Calentar la GPU
dummy = np.random.randn(1, 3, 224, 224).astype(np.float32)
for _ in range(10):
    session.run(None, {"input": dummy})

# Evaluar el rendimiento por tamaño de lote
for batch_size in [1, 4, 8, 16, 32, 64]:
    inputs = np.random.randn(batch_size, 3, 224, 224).astype(np.float32)
    
    start = time.time()
    n_iter = 100
    for _ in range(n_iter):
        session.run(None, {"input": inputs})
    elapsed = time.time() - start
    
    throughput = (batch_size * n_iter) / elapsed
    latency = (elapsed / n_iter) * 1000  # ms
    
    print(f"Lote {batch_size:3d}: {throughput:7.1f} img/seg, {latency:.1f} ms/lote")
```

***

## Paso 6 — Proveedor de ejecución TensorRT (máximo rendimiento)

Para GPUs NVIDIA, TensorRT EP ofrece un rendimiento aún mejor:

```python
import onnxruntime as ort
import numpy as np

# Configuración del proveedor de ejecución TensorRT
tensorrt_provider_options = {
    "trt_max_workspace_size": 4 * 1024 * 1024 * 1024,  # 4 GB
    "trt_fp16_enable": True,          # Habilita FP16 para una inferencia más rápida
    "trt_int8_enable": False,
    "trt_engine_cache_enable": True,   # Guarda en caché los motores compilados
    "trt_engine_cache_path": "/tmp/trt_cache",
    "trt_max_partition_iterations": 1000,
    "trt_min_subgraph_size": 1,
    "trt_timing_cache_enable": True,
}

providers = [
    ("TensorrtExecutionProvider", tensorrt_provider_options),
    ("CUDAExecutionProvider", {"device_id": 0}),
    "CPUExecutionProvider"
]

session = ort.InferenceSession("resnet50.onnx", providers=providers)
print("Proveedor activo:", session.get_providers()[0])

# La primera ejecución compila el motor TensorRT (puede tardar de 1 a 3 minutos)
# Las ejecuciones posteriores usan el motor en caché y son muy rápidas
```

{% hint style="warning" %}
**Compilación del motor TensorRT** ocurre en la primera inferencia y puede tardar de 1 a 5 minutos. Habilita el almacenamiento en caché (`trt_engine_cache_enable: True`) para que el motor compilado se reutilice entre sesiones.
{% endhint %}

***

## Paso 7 — Cuantización INT8 para máxima velocidad

```python
from onnxruntime.quantization import quantize_dynamic, quantize_static, QuantType
import onnxruntime as ort
import numpy as np

# Cuantización INT8 dinámica (no se necesitan datos de calibración)
quantize_dynamic(
    model_input="resnet50.onnx",
    model_output="resnet50_int8_dynamic.onnx",
    weight_type=QuantType.QInt8
)

# Cuantización INT8 estática (requiere datos de calibración)
from onnxruntime.quantization import CalibrationDataReader

class ImageCalibrationReader(CalibrationDataReader):
    def __init__(self, data_dir, input_name="input"):
        self.data_dir = data_dir
        self.input_name = input_name
        self.images = self._load_images()
        self.idx = 0
    
    def _load_images(self):
        # Cargar 100 imágenes de calibración
        import glob, torchvision.transforms as T
        from PIL import Image
        transform = T.Compose([T.Resize(256), T.CenterCrop(224), T.ToTensor()])
        images = []
        for path in glob.glob(f"{self.data_dir}/*.jpg")[:100]:
            img = Image.open(path).convert("RGB")
            images.append(transform(img).numpy())
        return images
    
    def get_next(self):
        if self.idx >= len(self.images):
            return None
        data = {self.input_name: self.images[self.idx:self.idx+1]}
        self.idx += 1
        return data

from onnxruntime.quantization import quantize_static, QuantFormat
quantize_static(
    model_input="resnet50.onnx",
    model_output="resnet50_int8_static.onnx",
    calibration_data_reader=ImageCalibrationReader("/data/calibration_images"),
    quant_format=QuantFormat.QDQ,
    weight_type=QuantType.QInt8
)
```

***

## Paso 8 — Crea una API de inferencia

```bash
cat > /workspace/onnx_api.py << 'EOF'
from fastapi import FastAPI, File, UploadFile
from fastapi.responses import JSONResponse
import onnxruntime as ort
import numpy as np
from PIL import Image
import io
import torchvision.transforms as transforms
import json

app = FastAPI(title="API de inferencia de ONNX Runtime")

# Cargar el modelo al inicio
session = ort.InferenceSession(
    "resnet50.onnx",
    providers=["CUDAExecutionProvider", "CPUExecutionProvider"]
)

# Cargar las etiquetas de clase de ImageNet
with open("imagenet_classes.json") as f:
    classes = json.load(f)

transform = transforms.Compose([
    transforms.Resize(256),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

@app.get("/health")
async def health():
    return {"status": "ok", "providers": session.get_providers()}

@app.post("/predict")
async def predict(file: UploadFile = File(...), topk: int = 5):
    image_data = await file.read()
    img = Image.open(io.BytesIO(image_data)).convert("RGB")
    tensor = transform(img).unsqueeze(0).numpy()
    
    outputs = session.run(None, {"input": tensor})[0][0]
    top_indices = outputs.argsort()[-topk:][::-1]
    
    results = [
        {"label": classes[str(i)], "score": float(outputs[i])}
        for i in top_indices
    ]
    return JSONResponse({"predictions": results})

if __name__ == "__main__":
    import uvicorn
    uvicorn.run(app, host="0.0.0.0", port=8080)
EOF

python3 /workspace/onnx_api.py &

# Probar la API
curl -X POST "http://localhost:8080/predict" \
    -H "accept: application/json" \
    -F "file=@test_image.jpg"
```

***

## Paso 9 — Monitoriza el uso de la GPU

```bash
# Monitorización de GPU en tiempo real durante la inferencia
watch -n 0.5 nvidia-smi

# O usa nvitop para una mejor interfaz
pip install nvitop
nvitop
```

***

## Puntos de referencia de rendimiento

| Modelo    | GPU      | Proveedor     | Rendimiento (inf/seg) |
| --------- | -------- | ------------- | --------------------- |
| ResNet50  | RTX 4090 | CUDA          | \~4,200               |
| ResNet50  | RTX 4090 | TensorRT FP16 | \~8,500               |
| BERT Base | RTX 4090 | CUDA          | \~380                 |
| BERT Base | RTX 4090 | TensorRT FP16 | \~720                 |
| YOLOv8n   | RTX 3090 | CUDA          | \~1,800               |
| YOLOv8x   | A100     | TensorRT FP16 | \~920                 |

***

## Solución de problemas

### Proveedor CUDA no disponible

```bash
# Verifica que esté instalado ORT con CUDA (no la versión solo-CPU)
pip uninstall onnxruntime
pip install onnxruntime-gpu

python3 -c "import onnxruntime as ort; print(ort.get_available_providers())"
```

### Errores de compilación de TensorRT

```bash
# Verifica la compatibilidad de la versión de TensorRT
python3 -c "import tensorrt; print(tensorrt.__version__)"

# Usa en su lugar CUDA EP
providers = ["CUDAExecutionProvider"]  # Omitir TensorRT EP
```

### Errores de desajuste de formas

```python
# Verifica las formas de entrada/salida del modelo
for input in session.get_inputs():
    print(f"Entrada: {input.name}, forma: {input.shape}, tipo: {input.type}")

for output in session.get_outputs():
    print(f"Salida: {output.name}, forma: {output.shape}, tipo: {output.type}")
```

***

## Avanzado: canalización multimodelo

```python
import onnxruntime as ort
import numpy as np

class MultiModelPipeline:
    def __init__(self):
        providers = ["CUDAExecutionProvider"]
        self.detector = ort.InferenceSession("detector.onnx", providers=providers)
        self.classifier = ort.InferenceSession("classifier.onnx", providers=providers)
    
    def run(self, image: np.ndarray) -> list:
        # Paso 1: detección de objetos
        boxes = self.detector.run(None, {"image": image})[0]
        
        results = []
        for box in boxes:
            # Recortar la región detectada
            crop = self._crop(image, box)
            
            # Paso 2: clasificar cada región
            label = self.classifier.run(None, {"input": crop})[0]
            results.append({"box": box.tolist(), "label": int(label.argmax())})
        
        return results
    
    def _crop(self, image, box):
        x1, y1, x2, y2 = box.astype(int)
        return image[:, :, y1:y2, x1:x2]

pipeline = MultiModelPipeline()
```

***

## Recursos adicionales

* [GitHub de ONNX Runtime](https://github.com/microsoft/onnxruntime)
* [Documentación de ONNX Runtime](https://onnxruntime.ai/docs/)
* [Hugging Face Optimum](https://huggingface.co/docs/optimum/)
* [ONNX Model Zoo](https://github.com/onnx/models) — modelos preexportados
* [Netron](https://netron.app/) — visualizador de modelos ONNX
* [API de Python de ONNX Runtime](https://onnxruntime.ai/docs/api/python/)

***

*ONNX Runtime en Clore.ai es la opción ideal para servicios de inferencia en producción que necesitan servir modelos de distintos frameworks con la máxima eficiencia de GPU.*

***

## Recomendaciones de GPU para Clore.ai

| Caso de uso              | GPU recomendada | Costo estimado en Clore.ai                |
| ------------------------ | --------------- | ----------------------------------------- |
| Desarrollo/Pruebas       | RTX 3090 (24GB) | $0.07–0.21/gpu/hr                         |
| Inferencia en producción | RTX 4090 (24GB) | $0.14–0.42/gpu/hr                         |
| Despliegue a gran escala | A100 80GB       | [bare metal](https://clore.ai/bare-metal) |

> 💡 Todos los ejemplos de esta guía pueden desplegarse en [Clore.ai](https://clore.ai/marketplace) servidores GPU. Explora las GPUs disponibles y alquila por hora: sin compromisos, acceso root completo.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.clore.ai/guides/guides_v2-es/devops-con-gpu/onnx-runtime.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
