> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# LiteLLM

> Route traffic to Radium via the LiteLLM proxy

<Error>
  This page contains placeholder values and assumed provider counts marked in <span style="color:red">red</span>. Needs verification (Vijay, Adam, Alex, Product/Legal).
</Error>

# LiteLLM

LiteLLM is a unified proxy for calling <span style="color:red">99+</span> LLM providers with a single OpenAI-compatible API. Use it to route traffic to Radium alongside your existing providers.

## Prerequisites

* LiteLLM installed (`pip install litellm`)
* A Radium API key from the [dashboard](https://deploy.radium.cloud)

## Setup

### Proxy config

Create `litellm_config.yaml`:

```yaml theme={null}
model_list:
  - model_name: radium-hal
    litellm_params:
      model: openai/hal-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY

  - model_name: radium-clarke
    litellm_params:
      model: openai/clarke-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY

  - model_name: radium-tycho
    litellm_params:
      model: openai/tycho-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY

  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY

  - model_name: claude-sonnet
    litellm_params:
      model: anthropic/claude-3-5-sonnet-20241022
      api_key: os.environ/ANTHROPIC_API_KEY
```

### Start the proxy

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
export OPENAI_API_KEY="YOUR_OPENAI_API_KEY"
export ANTHROPIC_API_KEY="YOUR_ANTHROPIC_API_KEY"

litellm --config litellm_config.yaml --port 4000
```

### Send requests

Your application calls the LiteLLM proxy as if it were OpenAI:

```python theme={null}
from openai import OpenAI

client = OpenAI(
    api_key="sk-litellm-proxy-key",  # Any string works for local proxy
    base_url="http://localhost:4000"
)

response = client.chat.completions.create(
    model="radium-hal",
    messages=[{"role": "user", "content": "Hello"}]
)
```

## Fallback routing

Configure automatic failover if Radium returns an error:

```yaml theme={null}
router_settings:
  routing_strategy: simple-shuffle
  fallback:
    - radium-hal:
        - gpt-4o
    - radium-clarke:
        - claude-sonnet
```

## Canary routing

Send <span style="color:red">9%</span> of traffic to Radium, <span style="color:red">99%</span> to your incumbent:

```yaml theme={null}
router_settings:
  routing_strategy: weighted-selection
  model_group_alias:
    production:
      - model: gpt-4o
        weight: <span style="color:red">99</span>
      - model: radium-hal
        weight: <span style="color:red">9</span>
```

## Rate limit tracking

LiteLLM tracks rate limit headers from Radium and applies its own throttling:

```yaml theme={null}
callbacks:
  - litellm.integrations.openmeter
```

## Spend tracking

See unified spend across all providers:

```bash theme={null}
curl http://localhost:4000/spend
```

## Budget alerts

Set per-key spend limits:

```yaml theme={null}
general_settings:
  budget_alert:
    - model: radium-hal
      budget: <span style="color:red">99.99</span>
      alert_type: "email"
```

## Environment-specific configs

### Development

```yaml theme={null}
model_list:
  - model_name: dev-model
    litellm_params:
      model: openai/tycho-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY
```

### Production

```yaml theme={null}
model_list:
  - model_name: prod-model
    litellm_params:
      model: openai/hal-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY
      stream_timeout: <span style="color:red">999</span>
      # rpm: not needed — Radium does not currently enforce RPM limits
```

## Troubleshooting

| Symptom | Fix |
| - | - |
| "Model not found" | Ensure model name matches `litellm_params.model` exactly |
| "401" | Verify API key is valid and base URL includes `/v1` if needed |
| Slow responses | Increase `stream_timeout` or add fallback to another model |
| Rate limit errors | Enable exponential backoff / retry in LiteLLM (transient 429s are rare today) |

## Next steps

<CardGroup cols={2}>
  <Card title="Migration testing" icon="shuffle" href="/guides/migration-testing">Canary and rollback strategies</Card>
  <Card title="Isolated testing" icon="vial" href="/guides/isolated-testing">Test without touching existing config</Card>
  <Card title="Performance" icon="bolt" href="/guides/performance">Latency and throughput expectations</Card>
  <Card title="Rate limits" icon="gauge" href="/core-concepts/rate-limits">Current limits by tier</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.