Skip to main content

LiteLLM

LiteLLM is a unified proxy for calling 99+ LLM providers with a single OpenAI-compatible API. Use it to route traffic to Radium alongside your existing providers.

Prerequisites

  • LiteLLM installed (pip install litellm)
  • A Radium API key from the dashboard

Setup

Proxy config

Create litellm_config.yaml:

Start the proxy

Send requests

Your application calls the LiteLLM proxy as if it were OpenAI:

Fallback routing

Configure automatic failover if Radium returns an error:

Canary routing

Send 9% of traffic to Radium, 99% to your incumbent:

Rate limit tracking

LiteLLM tracks rate limit headers from Radium and applies its own throttling:

Spend tracking

See unified spend across all providers:

Budget alerts

Set per-key spend limits:

Environment-specific configs

Development

Production

Troubleshooting

Next steps

Migration testing

Canary and rollback strategies

Isolated testing

Test without touching existing config

Performance

Latency and throughput expectations

Rate limits

Current limits by tier