> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Switching with fallback

> Route requests to the right model and fall back automatically on failure

# Switching with fallback

Radium supports multiple models. This recipe shows how to route requests intelligently — picking the right model for the task, and automatically falling back if one is overloaded or unavailable.

## The script

```python smart_router.py theme={null}
import os
from openai import OpenAI, RateLimitError, APIError

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

# Model tiers: primary model → fallback chain
MODEL_CHAIN = {
    "reasoning": ["clarke-1.0", "hal-1.0"],   # complex reasoning → fallback to dev
    "fast":      ["tycho-1.0", "clarke-1.0"], # speed → fallback to reliable
    "creative":  ["clarke-1.0", "hal-1.0"],   # creative tasks → fallback to dev
}


def classify_complexity(prompt: str) -> str:
    """Simple keyword-based routing. Replace with a real classifier in production."""
    reasoning_keywords = ["analyze", "compare", "explain why", "step by step", "prove"]
    fast_keywords = ["classify", "tag", "extract", "yes/no", "summarize in one sentence"]

    lower = prompt.lower()
    if any(kw in lower for kw in reasoning_keywords):
        return "reasoning"
    if any(kw in lower for kw in fast_keywords):
        return "fast"
    return "creative"


def chat(prompt: str, model_tier: str | None = None) -> dict:
    """
    Send a chat request with automatic model selection and fallback.
    Returns {"content": str, "model_used": str, "tier": str}
    """
    tier = model_tier or classify_complexity(prompt)
    chain = MODEL_CHAIN.get(tier, MODEL_CHAIN["creative"])

    messages = [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": prompt},
    ]

    for model in chain:
        try:
            response = client.chat.completions.create(
                model=model,
                messages=messages,
                temperature=0.7,
                max_tokens=512,
            )
            return {
                "content": response.choices[0].message.content,
                "model_used": model,
                "tier": tier,
            }
        except (RateLimitError, APIError) as e:
            print(f"  {model} failed ({type(e).__name__}), trying next...")
            continue

    raise RuntimeError("All models in the chain failed.")


# --- Run it ---
queries = [
    (None, "Summarize quantum computing in one sentence."),
    ("reasoning", "Compare and contrast supervised vs unsupervised learning."),
    ("fast", "Classify this sentiment: 'The movie was absolutely terrible.'"),
]

for tier_override, q in queries:
    result = chat(q, model_tier=tier_override)
    print(f"\n[{result['tier']}] → {result['model_used']}")
    print(result["content"])
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python smart_router.py
```

## Sample output

```
[fast] → tycho-1.0
Quantum computing leverages quantum bits to perform certain calculations exponentially faster than classical computers.

[reasoning] → clarke-1.0
Supervised learning uses labeled training data where each input has a known output...

[fast] → tycho-1.0
Negative
```

## Adding cost-based routing

Route based on estimated cost or latency requirements:

```python theme={null}
import tiktoken


def route_by_budget(prompt: str, max_cost_cents: float) -> str:
    encoder = tiktoken.encoding_for_model("gpt-4")
    tokens = len(encoder.encode(prompt))

    # Rough cost estimates (per 1M tokens) — populate from https://radium.cloud/pricing
    pricing = {
        "tycho-1.0": 0.00,
        "clarke-1.0": 0.00,
        "hal-1.0": 0.00,
    }

    for model, price in sorted(pricing.items(), key=lambda x: x[1]):
        est_cost = (tokens / 1_000_000) * price * 100  # in cents
        if est_cost <= max_cost_cents:
            return model

    return "tycho-1.0"  # cheapest fallback
```

## Tips

* **Use `tycho-1.0`** for high-volume, simple tasks — it's fastest and cheapest.
* **Use `clarke-1.0`** for general reasoning, coding, and longer-context work.
* **Use `hal-1.0`** for the hardest reasoning and most nuanced outputs.
* **Log which model served each request** so you can tune your routing rules over time.
* **Monitor latency per model** — measure in your own workload since latency varies by prompt length and time of day.

## Next steps

<CardGroup cols={2}>
  <Card title="Error handling & retries" icon="shield-halved" href="/examples/error-handling-retries">Add exponential backoff and circuit breakers</Card>
  <Card title="Batch processing" icon="layer-group" href="/examples/batch-processing">Process thousands of requests with smart routing</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.