> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Error handling & retries

> Build resilience with exponential backoff, rate-limit handling, and model fallback

# Error handling & retries

Production workloads hit rate limits, network hiccups, and transient errors. This recipe shows how to wrap your Radium calls with robust retry logic and automatic fallback to a backup model.

## The script

```python resilient_chat.py theme={null}
import os
import time
import random
from openai import OpenAI, APIError, RateLimitError, APIConnectionError

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


def chat_with_retry(
    messages,
    model="clarke-1.0",
    fallback_model="tycho-1.0",
    max_retries=4,
    base_delay=1.0,
    max_delay=60.0,
):
    """
    Send a chat completion with exponential-backoff retries.
    Falls back to `fallback_model` if the primary model keeps failing.
    """
    models_to_try = [model, fallback_model]

    for current_model in models_to_try:
        for attempt in range(max_retries):
            try:
                response = client.chat.completions.create(
                    model=current_model,
                    messages=messages,
                    temperature=0.7,
                    max_tokens=512,
                )
                return response.choices[0].message.content

            except RateLimitError:
                # 429 — back off and retry
                delay = min(base_delay * (2 ** attempt) + random.uniform(0, 1), max_delay)
                print(f"  Rate limited on {current_model}. Retrying in {delay:.1f}s...")
                time.sleep(delay)

            except APIConnectionError:
                # Network blip — short retry
                delay = min(base_delay * (2 ** attempt), max_delay)
                print(f"  Connection error on {current_model}. Retrying in {delay:.1f}s...")
                time.sleep(delay)

            except APIError as e:
                # 5xx or other API-level errors — retry a few times, then switch model
                if e.status_code and e.status_code >= 500:
                    delay = min(base_delay * (2 ** attempt), max_delay)
                    print(f"  API error {e.status_code} on {current_model}. Retrying in {delay:.1f}s...")
                    time.sleep(delay)
                else:
                    # 4xx client error — don't retry, just switch or raise
                    print(f"  Client error {e.status_code} on {current_model}. Switching model...")
                    break

        print(f"  Exhausted retries on {current_model}.")

    raise RuntimeError("All models failed. Check your API key or Radium status.")


# --- Run it ---
answer = chat_with_retry(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain recursion in one sentence."},
    ]
)
print(answer)
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python resilient_chat.py
```

## What it does

1. **Retries `RateLimitError`** with exponential backoff + jitter — the standard recipe for 429 responses.
2. **Retries `APIConnectionError`** with shorter backoff for transient network issues.
3. **Retries 5xx errors** from the API itself, but gives up on unrecoverable 4xx client errors.
4. **Falls back** to `tycho-1.0` if `clarke-1.0` is overloaded or unavailable.
5. **Fails loudly** only after both models are exhausted — no silent swallowing.

## Adding circuit-breaker logic

For production services, add a simple circuit breaker that stops hammering a struggling endpoint:

```python theme={null}
circuit_failures = 0
CIRCUIT_THRESHOLD = 5
CIRCUIT_TIMEOUT = 30  # seconds

def chat_with_circuit(messages):
    global circuit_failures
    if circuit_failures >= CIRCUIT_THRESHOLD:
        print("Circuit open — using fallback model only")
        return chat_with_retry(messages, model="tycho-1.0")

    try:
        result = chat_with_retry(messages)
        circuit_failures = 0
        return result
    except Exception:
        circuit_failures += 1
        raise
```

## Tips

* **Start with `max_retries=3`** and tune based on your traffic patterns.
* **Monitor `Retry-After` headers** from rate-limit responses — the example uses fixed backoff; production code should read the header.
* **Log every retry** with `model_name` and `attempt` so you can spot recurring patterns in your observability tool.
* **Reserve your fallback model** for a different tier (e.g. switch from `clarke-1.0` to `tycho-1.0`) so you don't hit the same capacity constraint.

## Next steps

<CardGroup cols={2}>
  <Card title="Switching with fallback" icon="code-compare" href="/examples/switching-with-fallback">Route requests intelligently across models</Card>
  <Card title="Tool calling agent" icon="robot" href="/examples/tool-calling-agent">Build an agent that recovers from tool failures</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.