Skip to main content

Error handling & retries

Production workloads hit rate limits, network hiccups, and transient errors. This recipe shows how to wrap your Radium calls with robust retry logic and automatic fallback to a backup model.

The script

resilient_chat.py

Run it

What it does

  1. Retries RateLimitError with exponential backoff + jitter — the standard recipe for 429 responses.
  2. Retries APIConnectionError with shorter backoff for transient network issues.
  3. Retries 5xx errors from the API itself, but gives up on unrecoverable 4xx client errors.
  4. Falls back to tycho-1.0 if clarke-1.0 is overloaded or unavailable.
  5. Fails loudly only after both models are exhausted — no silent swallowing.

Adding circuit-breaker logic

For production services, add a simple circuit breaker that stops hammering a struggling endpoint:

Tips

  • Start with max_retries=3 and tune based on your traffic patterns.
  • Monitor Retry-After headers from rate-limit responses — the example uses fixed backoff; production code should read the header.
  • Log every retry with model_name and attempt so you can spot recurring patterns in your observability tool.
  • Reserve your fallback model for a different tier (e.g. switch from clarke-1.0 to tycho-1.0) so you don’t hit the same capacity constraint.

Next steps

Switching with fallback

Route requests intelligently across models

Tool calling agent

Build an agent that recovers from tool failures