Skip to main content

Migration testing

Move from your incumbent to Radium in stages. Start with a small canary, measure output quality and cost, then scale up.

Overview

Phase 0: Pre-flight validation

Before routing any live traffic, run automated checks against both providers with identical prompts.
Review comparison-results.json for:
  • Output correctness on structured tasks
  • Reasoning quality on multi-step tasks
  • Token efficiency (lower is cheaper)

Phase 1: 5% canary with LiteLLM

LiteLLM is the simplest way to route a percentage of traffic to Radium without changing application code.

LiteLLM config

Application side

Your app still calls the same proxy endpoint. LiteLLM handles the split:

Monitoring

Track these metrics per provider during the canary:

Phase 2: 25% expansion

Gradually increase the Radium ratio in LiteLLM or via your own weighted router:
Add more traffic types:
  • Streaming completions
  • Tool-calling workflows
  • Structured-output pipelines
  • Multi-turn conversations

Phase 3: Full cutover

When Radium has proven reliable, switch the default to 100%. Keep the incumbent route live as a hot rollback:
Set an alert on Radium error rate. If it spikes above 2%, immediately flip the default back to the incumbent.

Rollback procedure

  1. Change the default model alias back to the incumbent
  2. Verify the incumbent is responding within 30 seconds
  3. Root-cause the Radium failure (check status page)
  4. Re-enable Radium once the issue is resolved

Side-by-side cost worksheet

Fill in real numbers from your dashboard and the Radium dashboard. Radium is typically cheaper per token; the worksheet proves the total cost for your specific workload.