> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Migration testing

> Canary deployments, A/B comparison and rollback strategies

# Migration testing

Move from your incumbent to Radium in stages. Start with a small canary, measure output quality and cost, then scale up.

## Overview

| Phase | Traffic split | Goal | Timeline |
| - | - | - | - |
| Phase 0 | 0% | Validate key pages, tool calls, streaming | Day 1 |
| Phase 1 | 5% | Canary on non-critical traffic | Week 1 |
| Phase 2 | 25% | Expand to more endpoints | Week 2-3 |
| Phase 3 | 100% | Full cutover with rollback ready | Week 4+ |

## Phase 0: Pre-flight validation

Before routing any live traffic, run automated checks against both providers with identical prompts.

```python theme={null}
import os, json
from openai import OpenAI

incumbent = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
radium = OpenAI(api_key=os.environ["RADIUM_API_KEY"], base_url="https://api.radium.cloud/v1")

def compare(prompt: str, incumbent_model: str, radium_model: str):
    r1 = incumbent.chat.completions.create(model=incumbent_model, messages=[{"role":"user","content":prompt}])
    r2 = radium.chat.completions.create(model=radium_model, messages=[{"role":"user","content":prompt}])
    return {
        "incumbent": r1.choices[0].message.content,
        "radium": r2.choices[0].message.content,
        "incumbent_tokens": r1.usage.total_tokens,
        "radium_tokens": r2.usage.total_tokens,
    }

results = []
for prompt in open("test-prompts.txt").readlines():
    results.append(compare(prompt.strip(), "gpt-4o", "hal-1.0"))

json.dump(results, open("comparison-results.json", "w"), indent=2)
```

Review `comparison-results.json` for:

* Output correctness on structured tasks
* Reasoning quality on multi-step tasks
* Token efficiency (lower is cheaper)

## Phase 1: 5% canary with LiteLLM

LiteLLM is the simplest way to route a percentage of traffic to Radium without changing application code.

### LiteLLM config

```yaml theme={null}
model_list:
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY
  - model_name: radium-hal
    litellm_params:
      model: openai/hal-1.0
      api_base: https://api.radium.cloud/v1
      api_key: os.environ/RADIUM_API_KEY

router_settings:
  routing_strategy: simple-shuffle
  model_group_alias:
    production: gpt-4o
    canary: radium-hal
```

### Application side

Your app still calls the same proxy endpoint. LiteLLM handles the split:

```python theme={null}
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LITELLM_PROXY_KEY"],
    base_url="http://localhost:4000"
)

# 95% hits gpt-4o, 5% hits hal-1.0
response = client.chat.completions.create(
    model="production",
    messages=[{"role": "user", "content": "Hello"}]
)
```

### Monitoring

Track these metrics per provider during the canary:

| Metric | Incumbent | Radium | Acceptable delta |
| - | - | - | - |
| P50 TTFT | record | record | \< 50% slower |
| Error rate | record | record | \< 1% |
| Token cost / 1K requests | record | record | cheaper or within 10% |
| User satisfaction score | record | record | no drop |

## Phase 2: 25% expansion

Gradually increase the Radium ratio in LiteLLM or via your own weighted router:

```python theme={null}
import random

def pick_provider():
    return "radium" if random.random() < 0.25 else "openai"
```

Add more traffic types:

* Streaming completions
* Tool-calling workflows
* Structured-output pipelines
* Multi-turn conversations

## Phase 3: Full cutover

When Radium has proven reliable, switch the default to 100%.

Keep the incumbent route live as a hot rollback:

```python theme={null}
def chat(messages, provider="radium"):
    if provider == "radium":
        client = radium_client
        model = "hal-1.0"
    else:
        client = incumbent_client
        model = "gpt-4o"
    return client.chat.completions.create(model=model, messages=messages)
```

Set an alert on Radium error rate. If it spikes above 2%, immediately flip the default back to the incumbent.

## Rollback procedure

1. Change the default model alias back to the incumbent
2. Verify the incumbent is responding within 30 seconds
3. Root-cause the Radium failure (check [status page](https://status.radium.cloud))
4. Re-enable Radium once the issue is resolved

## Side-by-side cost worksheet

| Dimension | Incumbent | Radium |
| - | - | - |
| Input tokens / 1K req | measure | measure |
| Output tokens / 1K req | measure | measure |
| Cost per 1K input | measure | see [radium.cloud/pricing](https://radium.cloud/pricing) |
| Cost per 1K output | measure | see [radium.cloud/pricing](https://radium.cloud/pricing) |
| Total cost / 1K req | calculate | calculate |

Fill in real numbers from your dashboard and the Radium dashboard. Radium is typically cheaper per token; the worksheet proves the total cost for your specific workload.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.