> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# A/B testing models

> Compare two models on the same prompt to find the best fit for your task

# A/B testing models

Not sure whether `tycho-1.0`, `clarke-1.0`, or `hal-1.0` is right for your use case? Run the same prompt through multiple models and score the outputs side by side.

## The script

```python ab_test.py theme={null}
import os
import asyncio
from openai import AsyncOpenAI, BadRequestError
from dataclasses import dataclass
from typing import Callable

client = AsyncOpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

MODELS = ["tycho-1.0", "clarke-1.0", "hal-1.0"]


@dataclass
class TrialResult:
    model: str
    content: str
    latency_ms: float
    prompt_tokens: int
    completion_tokens: int
    total_tokens: int


async def run_trial(prompt: str, model: str) -> TrialResult:
    import time
    start = time.perf_counter()
    try:
        response = await client.chat.completions.create(
            model=model,
            messages=[
                {"role": "system", "content": "You are a helpful assistant."},
                {"role": "user", "content": prompt},
            ],
            temperature=0.7,
            max_tokens=512,
        )
    except BadRequestError as e:
        return TrialResult(
            model=model,
            content=f"ERROR: {e}",
            latency_ms=0,
            prompt_tokens=0,
            completion_tokens=0,
            total_tokens=0,
        )

    latency = (time.perf_counter() - start) * 1000
    usage = response.usage
    return TrialResult(
        model=model,
        content=response.choices[0].message.content,
        latency_ms=latency,
        prompt_tokens=usage.prompt_tokens,
        completion_tokens=usage.completion_tokens,
        total_tokens=usage.total_tokens,
    )


def score_length(result: TrialResult) -> float:
    """Prefer concise answers (higher score = worse)."""
    return len(result.content)


def score_formatting(result: TrialResult) -> float:
    """Prefer answers with bullet points or numbered lists."""
    return 1.0 if any(marker in result.content for marker in ["- ", "1. ", "* "]) else 0.0


def print_comparison(results: list[TrialResult]):
    print(f"\n{'Model':<15} {'Latency':>10} {'Prompt':>8} {'Completion':>12} {'Total':>8}")
    print("-" * 60)
    for r in results:
        print(
            f"{r.model:<15} "
            f"{r.latency_ms:>9.0f}ms "
            f"{r.prompt_tokens:>8} "
            f"{r.completion_tokens:>12} "
            f"{r.total_tokens:>8}"
        )

    print("\n--- Outputs ---")
    for r in results:
        print(f"\n[{r.model}]:\n{r.content[:300]}...")


async def main():
    prompt = (
        "Explain how HTTP/3 works and why it's faster than HTTP/2. "
        "Use bullet points."
    )

    print(f"Prompt: {prompt}\n")
    results = await asyncio.gather(*[run_trial(prompt, m) for m in MODELS])

    # Custom scoring
    for scorer in [score_length, score_formatting]:
        results.sort(key=scorer, reverse=(scorer.__name__ == "score_formatting"))
        winner = results[0]
        print(f"\n🏆 Winner by {scorer.__name__}: {winner.model}")

    print_comparison(results)


if __name__ == "__main__":
    asyncio.run(main())
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python ab_test.py
```

## What it does

1. **Runs the same prompt** through `tycho-1.0`, `clarke-1.0`, and `hal-1.0` concurrently.
2. **Measures latency** precisely with `time.perf_counter()`.
3. **Records token usage** for cost comparison.
4. **Scores outputs** with custom heuristics — length, formatting quality, or anything else you define.
5. **Prints a side-by-side comparison** table with latency, tokens, and truncated responses.

## Adding an LLM-as-judge scorer

Use a stronger model to evaluate outputs from weaker ones:

```python theme={null}
async def llm_judge_score(prompt: str, result: TrialResult) -> float:
    eval_prompt = (
        f"Rate the following answer to the question on a scale of 1-10.\n\n"
        f"Question: {prompt}\n\n"
        f"Answer: {result.content}\n\n"
        f"Respond with only the number."
    )
    response = await client.chat.completions.create(
        model="clarke-1.0",
        messages=[{"role": "user", "content": eval_prompt}],
        temperature=0.0,
        max_tokens=5,
    )
    try:
        return float(response.choices[0].message.content.strip())
    except ValueError:
        return 0.0
```

## Tips

* **Run at least 10 trials per model** — single-shot comparisons can be noisy.
* **Score on what matters** to your product: accuracy, conciseness, latency, cost, or safety.
* **Keep a golden dataset** of prompts you run on every model update to catch regressions.
* **Store results in a spreadsheet or DB** so you can trend quality over time.

## Next steps

<CardGroup cols={2}>
  <Card title="Switching with fallback" icon="code-compare" href="/examples/switching-with-fallback">Route requests to the winning model in production</Card>
  <Card title="Monitoring usage" icon="chart-line" href="/examples/monitoring-usage">Track quality and cost metrics over time</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.