> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmarks

> Evaluation results and side-by-side comparisons

<Error>
  This page contains placeholder values, assumptions, or outdated info marked in <span style="color:red">red</span>. These need verification from stakeholders (Vijay, Adam, Alex, Product/Legal).
</Error>

# Benchmarks

Radium models are evaluated on standard academic and real-world benchmarks. This page shows how they compare to incumbent models on tasks that matter to production workloads.

## Academic benchmarks

### Reasoning and coding

| Benchmark | hal-1.0 | clarke-1.0 | tycho-1.0 | GPT-4o | Claude 3.5 Sonnet |
| - | - | - | - | - | - |
| MMLU (5-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| HumanEval (0-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| MBPP (3-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| MATH (4-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| GPQA Diamond | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |

### Multilingual and long-context

| Benchmark | hal-1.0 | clarke-1.0 | tycho-1.0 | GPT-4o | Claude 3.5 Sonnet |
| - | - | - | - | - | - |
| MGSM (8-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| LongContext (128K) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| DROP (3-shot) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |

## Real-world task performance

We also evaluate on production-like tasks that aren't captured by academic benchmarks:

| Task | hal-1.0 | clarke-1.0 | tycho-1.0 | GPT-4o | Claude 3.5 Sonnet |
| - | - | - | - | - | - |
| JSON extraction (100 samples) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| Multi-step tool calling | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| Code review ( PR descriptions) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| SQL generation (Spider dev) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |
| Document summarization (Rouge-L) | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> | <span style="color:red">99.9%</span> |

## Side-by-side examples

### Example 1: Reasoning task

**Prompt:** "A train travels 60 mph for 2 hours, then 40 mph for 3 hours. What is the average speed for the entire trip?"

| **hal-1.0** | **GPT-4o** |
| - | - |
| Average speed = total distance / total time = (60×2 + 40×3) / (2+3) = 240/5 = **48 mph** | Same correct answer |

All models answer correctly on this prompt.

### Example 2: Tool-calling task

**Prompt:** "What's the weather in Tokyo? Use the get\_weather tool."

| **clarke-1.0** | **Claude 3.5 Sonnet** |
| - | - |
| Calls `get_weather(location="Tokyo")` correctly | Same correct tool call |

## How we test

* **Temperature:** 0 for deterministic tasks, 0.7 for open-ended tasks
* **Max tokens:** <span style="color:red">99999</span> unless the benchmark specifies otherwise
* **Seed:** <span style="color:red">9999</span> where reproducibility is required
* **Date tested:** <span style="color:red">2026-01-01</span>

## Reproducing the results

Run the suite against your own key:

```bash theme={null}
export RADIUM_API_KEY="YOUR_KEY"
python run_evals.py --model hal-1.0 --subset reasoning
```

## Limitations

* Benchmark scores vary by prompt phrasing and evaluation method
* Real-world performance depends on your specific use case
* We recommend running your own evals on your own data before making a decision


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.