> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Performance

> TTFT, latency and throughput expectations

# Performance

Radium is optimized for low latency and high throughput. Understand what to expect and how to measure it.

## Time to first token (TTFT)

TTFT is the time from sending a request to receiving the first streamed token.

Radium measures TTFT internally and will publish live P50/P99 numbers on the [Radium status page](https://status.radium.cloud) in a future release.

Longer prompts increase TTFT linearly. Cache hits (repeated prompts) reduce TTFT by approximately 50%.

## Measuring latency in your code

```python theme={null}
import time
from openai import OpenAI

client = OpenAI(api_key="YOUR_RADIUM_API_KEY", base_url="https://api.radium.cloud/v1")

start = time.time()
response = client.chat.completions.create(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

first_token_time = None
for chunk in response:
    if first_token_time is None and chunk.choices[0].delta.content:
        first_token_time = time.time() - start
        print(f"TTFT: {first_token_time * 1000:.0f}ms")
```

## Temperature-0 determinism

At `temperature=0` and `seed` fixed, Radium produces deterministic output for a given model version. Minor variance (\< 1% token difference) may occur across model hot-swaps or infrastructure updates.

For maximum reproducibility, pin the model version in your request headers and pass a fixed `seed`.

## Streaming performance

Always use streaming for latency-sensitive applications. It reduces perceived TTFT to near-zero and allows partial processing of long outputs.

```python theme={null}
# Non-streaming: wait for entire response
response = client.chat.completions.create(model="hal-1.0", messages=messages)
text = response.choices[0].message.content

# Streaming: first token arrives quickly, rest follows
for chunk in client.chat.completions.create(model="hal-1.0", messages=messages, stream=True):
    print(chunk.choices[0].delta.content, end="")
```

## Batch processing

For offline workloads, use batch to amortize overhead:

```python theme={null}
import asyncio

async def batch_requests(prompts, model="tycho-1.0"):
    tasks = [
        asyncio.to_thread(client.chat.completions.create, model=model, messages=[{"role": "user", "content": p}])
        for p in prompts
    ]
    return await asyncio.gather(*tasks)

results = asyncio.run(batch_requests(prompts))
```

## Performance tuning checklist

* [ ] Use streaming for interactive UIs
* [ ] Set `temperature=0` for deterministic tasks
* [ ] Use `hal-1.0` for fastest TTFT
* [ ] Use `tycho-1.0` for lower cost on simpler tasks
* [ ] Enable prompt caching for repeated queries
* [ ] Batch offline workloads
* [ ] Keep prompts within each model's context window (Tycho 125K, Hal 250K, Clarke 1M)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.