Performance
Radium is optimized for low latency and high throughput. Understand what to expect and how to measure it.Time to first token (TTFT)
TTFT is the time from sending a request to receiving the first streamed token. Radium measures TTFT internally and will publish live P50/P99 numbers on the Radium status page in a future release. Longer prompts increase TTFT linearly. Cache hits (repeated prompts) reduce TTFT by approximately 50%.Measuring latency in your code
Temperature-0 determinism
Attemperature=0 and seed fixed, Radium produces deterministic output for a given model version. Minor variance (< 1% token difference) may occur across model hot-swaps or infrastructure updates.
For maximum reproducibility, pin the model version in your request headers and pass a fixed seed.
Streaming performance
Always use streaming for latency-sensitive applications. It reduces perceived TTFT to near-zero and allows partial processing of long outputs.Batch processing
For offline workloads, use batch to amortize overhead:Performance tuning checklist
- Use streaming for interactive UIs
- Set
temperature=0for deterministic tasks - Use
hal-1.0for fastest TTFT - Use
tycho-1.0for lower cost on simpler tasks - Enable prompt caching for repeated queries
- Batch offline workloads
- Keep prompts within each model’s context window (Tycho 125K, Hal 250K, Clarke 1M)