Skip to main content

Spend controls

Limit your spend at the request level (via max_tokens) and handle oversized prompts gracefully before you ship to production.
Budget alerts, daily hard/soft limits, and auto-suspend are on the roadmap. For now the only spend guards are per-request max_tokens and manual monitoring through the dashboard.

Per-request token cap

Set max_tokens on every request to prevent runaway generation from unbounded prompts or accidental infinite loops.
The absolute context-window limit depends on the model you choose:

Rejecting oversized requests

Radium returns 400 with code context_length_exceeded if the prompt exceeds the model’s context window. Handle this gracefully:

Usage audit

Export a CSV of recent requests from the dashboard for custom analysis:

Pre-production checklist

Before shipping to users, verify:
  • max_tokens is set on every request
  • Your app handles context_length_exceeded errors
  • Usage is reviewed regularly in the dashboard
  • Budget guardrails (alerts, auto-suspend) — coming soon on our roadmap