> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Spend controls

> Per-request limits, context-window guardrails and budget management

# Spend controls

Limit your spend at the request level (via `max_tokens`) and handle oversized prompts gracefully before you ship to production.

<Note>
  Budget alerts, daily hard/soft limits, and auto-suspend are on the roadmap. For now the only spend guards are per-request `max_tokens` and manual monitoring through the dashboard.
</Note>

## Per-request token cap

Set `max_tokens` on every request to prevent runaway generation from unbounded prompts or accidental infinite loops.

```python theme={null}
from openai import OpenAI
client = OpenAI(api_key="YOUR_RADIUM_API_KEY", base_url="https://api.radium.cloud/v1")

response = client.chat.completions.create(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "Write a chapter"}],
    max_tokens=4096,
)
```

The absolute context-window limit depends on the model you choose:

| Model | Context window |
| - | - |
| hal-1.0 | 250,000 tokens |
| clarke-1.0 | 1,000,000 tokens |
| tycho-1.0 | 125,000 tokens |

## Rejecting oversized requests

Radium returns `400` with code `context_length_exceeded` if the prompt exceeds the model's context window. Handle this gracefully:

```python theme={null}
try:
    response = client.chat.completions.create(model="clarke-1.0", messages=messages)
except APIError as e:
    if e.code == "context_length_exceeded":
        # Truncate or split the prompt
        messages = messages[-10:]  # Keep last 10 messages
        response = client.chat.completions.create(model="clarke-1.0", messages=messages)
```

## Usage audit

Export a CSV of recent requests from the dashboard for custom analysis:

| Column | Description |
| - | - |
| `timestamp` | UTC request time |
| `model` | Model ID used |
| `input_tokens` | Tokens in the prompt |
| `output_tokens` | Tokens in the response |
| `cost_usd` | Estimated cost for this request |
| `status_code` | HTTP status |

## Pre-production checklist

Before shipping to users, verify:

* [ ] `max_tokens` is set on every request
* [ ] Your app handles `context_length_exceeded` errors
* [ ] Usage is reviewed regularly in the dashboard
* [ ] Budget guardrails (alerts, auto-suspend) — [coming soon on our roadmap](#)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.