> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Capacity

> Queue depth, concurrency limits and throttling signals

# Capacity

Understand how Radium handles traffic and what to do when you encounter transient capacity issues.

<Note>
  Radium does not currently enforce fixed concurrency limits, queue depth limits, or tier-based capacity quotas. Capacity controls may be introduced in the future.
</Note>

## Detecting throttling

Every response includes these headers so you can monitor infrastructure health:

| Header | Description |
| - | - |
| `X-Queue-Depth` | Current queue depth for this model (always `0` today) |
| `X-Queue-Wait-Ms` | Estimated wait time if queued (always `0` today) |

```python theme={null}
import requests

response = requests.post(
    "https://api.radium.cloud/v1/chat/completions",
    headers={"Authorization": f"Bearer {key}"},
    json={"model": "hal-1.0", "messages": [{"role": "user", "content": "Hello"}]}
)

print("Queue depth:", response.headers.get("X-Queue-Depth"))
print("Queue wait (ms):", response.headers.get("X-Queue-Wait-Ms"))
```

## Handling transient errors

During traffic spikes or infrastructure maintenance you may receive a `429 Too Many Requests` or `503 Service Unavailable`. These are retryable.

| Status | Meaning | Action |
| - | - | - |
| `429` | Rate limit hit (rare today) | Back off with exponential jitter |
| `503` | Temporary capacity issue | Retry with exponential backoff |
| TTFT spikes | Model overloaded | Check status page, consider failover |

## Retry strategy

Use exponential backoff with jitter when you encounter `429` or `503`:

```python theme={null}
import random, time

def with_retry(fn, max_retries=5):
    for attempt in range(max_retries):
        try:
            return fn()
        except APIError as e:
            if e.status_code in (429, 503) and attempt < max_retries - 1:
                delay = (2 ** attempt) + random.random()
                time.sleep(delay)
            else:
                raise
```

## Managing concurrency in your code

Even without fixed limits on our side, you should cap concurrent requests in your own integration to avoid overwhelming downstream dependencies. A semaphore is the standard pattern:

```python theme={null}
import asyncio

SEMAPHORE = asyncio.Semaphore(10)  # adjust to your own safe concurrency level

async def safe_request(client, ...):
    async with SEMAPHORE:
        return await client.chat.completions.create(...)
```

## Capacity planning

Estimate your required concurrency before going live:

```
Required concurrent = (requests_per_second × avg_response_time_seconds)
```

If you expect sustained high-volume traffic, [contact sales](https://radium.cloud/contact) so we can plan capacity together.

## Status page

Real-time capacity and incident information is available at [status.radium.cloud](https://status.radium.cloud).

Subscribe to the RSS feed or webhook for automated alerting:

```bash theme={null}
curl https://status.radium.cloud/feed
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.