> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitoring usage

> Track token usage, latency, and cost per request for observability and budgeting

# Monitoring usage

You can't optimize what you don't measure. This recipe shows how to wrap every Radium call with telemetry — token counts, latency, cost estimates, and error rates — and export them to your observability stack.

## The script

```python monitor.py theme={null}
import os
import time
import json
from dataclasses import dataclass, asdict
from datetime import datetime, timezone
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

# Pricing (per 1M tokens) — populate from https://radium.cloud/pricing
PRICING = {
    "tycho-1.0": {"prompt": 0.00, "completion": 0.00},
    "clarke-1.0": {"prompt": 0.00, "completion": 0.00},
    "hal-1.0": {"prompt": 0.00, "completion": 0.00},
}


@dataclass
class RequestLog:
    timestamp: str
    model: str
    latency_ms: float
    prompt_tokens: int
    completion_tokens: int
    total_tokens: int
    estimated_cost_usd: float
    status: str  # "success" | "error"
    error_type: str | None


def estimate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float:
    rates = PRICING.get(model, PRICING["clarke-1.0"])
    return (
        (prompt_tokens / 1_000_000) * rates["prompt"]
        + (completion_tokens / 1_000_000) * rates["completion"]
    )


def logged_chat(messages, model="clarke-1.0", **kwargs) -> tuple[str, RequestLog]:
    """
    Wraps chat.completions.create with logging.
    Returns (content, log_entry).
    """
    start = time.perf_counter()
    status = "success"
    error_type = None
    usage = None

    try:
        response = client.chat.completions.create(
            model=model,
            messages=messages,
            **kwargs,
        )
        content = response.choices[0].message.content
        usage = response.usage
    except Exception as e:
        status = "error"
        error_type = type(e).__name__
        content = ""

    latency = (time.perf_counter() - start) * 1000

    log = RequestLog(
        timestamp=datetime.now(timezone.utc).isoformat(),
        model=model,
        latency_ms=latency,
        prompt_tokens=usage.prompt_tokens if usage else 0,
        completion_tokens=usage.completion_tokens if usage else 0,
        total_tokens=usage.total_tokens if usage else 0,
        estimated_cost_usd=estimate_cost(
            model,
            usage.prompt_tokens if usage else 0,
            usage.completion_tokens if usage else 0,
        ) if usage else 0.0,
        status=status,
        error_type=error_type,
    )
    return content, log


# --- Run it ---
logs: list[RequestLog] = []

for prompt in [
    "Summarize the theory of relativity in one paragraph.",
    "List 3 benefits of exercise.",
    "Explain recursion with a Python example.",
]:
    content, log = logged_chat(
        messages=[{"role": "user", "content": prompt}],
        model="clarke-1.0",
        max_tokens=256,
    )
    logs.append(log)
    print(f"[{log.model}] {log.status} — {log.total_tokens} tokens, ${log.estimated_cost_usd:.6f}, {log.latency_ms:.0f}ms")

# Export to JSONL
with open("usage_logs.jsonl", "w") as f:
    for log in logs:
        f.write(json.dumps(asdict(log)) + "\n")

# Aggregate summary
if logs:
    total_cost = sum(l.estimated_cost_usd for l in logs)
    avg_latency = sum(l.latency_ms for l in logs) / len(logs)
    total_tokens = sum(l.total_tokens for l in logs)
    error_rate = sum(1 for l in logs if l.status == "error") / len(logs)

    print(f"\nSummary: {len(logs)} requests, {total_tokens} tokens, ${total_cost:.4f}, avg {avg_latency:.0f}ms, {error_rate:.0%} errors")
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python monitor.py
```

## Sample output

```
[clarke-1.0] success — 187 tokens, $0.000748, 423ms
[clarke-1.0] success — 89 tokens, $0.000356, 312ms
[clarke-1.0] success — 156 tokens, $0.000624, 389ms

Summary: 3 requests, 432 tokens, $0.0017, avg 375ms, 0% errors
```

## Shipping to an observability tool

Export `usage_logs.jsonl` to Datadog, Grafana, or any metrics backend:

```python export_to_prometheus.py theme={null}
# Pseudocode — wire into your existing metrics client
from prometheus_client import Counter, Histogram

llm_requests = Counter("llm_requests_total", "Total LLM requests", ["model", "status"])
llm_tokens = Counter("llm_tokens_total", "Total tokens", ["model", "type"])
llm_latency = Histogram("llm_latency_seconds", "Request latency", ["model"])

for log in logs:
    llm_requests.labels(model=log.model, status=log.status).inc()
    llm_tokens.labels(model=log.model, type="prompt").inc(log.prompt_tokens)
    llm_tokens.labels(model=log.model, type="completion").inc(log.completion_tokens)
    llm_latency.labels(model=log.model).observe(log.latency_ms / 1000)
```

## Tips

* **Log every request** — even cached ones (mark them `cached=True`) so your dashboards are complete.
* **Alert on error rate > 1%** and latency p99 > 5s.
* **Tag by model and use-case** so you can drill into which workflows are expensive.
* **Keep pricing maps in config** — update them when rates change (see [radium.cloud/pricing](https://radium.cloud/pricing)).
* **Store raw request/response** (sampled at 1%) for debugging quality regressions.

## Next steps

<CardGroup cols={2}>
  <Card title="Caching responses" icon="clock-rotate-left" href="/examples/caching-responses">Reduce costs with intelligent caching</Card>
  <Card title="A/B testing models" icon="scale-balanced" href="/examples/ab-testing-models">Compare models on quality and cost</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.