---
name: radium
description: Use when building LLM applications, integrating models into frameworks, switching from OpenAI or Anthropic, implementing tool calling agents, structured outputs, or streaming responses. Agents should reach for this skill when working with chat completions, model selection, API integration, error handling, or framework-specific implementations.
metadata:
    mintlify-proj: radium
    version: "1.0"
---

# Radium Skill Reference

## Product summary

Radium is an OpenAI- and Anthropic-compatible LLM API serving three tuned model tiers: Hal 1.0 (reasoning-class), Clarke 1.0 (Sonnet-class, default), and Tycho 1.0 (Haiku-class, high-volume). Send requests to `https://api.radium.cloud/v1` (OpenAI format) or `https://api.radium.cloud` (Anthropic format) using your API key. All three models support streaming, tool calling, JSON mode, and structured outputs. Switch models by changing the `model` parameter; everything else stays the same. Access the full documentation at https://radium-e5ab5fa4.mintlify.app.

## When to use

Reach for this skill when:
- Building chat applications or agents that need LLM inference
- Migrating from OpenAI (gpt-5, gpt-4.1) or Anthropic (Claude 4 Sonnet, Haiku)
- Implementing tool calling agents or multi-step workflows
- Extracting structured JSON from unstructured text
- Streaming responses to users in real time
- Integrating with frameworks (LangChain, LangGraph, Vercel AI SDK, PydanticAI, CrewAI, etc.)
- Comparing model performance on the same prompt
- Handling rate limits, errors, or token budgets

## Quick reference

### API endpoints and authentication

| Format | Base URL | Auth header | Model strings |
|---|---|---|---|
| OpenAI | `https://api.radium.cloud/v1` | `Authorization: Bearer $RADIUM_API_KEY` | `hal-1.0`, `clarke-1.0`, `tycho-1.0` |
| Anthropic | `https://api.radium.cloud` | `x-api-key: $RADIUM_API_KEY` | `hal-1.0`, `clarke-1.0`, `tycho-1.0` |

### Model selection

| Model | Class | Best for | Context | Reasoning | Vision |
|---|---|---|---|---|---|
| **Hal 1.0** | Opus-class | Complex reasoning, multi-step agents, deep code analysis | 250K tokens | Yes | Yes |
| **Clarke 1.0** | Sonnet-class | RAG, copilots, tool calls, structured outputs (default) | 1M tokens | Yes | Yes |
| **Tycho 1.0** | Haiku-class | Classification, extraction, summarization, routing | 125K tokens | No | No |

### Common parameters

```python
response = client.chat.completions.create(
    model="clarke-1.0",           # Required: hal-1.0, clarke-1.0, or tycho-1.0
    messages=[...],               # Required: [{role, content}, ...]
    max_tokens=512,               # Recommended: prevent runaway generation
    temperature=0.7,              # Optional: 0–2, default 1
    top_p=1.0,                    # Optional: 0–1, default 1
    stream=False,                 # Optional: set True for streaming
    tools=[...],                  # Optional: tool definitions for tool calling
    response_format={...},        # Optional: JSON mode or JSON schema
)
```

### Error codes and retry guidance

| Status | Type | Retryable | Action |
|---|---|---|---|
| 400 | `invalid_request_error` | No | Check request shape and parameters |
| 401 | `authentication_error` | No | Verify API key and header format |
| 404 | `not_found_error` | No | Check base URL and model string |
| 429 | `rate_limit_error` | Yes | Exponential backoff with jitter |
| 500 | `internal_error` | Yes | Retry with backoff, then contact support |

## Decision guidance

### When to use each model

| Scenario | Model | Reason |
|---|---|---|
| Complex reasoning, multi-step planning, code generation | Hal 1.0 | Highest quality, handles deep analysis |
| RAG, copilots, tool calling, most production workloads | Clarke 1.0 | Best balance of quality, latency, cost |
| High-volume classification, extraction, routing | Tycho 1.0 | Lowest cost, fastest latency |
| Unsure which to pick | Clarke 1.0 | Safe default for most tasks |

### When to use OpenAI vs Anthropic format

| Scenario | Format | Reason |
|---|---|---|
| Using OpenAI SDK or LangChain ChatOpenAI | OpenAI | Native compatibility, no translation needed |
| Using Anthropic SDK or PydanticAI | Anthropic | Native compatibility, no translation needed |
| Using LiteLLM or OpenRouter | Either | Both formats work; pick what your router expects |
| Migrating from OpenAI | OpenAI | Minimal code changes, same parameter names |
| Migrating from Anthropic | Anthropic | Minimal code changes, same parameter names |

### When to use streaming vs non-streaming

| Scenario | Approach | Reason |
|---|---|---|
| User-facing chat or copilot | Streaming | Reduces perceived latency, shows progress |
| Batch processing or background jobs | Non-streaming | Simpler error handling, full response at once |
| Tool calling agents | Non-streaming | Easier to parse tool_calls from complete response |
| Structured extraction | Non-streaming | Full JSON response needed for validation |

### When to use JSON mode vs JSON schema

| Scenario | Approach | Reason |
|---|---|---|
| Need valid JSON but flexible structure | JSON mode | Simpler, no schema overhead |
| Need strict field validation | JSON schema | Guarantees structure matches your code |
| Extracting to Pydantic models | JSON schema | Type safety, automatic parsing |
| High-volume extraction | JSON mode | Slightly faster, less overhead |

## Workflow

### 1. Set up authentication

```bash
export RADIUM_API_KEY="your-api-key-from-dashboard"
```

Create a key at https://deploy.radium.cloud → API Keys → Create key. Never commit keys to version control; use environment variables or a secrets manager.

### 2. Choose your SDK and base URL

**OpenAI SDK:**
```python
from openai import OpenAI
client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1"
)
```

**Anthropic SDK:**
```python
from anthropic import Anthropic
client = Anthropic(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud"
)
```

### 3. Pick a model

Start with `clarke-1.0` unless you have a specific reason to use Hal or Tycho. You can switch models later by changing the `model` parameter.

### 4. Send a request

```python
response = client.chat.completions.create(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "Your prompt here"}],
    max_tokens=512,
)
print(response.choices[0].message.content)
```

### 5. Handle errors and retry

```python
import time
import random

def with_retry(fn, max_retries=5):
    for attempt in range(max_retries):
        try:
            return fn()
        except Exception as e:
            status = getattr(e, "status_code", None)
            if status in (429, 503) and attempt < max_retries - 1:
                delay = (2 ** attempt) + random.random()
                time.sleep(delay)
            else:
                raise
```

### 6. For tool calling agents

Define tools as JSON schemas, send them in the `tools` parameter, execute the model's tool calls, and return results as `tool` messages:

```python
response = client.chat.completions.create(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "What's the weather?"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather",
            "parameters": {
                "type": "object",
                "properties": {"location": {"type": "string"}},
                "required": ["location"]
            }
        }
    }],
)

# Check if model called a tool
if response.choices[0].message.tool_calls:
    for tool_call in response.choices[0].message.tool_calls:
        # Execute tool, append result as tool message
        result = execute_tool(tool_call.function.name, tool_call.function.arguments)
        messages.append({"role": "tool", "tool_call_id": tool_call.id, "content": result})
    # Send results back to model
    response = client.chat.completions.create(model="clarke-1.0", messages=messages)
```

### 7. For structured extraction

Use JSON schema mode with Pydantic or raw JSON mode:

```python
from pydantic import BaseModel

class Person(BaseModel):
    name: str
    age: int

response = client.beta.chat.completions.parse(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "Extract person info from: ..."}],
    response_format=Person,
)
person = response.choices[0].message.parsed
```

## Common gotchas

- **Wrong base URL**: OpenAI format uses `/v1` suffix; Anthropic format does not. Double-check your base URL matches your SDK.
- **Model string typos**: Use exact strings `hal-1.0`, `clarke-1.0`, `tycho-1.0`. Using `gpt-*` or `claude-*` returns 404.
- **Missing max_tokens**: Set `max_tokens` on every request to prevent runaway generation. Recommended minimum is 512.
- **API key in client code**: Never hardcode or commit API keys. Use environment variables or a secrets manager.
- **Forgetting tool message format**: After executing a tool, append a `tool` message with `tool_call_id` and `content`. Missing this breaks the agent loop.
- **Streaming with tool calls**: Tool call streaming uses partial JSON deltas. Parse the final chunk with `finish_reason: "tool_calls"` to get the complete tool call.
- **Context window exceeded**: Clarke 1.0 has 1M tokens; Hal has 250K; Tycho has 125K. Radium returns 400 with `context_length_exceeded` if you exceed the limit. Handle this by truncating or splitting messages.
- **Parallel tool calls**: All models support parallel tool calls. Iterate over `message.tool_calls` and execute each one before sending results back.
- **JSON schema validation**: Complex nested schemas with recursive references may not be fully supported. Test your schema on a representative sample before production.
- **Anthropic format response reading**: When using the Messages API, read the last content block (`response.content[-1].text`), not the first, because reasoning blocks may precede text.

## Verification checklist

Before submitting work with Radium:

- [ ] API key is set in environment variables, not hardcoded
- [ ] Base URL matches your SDK format (OpenAI vs Anthropic)
- [ ] Model string is one of: `hal-1.0`, `clarke-1.0`, `tycho-1.0`
- [ ] `max_tokens` is set on every request (recommended: 512 or higher)
- [ ] Error handling includes retry logic for 429 and 503 responses
- [ ] Tool calling agents append tool results as `tool` messages with `tool_call_id`
- [ ] Structured extraction uses JSON schema or JSON mode with proper validation
- [ ] Streaming responses are parsed correctly (check `delta.content` for OpenAI, iterate chunks for Anthropic)
- [ ] Context window limits are respected (250K for Hal, 1M for Clarke, 125K for Tycho)
- [ ] Sensitive data (API keys, passwords, SSH keys) is not logged or sent in prompts

## Resources

- **Full documentation index**: https://radium-e5ab5fa4.mintlify.app/llms.txt — comprehensive page-by-page navigation for all topics
- **Quickstart**: https://radium-e5ab5fa4.mintlify.app/quickstart — send your first request in four steps
- **Models overview**: https://radium-e5ab5fa4.mintlify.app/models/overview — compare Hal, Clarke, and Tycho
- **Tool calling**: https://radium-e5ab5fa4.mintlify.app/core-concepts/tool-calling — implement agents with external functions
- **Frameworks**: https://radium-e5ab5fa4.mintlify.app/tools/overview — integrate with LangChain, LangGraph, Vercel AI SDK, and others

---

> For additional documentation and navigation, see: https://radium-e5ab5fa4.mintlify.app/llms.txt