> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming response

> Stream tokens from Radium in real time

# Streaming response

Streaming returns tokens as the model generates them, which makes apps feel faster and lets you show progress to users.

## Basic streaming

```python stream.py theme={null}
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

stream = client.chat.completions.create(
    model="clarke-1.0",
    messages=[{"role": "user", "content": "Write a haiku about debugging."}],
    stream=True,
)

for chunk in stream:
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)
print()
```

<Note>
  The reasoning-chunk field name and the <code>extra\_body</code> parameter below are inferred from Anthropic-compatible patterns and have not yet been verified against the live Radium API. Confirm with Engineering before publishing.
</Note>

## Streaming with reasoning

When `clarke-1.0` or `hal-1.0` produces a reasoning chain before the final answer, streaming events include a `reasoning_content` field:

```python stream_reasoning.py theme={null}
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

stream = client.chat.completions.create(
    model="hal-1.0",
    messages=[{"role": "user", "content": "Explain the halting problem."}],
    stream=True,
    extra_body={"reasoning": {"enabled": True}},
)

reasoning_buffer = ""
answer_buffer = ""

for chunk in stream:
    delta = chunk.choices[0].delta
    if getattr(delta, "reasoning_content", None):
        reasoning_buffer += delta.reasoning_content
        print(f"\r[Thinking... {len(reasoning_buffer)} chars]", end="", flush=True)
    if delta.content:
        answer_buffer += delta.content
        print(delta.content, end="", flush=True)

print(f"\n\nReasoning: {reasoning_buffer[:200]}...")
```

## Streaming in a web app

A minimal FastAPI endpoint that streams to a frontend:

```python api.py theme={null}
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import OpenAI
import os
import json

app = FastAPI()
client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

@app.post("/chat")
async def chat(message: str):
    stream = client.chat.completions.create(
        model="clarke-1.0",
        messages=[{"role": "user", "content": message}],
        stream=True,
    )

    def event_generator():
        for chunk in stream:
            content = chunk.choices[0].delta.content
            if content:
                yield f"data: {json.dumps({'text': content})}\n\n"
        yield "data: [DONE]\n\n"

    return StreamingResponse(
        event_generator(),
        media_type="text/event-stream",
    )
```

Run with:

```bash theme={null}
uvicorn api:app --reload
```

## Next steps

<CardGroup cols={2}>
  <Card title="Tool calling agent" icon="robot" href="/examples/tool-calling-agent">Build an agent that calls functions</Card>
  <Card title="RAG chatbot" icon="messages" href="/examples/rag-chatbot">Add documents to your chatbot</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.