Skip to main content

Streaming response

Streaming returns tokens as the model generates them, which makes apps feel faster and lets you show progress to users.

Basic streaming

stream.py
The reasoning-chunk field name and the extra_body parameter below are inferred from Anthropic-compatible patterns and have not yet been verified against the live Radium API. Confirm with Engineering before publishing.

Streaming with reasoning

When clarke-1.0 or hal-1.0 produces a reasoning chain before the final answer, streaming events include a reasoning_content field:
stream_reasoning.py

Streaming in a web app

A minimal FastAPI endpoint that streams to a frontend:
api.py
Run with:

Next steps

Tool calling agent

Build an agent that calls functions

RAG chatbot

Add documents to your chatbot