> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# RAG chatbot

> Build a retrieval-augmented generation chatbot with Radium

<Error>
  This recipe calls <code>client.embeddings.create()</code> through the Radium API, but the <a href="/switching/from-openai">OpenAI compatibility guide</a> states that <code>POST /v1/embeddings</code> is currently <strong>not supported</strong>. This entire example needs rewriting once embeddings become available, or it should be retargeted to a third-party embedding provider.
</Error>

# RAG chatbot

Retrieval-Augmented Generation (RAG) grounds your chatbot in documents it can search at query time. This example uses an in-memory vector store, but the pattern works with any retrieval backend.

## Setup

```bash theme={null}
pip install openai numpy
```

## The full RAG pipeline

```python rag.py theme={null}
import os
import numpy as np
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

# 1. Your knowledge base
documents = [
    "Radium serves frontier-class models through OpenAI- and Anthropic-compatible endpoints.",
    "Clarke is the default model for RAG, copilots, and production workloads.",
    "Hal is the reasoning model for complex multi-step agents and code generation.",
    "Tycho is the fast, cost-effective model for classification and extraction.",
    "All Radium models support streaming, tool calling, and structured outputs.",
]


# 2. Simple in-memory vector store
class VectorStore:
    def __init__(self):
        self.documents = []
        self.embeddings = []

    def add(self, text: str):
        response = client.embeddings.create(
            model="clarke-1.0",  # or your preferred embedding model
            input=text,
        )
        embedding = np.array(response.data[0].embedding)
        self.documents.append(text)
        self.embeddings.append(embedding)

    def search(self, query: str, top_k: int = 2) -> list[str]:
        response = client.embeddings.create(
            model="clarke-1.0",
            input=query,
        )
        query_embedding = np.array(response.data[0].embedding)

        # Cosine similarity
        embeddings_matrix = np.stack(self.embeddings)
        similarities = embeddings_matrix @ query_embedding
        top_indices = np.argsort(similarities)[-top_k:][::-1]

        return [self.documents[i] for i in top_indices]


# 3. Build the index
store = VectorStore()
for doc in documents:
    store.add(doc)

print(f"Indexed {len(documents)} documents.\n")


# 4. Chat with retrieval
def chat(query: str) -> str:
    # Retrieve relevant context
    context = store.search(query, top_k=2)
    context_block = "\n\n".join(f"- {d}" for d in context)

    # Build the prompt
    messages = [
        {
            "role": "system",
            "content": (
                "You are a helpful assistant. Use the provided context to answer. "
                "If the context does not contain the answer, say so."
            ),
        },
        {
            "role": "user",
            "content": f"Context:\n{context_block}\n\nQuestion: {query}",
        },
    ]

    response = client.chat.completions.create(
        model="clarke-1.0",
        messages=messages,
        temperature=0.3,
        max_tokens=512,
    )

    return response.choices[0].message.content


# 5. Ask questions
if __name__ == "__main__":
    questions = [
        "Which model should I use for a coding assistant?",
        "What features do all Radium models support?",
        "Does Radium support image generation?",
    ]

    for q in questions:
        print(f"Q: {q}")
        print(f"A: {chat(q)}\n")
```

## Expected output

```
Indexed 5 documents.

Q: Which model should I use for a coding assistant?
A: Hal is the reasoning model for complex multi-step agents and code generation, making it the best choice for a coding assistant.

Q: What features do all Radium models support?
A: All Radium models support streaming, tool calling, and structured outputs.

Q: Does Radium support image generation?
A: The provided context does not mention image generation.
```

## Production considerations

| Component | In-memory (demo) | Production |
| - | - | - |
| Vector store | `numpy` arrays | Pinecone, Weaviate, Qdrant, pgvector |
| Embeddings | Radium API | Dedicated embedding model or service |
| Document ingestion | Static list | Crawler + chunking pipeline |
| Re-ranking | None | Cross-encoder re-ranker |
| Caching | None | Redis for embedding cache |

## Streaming RAG responses

Add `stream=True` to the chat completion and iterate over chunks to show the answer as it generates:

```python theme={null}
stream = client.chat.completions.create(
    model="clarke-1.0",
    messages=messages,
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
```

## Next steps

<CardGroup cols={2}>
  <Card title="Streaming response" icon="bolt" href="/examples/streaming-response">Stream tokens in real time</Card>
  <Card title="Tool calling agent" icon="robot" href="/examples/tool-calling-agent">Give your chatbot tools to act</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.