Skip to main content

RAG chatbot

Retrieval-Augmented Generation (RAG) grounds your chatbot in documents it can search at query time. This example uses an in-memory vector store, but the pattern works with any retrieval backend.

Setup

The full RAG pipeline

rag.py

Expected output

Production considerations

Streaming RAG responses

Add stream=True to the chat completion and iterate over chunks to show the answer as it generates:

Next steps

Streaming response

Stream tokens in real time

Tool calling agent

Give your chatbot tools to act