> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Multi-turn conversation

> Maintain message history, manage context limits, and summarize long threads

# Multi-turn conversation

Stateless LLMs have no memory — you build it by sending the full message history with every request. This recipe shows how to manage a conversation thread, trim it when it gets too long, and summarize past turns to stay within context limits.

## The script

```python chat_session.py theme={null}
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


class Conversation:
    def __init__(self, model="clarke-1.0", max_history=10, summary_trigger=8):
        self.model = model
        self.messages = [
            {"role": "system", "content": "You are a helpful assistant."}
        ]
        self.max_history = max_history      # hard cap on turns kept
        self.summary_trigger = summary_trigger  # when to summarize old turns

    def ask(self, user_message: str) -> str:
        # 1. Add user message
        self.messages.append({"role": "user", "content": user_message})

        # 2. Summarize if history is getting long
        if len(self.messages) > self.summary_trigger + 1:
            self._summarize()

        # 3. Trim if still over limit
        self._trim()

        # 4. Send to model
        response = client.chat.completions.create(
            model=self.model,
            messages=self.messages,
            temperature=0.7,
            max_tokens=512,
        )

        assistant_reply = response.choices[0].message.content
        self.messages.append({"role": "assistant", "content": assistant_reply})
        return assistant_reply

    def _trim(self):
        """Keep only the system message + last `max_history` turns."""
        # +2 because system msg + pairs of [user, assistant]
        while len(self.messages) > self.max_history * 2 + 1:
            # Don't remove system message at index 0
            self.messages.pop(1)

    def _summarize(self):
        """Replace old user/assistant turns with a summary via the model."""
        # Everything except system message and the most recent 2 turns
        old_turns = self.messages[1:-4]
        if not old_turns:
            return

        summary_request = [
            {"role": "system", "content": "Summarize the following conversation briefly."},
            {"role": "user", "content": "\n".join(f"{m['role']}: {m['content']}" for m in old_turns)},
        ]
        summary_resp = client.chat.completions.create(
            model=self.model,
            messages=summary_request,
            temperature=0.3,
            max_tokens=256,
        )
        summary = summary_resp.choices[0].message.content

        # Replace old turns with a system-level summary
        self.messages = [
            self.messages[0],  # original system prompt
            {"role": "system", "content": f"Conversation summary so far: {summary}"},
        ] + self.messages[-4:]  # keep last 2 user/assistant pairs


# --- Run it ---
chat = Conversation(model="clarke-1.0")

for msg in [
    "What is machine learning?",
    "Can you give me a concrete example?",
    "Who invented the perceptron?",
    "When was that exactly?",
    "How does a neural network learn?",
    "What is backpropagation?",
    "Tell me about transformers.",
    "Who wrote the Attention Is All You Need paper?",
    "Summarize everything we discussed.",
]:
    print(f"\nUser: {msg}")
    reply = chat.ask(msg)
    print(f"Assistant: {reply}")

print(f"\nFinal conversation has {len(chat.messages)} messages.")
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python chat_session.py
```

## What it does

1. **Maintains a `messages` list** that grows with every user/assistant exchange.
2. **Triggers summarization** once the thread exceeds `summary_trigger` turns — asks the model to compress old context into a single system message.
3. **Hard-trims history** to `max_history` turns if it still grows too large, preventing context-window overflow.
4. **Keeps recent context intact** — summarization always preserves the last 2 turns so the model doesn't lose the current thread.

## Persisting across restarts

Save and load conversations with JSON so users don't lose context when your app restarts:

```python theme={null}
import json

# Save
with open("session.json", "w") as f:
    json.dump(chat.messages, f)

# Load
with open("session.json", "r") as f:
    chat.messages = json.load(f)
```

## Tips

* **Use `clarke-1.0`** for the main chat — its 1,000,000-token context window handles long threads well.
* **Keep `temperature` low (0.3)** during summarization so summaries stay factual.
* **Track token count** via `response.usage.total_tokens` if you want to summarize based on tokens rather than turns.
* **Store user IDs** alongside saved sessions if you're building a multi-user app.

## Next steps

<CardGroup cols={2}>
  <Card title="Structured extraction" icon="table" href="/examples/structured-extraction">Extract structured data from conversation turns</Card>
  <Card title="RAG chatbot" icon="book-open" href="/examples/rag-chatbot">Ground conversations in your own documents</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.