Multi-turn conversation
Stateless LLMs have no memory — you build it by sending the full message history with every request. This recipe shows how to manage a conversation thread, trim it when it gets too long, and summarize past turns to stay within context limits.The script
chat_session.py
Run it
What it does
- Maintains a
messageslist that grows with every user/assistant exchange. - Triggers summarization once the thread exceeds
summary_triggerturns — asks the model to compress old context into a single system message. - Hard-trims history to
max_historyturns if it still grows too large, preventing context-window overflow. - Keeps recent context intact — summarization always preserves the last 2 turns so the model doesn’t lose the current thread.
Persisting across restarts
Save and load conversations with JSON so users don’t lose context when your app restarts:Tips
- Use
clarke-1.0for the main chat — its 1,000,000-token context window handles long threads well. - Keep
temperaturelow (0.3) during summarization so summaries stay factual. - Track token count via
response.usage.total_tokensif you want to summarize based on tokens rather than turns. - Store user IDs alongside saved sessions if you’re building a multi-user app.
Next steps
Structured extraction
Extract structured data from conversation turns
RAG chatbot
Ground conversations in your own documents