Skip to main content

Multi-turn conversation

Stateless LLMs have no memory — you build it by sending the full message history with every request. This recipe shows how to manage a conversation thread, trim it when it gets too long, and summarize past turns to stay within context limits.

The script

chat_session.py

Run it

What it does

  1. Maintains a messages list that grows with every user/assistant exchange.
  2. Triggers summarization once the thread exceeds summary_trigger turns — asks the model to compress old context into a single system message.
  3. Hard-trims history to max_history turns if it still grows too large, preventing context-window overflow.
  4. Keeps recent context intact — summarization always preserves the last 2 turns so the model doesn’t lose the current thread.

Persisting across restarts

Save and load conversations with JSON so users don’t lose context when your app restarts:

Tips

  • Use clarke-1.0 for the main chat — its 1,000,000-token context window handles long threads well.
  • Keep temperature low (0.3) during summarization so summaries stay factual.
  • Track token count via response.usage.total_tokens if you want to summarize based on tokens rather than turns.
  • Store user IDs alongside saved sessions if you’re building a multi-user app.

Next steps

Structured extraction

Extract structured data from conversation turns

RAG chatbot

Ground conversations in your own documents