> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Caching responses

> Cache LLM responses to reduce costs and latency on repeated prompts

# Caching responses

Many prompts repeat — FAQs, classification labels, standard summaries. Caching identical or similar requests can cut costs by 30-70% and eliminate latency for cache hits.

## The script

```python cache.py theme={null}
import os
import hashlib
import json
import time
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

CACHE_DIR = ".radium_cache"
os.makedirs(CACHE_DIR, exist_ok=True)


def _cache_key(messages: list, model: str, temperature: float, max_tokens: int) -> str:
    """Deterministic hash of request parameters."""
    payload = json.dumps({
        "model": model,
        "messages": messages,
        "temperature": temperature,
        "max_tokens": max_tokens,
    }, sort_keys=True)
    return hashlib.sha256(payload.encode()).hexdigest()


def cached_chat(
    messages,
    model="clarke-1.0",
    temperature=0.7,
    max_tokens=512,
    use_cache: bool = True,
) -> dict:
    """
    Chat completion with local file caching.
    Returns {"content": str, "cached": bool, "latency_ms": float}
    """
    key = _cache_key(messages, model, temperature, max_tokens)
    cache_path = os.path.join(CACHE_DIR, f"{key}.json")

    # 1. Check cache
    if use_cache and os.path.exists(cache_path):
        with open(cache_path) as f:
            cached = json.load(f)
        return {"content": cached["content"], "cached": True, "latency_ms": 0.0}

    # 2. Hit the API
    start = time.perf_counter()
    response = client.chat.completions.create(
        model=model,
        messages=messages,
        temperature=temperature,
        max_tokens=max_tokens,
    )
    latency = (time.perf_counter() - start) * 1000
    content = response.choices[0].message.content

    # 3. Write to cache
    with open(cache_path, "w") as f:
        json.dump({"content": content, "timestamp": time.time()}, f)

    return {"content": content, "cached": False, "latency_ms": latency}


# --- Run it ---
for i in range(3):
    result = cached_chat(
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": "What is the capital of France?"},
        ]
    )
    print(f"Run {i+1}: cached={result['cached']}, latency={result['latency_ms']:.1f}ms")
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python cache.py
```

## Sample output

```
Run 1: cached=False, latency=412.3ms
Run 2: cached=True, latency=0.0ms
Run 3: cached=True, latency=0.0ms
```

<Warning>
  The semantic-caching example below calls <code>client.embeddings.create()</code> through the Radium API. The <a href="/switching/from-openai">OpenAI compatibility guide</a> states that <code>POST /v1/embeddings</code> is currently <strong>not supported</strong>. Retarget to a third-party embedding provider before using this pattern in production.
</Warning>

## Semantic caching with embeddings

Exact-match caching misses when phrasing changes. Use embeddings to find similar past queries:

```python semantic_cache.py theme={null}
import numpy as np
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

embedding_cache = {}  # query_text -> (embedding, answer)


def get_embedding(text: str) -> list[float]:
    resp = client.embeddings.create(model="text-embedding-3-small", input=text)
    return resp.data[0].embedding


def cosine_similarity(a: list[float], b: list[float]) -> float:
    a, b = np.array(a), np.array(b)
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))


def semantic_cached_chat(user_query: str, threshold: float = 0.92) -> dict:
    query_emb = get_embedding(user_query)

    # Search for similar cached queries
    for cached_text, (cached_emb, answer) in embedding_cache.items():
        if cosine_similarity(query_emb, cached_emb) >= threshold:
            return {"content": answer, "cached": True, "matched": cached_text}

    # No match — call the LLM
    response = client.chat.completions.create(
        model="clarke-1.0",
        messages=[{"role": "user", "content": user_query}],
        max_tokens=512,
    )
    answer = response.choices[0].message.content
    embedding_cache[user_query] = (query_emb, answer)
    return {"content": answer, "cached": False, "matched": None}
```

## Cache invalidation strategies

| Strategy | When to use |
| - | - |
| TTL (time-to-live) | `timestamp + 86400` — cache for 24h |
| Version stamp | Append a `prompt_version` to the cache key |
| Manual flush | Delete `.radium_cache/` when deploying new prompts |
| Model-aware | Include model name in key so `clarke-1.0` ≠ `tycho-1.0` |

## Tips

* **Use exact-match caching** for deterministic tasks like classification, extraction, and FAQ answering.
* **Use semantic caching** for open-ended Q\&A where users rephrase the same question.
* **Never cache** requests with `temperature > 0` unless you want identical outputs — stochasticity defeats the cache.
* **Monitor hit rate** — a 50%+ hit rate usually means caching is worth the complexity.

## Next steps

<CardGroup cols={2}>
  <Card title="Monitoring usage" icon="chart-line" href="/examples/monitoring-usage">Track cache hit rates and cost savings</Card>
  <Card title="Batch processing" icon="layer-group" href="/examples/batch-processing">Process large datasets with caching</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.