> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Embeddings & vector search

> Generate embeddings and search semantically across your documents

<Error>
  This recipe calls <code>client.embeddings.create()</code> through the Radium API, but the <a href="/switching/from-openai">OpenAI compatibility guide</a> states that <code>POST /v1/embeddings</code> is currently <strong>not supported</strong>. This entire example needs rewriting once embeddings become available, or it should be retargeted to a third-party embedding provider.
</Error>

# Embeddings & vector search

Turn text into dense vectors and search by meaning, not keywords. This recipe shows how to generate embeddings with Radium, store them locally, and perform cosine-similarity search.

## The script

```python embeddings_search.py theme={null}
import os
import json
import numpy as np
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

EMBEDDING_MODEL = "text-embedding-3-small"


def embed(texts: list[str]) -> list[list[float]]:
    """Generate embeddings for a list of texts."""
    response = client.embeddings.create(model=EMBEDDING_MODEL, input=texts)
    return [item.embedding for item in response.data]


def cosine_similarity(a: np.ndarray, b: np.ndarray) -> float:
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))


def search(query: str, documents: list[str], top_k: int = 3) -> list[tuple[str, float]]:
    """Return top-k documents most similar to the query."""
    doc_embeddings = embed(documents)
    query_embedding = embed([query])[0]
    q_vec = np.array(query_embedding)

    scores = []
    for doc, emb in zip(documents, doc_embeddings):
        sim = cosine_similarity(q_vec, np.array(emb))
        scores.append((doc, sim))

    scores.sort(key=lambda x: x[1], reverse=True)
    return scores[:top_k]


# --- Dataset ---
documents = [
    "Radium provides fast, reliable LLM inference with automatic failover.",
    "Open-source frameworks like LangChain simplify building LLM applications.",
    "Vector databases such as Pinecone and Weaviate power semantic search.",
    "Radium's model router intelligently selects the best model per request.",
    "Transformer models revolutionized natural language processing since 2017.",
    "Radium supports OpenAI- and Anthropic-compatible APIs out of the box.",
]

# --- Run it ---
query = "How does Radium handle model selection?"
results = search(query, documents, top_k=2)

print(f"Query: {query}\n")
for doc, score in results:
    print(f"Score: {score:.3f} — {doc}")
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python embeddings_search.py
```

## Sample output

```
Query: How does Radium handle model selection?

Score: 0.912 — Radium's model router intelligently selects the best model per request.
Score: 0.847 — Radium provides fast, reliable LLM inference with automatic failover.
```

## Persisting to a vector store

For production, store embeddings in a proper vector database instead of in-memory lists:

```python vector_store.py theme={null}
import faiss

dimension = 1536  # text-embedding-3-small
doc_embeddings = embed(documents)

index = faiss.IndexFlatIP(dimension)  # inner product = cosine if normalized
index.add(np.array(doc_embeddings).astype("float32"))

# Search
query_emb = np.array(embed([query])[0]).astype("float32")
scores, indices = index.search(query_emb.reshape(1, -1), k=3)
for score, idx in zip(scores[0], indices[0]):
    print(f"{score:.3f} — {documents[idx]}")
```

## Chunking long documents

Large documents exceed the embedding model's context limit. Split them into chunks first:

```python theme={null}
def chunk_text(text: str, chunk_size: int = 500, overlap: int = 50) -> list[str]:
    chunks = []
    start = 0
    while start < len(text):
        end = start + chunk_size
        chunks.append(text[start:end])
        start = end - overlap
    return chunks
```

## Tips

* **`text-embedding-3-small`** is the best cost/quality tradeoff for most use cases.
* **Always normalize embeddings** before cosine similarity if your vector DB expects it.
* **Chunk size 256-512 tokens** works well for retrieval — smaller chunks are more precise, larger ones retain more context.
* **Include metadata** with each chunk (source URL, page number) so search results are actionable.
* **Update embeddings incrementally** — only re-embed changed documents, not the whole corpus.

## Next steps

<CardGroup cols={2}>
  <Card title="RAG chatbot" icon="book-open" href="/examples/rag-chatbot">Combine embeddings with LLM generation for grounded Q\&A</Card>
  <Card title="Caching responses" icon="clock-rotate-left" href="/examples/caching-responses">Use semantic similarity for intelligent caching</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.