Embeddings & vector search
Turn text into dense vectors and search by meaning, not keywords. This recipe shows how to generate embeddings with Radium, store them locally, and perform cosine-similarity search.The script
embeddings_search.py
Run it
Sample output
Persisting to a vector store
For production, store embeddings in a proper vector database instead of in-memory lists:vector_store.py
Chunking long documents
Large documents exceed the embedding model’s context limit. Split them into chunks first:Tips
text-embedding-3-smallis the best cost/quality tradeoff for most use cases.- Always normalize embeddings before cosine similarity if your vector DB expects it.
- Chunk size 256-512 tokens works well for retrieval — smaller chunks are more precise, larger ones retain more context.
- Include metadata with each chunk (source URL, page number) so search results are actionable.
- Update embeddings incrementally — only re-embed changed documents, not the whole corpus.
Next steps
RAG chatbot
Combine embeddings with LLM generation for grounded Q&A
Caching responses
Use semantic similarity for intelligent caching