Skip to main content

Multimodal images

Radium supports vision-capable models. Pass images as base64-encoded data URLs or public URLs to get descriptions, OCR, visual classifications, and more.

The script

vision.py

Run it

Processing multiple images in parallel

batch_vision.py

Tips

  • Use clarke-1.0 for vision tasks — it has strong image understanding.
  • Compress images to < 1MB before base64 encoding to stay within payload limits (unverified limit).
  • Use "detail": "low" for simple classification (faster, fewer tokens) and "high"** for OCR or detailed analysis (OpenAI-specific parameters — verify Radium support).
  • Video frames — extract 1 frame per second, then batch classify with the example above.
  • URL vs base64 — URLs reduce payload size but add a network hop; base64 is more reliable for local files.

Next steps

Batch processing

Process large image datasets in parallel

Structured extraction

Extract structured data from image descriptions