> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal images

> Send images to Radium models and describe, classify, or extract text from them

# Multimodal images

Radium supports vision-capable models. Pass images as base64-encoded data URLs or public URLs to get descriptions, OCR, visual classifications, and more.

## The script

```python vision.py theme={null}
import os
import base64
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


def encode_image(path: str) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")


def describe_image(image_path: str) -> str:
    b64 = encode_image(image_path)
    response = client.chat.completions.create(
        model="clarke-1.0",
        messages=[
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": "Describe this image in detail."},
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{b64}",
                            "detail": "high",
                        },
                    },
                ],
            }
        ],
        max_tokens=1024,
    )
    return response.choices[0].message.content


def classify_image(image_path: str, labels: list[str]) -> str:
    b64 = encode_image(image_path)
    label_str = ", ".join(labels)
    response = client.chat.completions.create(
        model="clarke-1.0",
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": f"Classify this image into one of these categories: {label_str}. Respond with only the category name.",
                    },
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{b64}",
                        },
                    },
                ],
            }
        ],
        temperature=0.0,
        max_tokens=50,
    )
    return response.choices[0].message.content.strip()


def classify_from_url(image_url: str, labels: list[str]) -> str:
    """Use a public URL instead of base64 encoding."""
    label_str = ", ".join(labels)
    response = client.chat.completions.create(
        model="clarke-1.0",
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": f"Classify this image: {label_str}. Output only the category.",
                    },
                    {
                        "type": "image_url",
                        "image_url": {"url": image_url},
                    },
                ],
            }
        ],
        temperature=0.0,
        max_tokens=50,
    )
    return response.choices[0].message.content.strip()


# --- Run it ---
if os.path.exists("photo.jpg"):
    print("Description:")
    print(describe_image("photo.jpg"))

    print("\nClassification:")
    print(classify_image("photo.jpg", ["indoor", "outdoor", "aerial"]))
```

## Run it

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
python vision.py
```

## Processing multiple images in parallel

```python batch_vision.py theme={null}
import asyncio
from openai import AsyncOpenAI
import os

client = AsyncOpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

SEMAPHORE = asyncio.Semaphore(10)


async def classify_one(image_path: str) -> dict:
    import base64
    with open(image_path, "rb") as f:
        b64 = base64.b64encode(f.read()).decode()

    async with SEMAPHORE:
        response = await client.chat.completions.create(
            model="clarke-1.0",
            messages=[
                {
                    "role": "user",
                    "content": [
                        {"type": "text", "text": "Is this image safe for work? Answer only Yes or No."},
                        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
                    ],
                }
            ],
            temperature=0.0,
            max_tokens=10,
        )
    return {"file": image_path, "answer": response.choices[0].message.content.strip()}


async def main():
    images = [f"img_{i}.png" for i in range(1, 11)]
    results = await asyncio.gather(*[classify_one(p) for p in images])
    for r in results:
        print(f"{r['file']}: {r['answer']}")


if __name__ == "__main__":
    asyncio.run(main())
```

## Tips

* **Use `clarke-1.0`** for vision tasks — it has strong image understanding.
* **Compress images** to <span style="color:red">\< 1MB</span> before base64 encoding to stay within payload limits (unverified limit).
* **Use <span style="color:red">`"detail": "low"`</span>** for simple classification (faster, fewer tokens) and <span style="color:red">`"high"`</span>\*\* for OCR or detailed analysis (OpenAI-specific parameters — verify Radium support).
* **Video frames** — extract 1 frame per second, then batch classify with the example above.
* **URL vs base64** — URLs reduce payload size but add a network hop; base64 is more reliable for local files.

## Next steps

<CardGroup cols={2}>
  <Card title="Batch processing" icon="layer-group" href="/examples/batch-processing">Process large image datasets in parallel</Card>
  <Card title="Structured extraction" icon="table" href="/examples/structured-extraction">Extract structured data from image descriptions</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.