> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Structured extraction

> Extract typed JSON data from unstructured text

# Structured extraction

Use structured outputs to extract typed data from unstructured text. Instead of parsing free-form prose, you get a validated JSON object that maps directly to your application code.

<Note>
  <code>client.beta.chat.completions.parse()</code> is a beta OpenAI SDK method that enables Pydantic-native structured output. Confirm with Engineering that Radium supports the underlying <code>response\_format</code> schema-constraint behaviour used by this method.
</Note>

## With Pydantic and OpenAI SDK

```python extract.py theme={null}
from openai import OpenAI
from pydantic import BaseModel, Field
import os

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


# 1. Define the schema you want extracted
class Person(BaseModel):
    name: str = Field(description="Full name of the person")
    age: int = Field(description="Age in years")
    occupation: str = Field(description="Job or profession")
    skills: list[str] = Field(description="List of professional skills")


# 2. Pass it as response_format
response = client.beta.chat.completions.parse(
    model="clarke-1.0",
    messages=[
        {
            "role": "user",
            "content": (
                "Extract person info from this bio:\n\n"
                "Ada Lovelace, born 1815, was an English mathematician and writer. "
                "She is often regarded as the first computer programmer because she "
                "wrote the first algorithm intended for Charles Babbage's Analytical Engine. "
                "She was skilled in mathematics, logic, and poetry."
            ),
        }
    ],
    response_format=Person,
)

# 3. Access the parsed object directly
person = response.choices[0].message.parsed
print(person.model_dump_json(indent=2))
```

## Output

```json theme={null}
{
  "name": "Ada Lovelace",
  "age": 209,
  "occupation": "Mathematician and writer",
  "skills": ["mathematics", "logic", "poetry", "programming"]
}
```

## With JSON mode (no Pydantic)

If you prefer not to use the beta parser, use `response_format={"type": "json_object"}`:

```python extract_json.py theme={null}
import json
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

response = client.chat.completions.create(
    model="clarke-1.0",
    messages=[
        {
            "role": "system",
            "content": "You extract person data from text. Respond with valid JSON only.",
        },
        {
            "role": "user",
            "content": "Dr. Alan Turing (1912–1954) was a mathematician, logician, and cryptanalyst. He broke the Enigma code and founded computer science and artificial intelligence.",
        },
    ],
    response_format={"type": "json_object"},
    max_tokens=512,
)

raw = response.choices[0].message.content
print(json.dumps(json.loads(raw), indent=2))
```

## Extracting from a batch of documents

Process many documents and collect structured results:

```python batch_extract.py theme={null}
from openai import OpenAI
from pydantic import BaseModel, Field
import os

client = OpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


class Company(BaseModel):
    name: str
    industry: str
    founded_year: int = Field(alias="foundedYear")
    headquarters: str
    employees: int | None = None


documents = [
    "OpenAI was founded in 2015 by Sam Altman and others. Based in San Francisco, it builds AGI with ~1,500 employees.",
    "Anthropic, founded in 2021 by Dario and Daniela Amodei, is an AI safety company headquartered in San Francisco.",
]

results = []
for doc in documents:
    response = client.beta.chat.completions.parse(
        model="clarke-1.0",
        messages=[{"role": "user", "content": f"Extract company info:\n\n{doc}"}],
        response_format=Company,
    )
    results.append(response.choices[0].message.parsed)

for c in results:
    print(f"{c.name} ({c.founded_year}) — {c.industry}")
```

## Tips

* Use **`clarke-1.0`** for extraction — it balances instruction-following with speed and cost.
* Use **`tycho-1.0`** for high-volume, simple extraction tasks like classification or entity tagging.
* Include a **system message** that instructs the model to respond with JSON only.
* Set **`max_tokens`** conservatively to avoid runaway generations.

## Next steps

<CardGroup cols={2}>
  <Card title="Batch processing" icon="layer-group" href="/examples/batch-processing">Process thousands of documents in parallel</Card>
  <Card title="Tool calling agent" icon="robot" href="/examples/tool-calling-agent">Use extraction inside an agent loop</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.