> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Async webhooks

> Process LLM requests asynchronously with webhooks for long-running jobs

# Async webhooks

Some LLM tasks take too long for a synchronous HTTP response. This recipe shows how to accept a request, enqueue the work, call Radium in the background, and deliver results via a webhook callback.

## The script (FastAPI + background tasks)

```python webhook_server.py theme={null}
import os
import asyncio
import httpx
from fastapi import FastAPI, BackgroundTasks
from pydantic import BaseModel, HttpUrl
from openai import AsyncOpenAI

client = AsyncOpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)

app = FastAPI()


class SummarizeRequest(BaseModel):
    text: str
    webhook_url: HttpUrl
    request_id: str


async def _process_and_notify(req: SummarizeRequest):
    """Call Radium and POST the result to the webhook URL."""
    try:
        response = await client.chat.completions.create(
            model="clarke-1.0",
            messages=[
                {
                    "role": "system",
                    "content": "Summarize the following text in 2-3 sentences.",
                },
                {"role": "user", "content": req.text},
            ],
            max_tokens=256,
        )
        summary = response.choices[0].message.content
        status = "success"
    except Exception as e:
        summary = ""
        status = "error"

    payload = {
        "request_id": req.request_id,
        "status": status,
        "summary": summary,
    }

    async with httpx.AsyncClient() as http:
        await http.post(str(req.webhook_url), json=payload, timeout=30)


@app.post("/summarize")
async def summarize(req: SummarizeRequest, background_tasks: BackgroundTasks):
    """Enqueue summarization and return immediately."""
    background_tasks.add_task(_process_and_notify, req)
    return {"request_id": req.request_id, "status": "enqueued"}
```

## Run it

```bash theme={null}
pip install fastapi uvicorn httpx
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
uvicorn webhook_server:app --port 8000
```

## Submit a job

```bash theme={null}
curl -X POST http://localhost:8000/summarize \
  -H "Content-Type: application/json" \
  -d '{
    "text": "The transformer architecture, introduced in Attention Is All You Need (2017), revolutionized NLP by replacing recurrent layers with self-attention mechanisms...",
    "webhook_url": "https://your-app.com/webhooks/radium",
    "request_id": "job-123"
  }'
```

## Receive the webhook

Your endpoint will receive:

```json theme={null}
{
  "request_id": "job-123",
  "status": "success",
  "summary": "The transformer architecture replaced recurrent layers with self-attention, enabling parallel training and state-of-the-art NLP results."
}
```

## Production queue (Celery + Redis)

For high throughput, replace FastAPI background tasks with a proper queue:

```python celery_worker.py theme={null}
from celery import Celery
from openai import AsyncOpenAI
import os

app = Celery("radium", broker="redis://localhost:6379/0")
client = AsyncOpenAI(
    api_key=os.environ["RADIUM_API_KEY"],
    base_url="https://api.radium.cloud/v1",
)


@app.task
def summarize_task(text: str, webhook_url: str, request_id: str):
    import asyncio
    asyncio.run(_do_summarize(text, webhook_url, request_id))


async def _do_summarize(text, webhook_url, request_id):
    response = await client.chat.completions.create(
        model="clarke-1.0",
        messages=[{"role": "user", "content": f"Summarize:\n\n{text}"}],
        max_tokens=256,
    )
    summary = response.choices[0].message.content
    # POST to webhook_url...
```

## Polling fallback

Not every client supports webhooks. Offer a status endpoint:

```python theme={null}
jobs = {}

@app.post("/summarize")
async def summarize_poll(req: SummarizeRequest, background_tasks: BackgroundTasks):
    jobs[req.request_id] = {"status": "pending", "result": None}
    background_tasks.add_task(_process_and_store, req)
    return {"request_id": req.request_id, "status": "pending"}


@app.get("/status/{request_id}")
async def status(request_id: str):
    return jobs.get(request_id, {"status": "not_found"})
```

## Tips

* **Return `request_id` immediately** — never block the HTTP connection waiting for an LLM response.
* **Retry webhooks** with exponential backoff — your user's endpoint might be temporarily down.
* **Sign webhooks** with an HMAC signature so clients can verify the payload came from you.
* **Use `tycho-1.0`** for fast, cheap async tasks like classification and tagging.
* **Monitor queue depth** — if jobs pile up, scale workers or add a circuit breaker.

## Next steps

<CardGroup cols={2}>
  <Card title="Error handling & retries" icon="shield-halved" href="/examples/error-handling-retries">Add resilience to async workers</Card>
  <Card title="Monitoring usage" icon="chart-line" href="/examples/monitoring-usage">Track async job metrics and costs</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.