Viro API

Chat & Streaming

POST https://api.viro.app/v1/chat/completions — authenticated with Authorization: Bearer viro_sk_..., same request/response shape as OpenAI's Chat Completions API.

Basic request
curl https://api.viro.app/v1/chat/completions \
  -H "Authorization: Bearer $VIRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "viro/gpt-oss-120b",
    "messages": [{"role": "user", "content": "Explain photosynthesis."}]
  }'
from openai import OpenAI

client = OpenAI(api_key="viro_sk_live_...", base_url="https://api.viro.app/v1")

response = client.chat.completions.create(
    model="viro/gpt-oss-120b",
    messages=[{"role": "user", "content": "Explain photosynthesis."}],
)
print(response.choices[0].message.content)
print(response.viro)  # {"request_id": "...", "renewable_verified": true, ...}
Model routing

Pass viro/optimized, viro/frontier, or viro/clean instead of a specific model and Viro picks the underlying model per request. Same request and response shape — streaming, tool calling, and structured output all work normally.

response = client.chat.completions.create(
    model="viro/optimized",
    messages=[{"role": "user", "content": "Rewrite this sentence professionally: thanks for the email"}],
)

See Routing for how each router picks a model, the response headers that show what ran, and billing/fallback behavior.

Streaming

Set stream: true for server-sent events. The final chunk before data: [DONE] carries the viro provenance block instead of a content delta.

stream = client.chat.completions.create(
    model="viro/gpt-oss-120b",
    messages=[{"role": "user", "content": "Count to 5."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
Reasoning

Models that think before answering return that thinking on message.reasoning, separate from content. Streaming delivers it as delta.reasoning, which arrives before the content deltas — on a thinking model that is most of the wait, so rendering it is the difference between a progress indicator and a blank screen.

response = client.chat.completions.create(
    model="viro/gpt-oss-120b",
    messages=[{"role": "user", "content": "17 sheep, all but 9 run away. How many left?"}],
)
print(response.choices[0].message.reasoning)  # the chain of thought
print(response.choices[0].message.content)    # the answer

Reasoning is returned whenever the model produces it. To leave it out — it can be long, and you are charged for those tokens either way, since they are already counted in completion_tokens — send reasoning: {"exclude": true}. include_reasoning: false is accepted as a deprecated alias for the same thing.

response = client.chat.completions.create(
    model="viro/gpt-oss-120b",
    messages=[{"role": "user", "content": "17 sheep, all but 9 run away. How many left?"}],
    extra_body={"reasoning": {"exclude": True}},
)
Tool calling

Standard OpenAI function calling — pass tools and read message.tool_calls back. Works on every chat model, including Anthropic (translated to Claude's native tool-use format under the hood).

response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "What's the weather in Boston?"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a location",
            "parameters": {
                "type": "object",
                "properties": {"location": {"type": "string"}},
                "required": ["location"],
            },
        },
    }],
)
print(response.choices[0].message.tool_calls)
Vision (image & document input)

On vision-capable models, pass an array for content mixing text and image_url parts — standard OpenAI format. Useful for photo/chart/receipt reading, and for document analysis by sending each page as an image. image_url.url accepts either an http(s) URL or a data:image/...;base64,... data URI.

response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)
print(response.choices[0].message.content)

Base64 works the same way — swap the URL for a data URI:

import base64

with open("receipt.jpg", "rb") as f:
    b64 = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="openai/gpt-5.6-terra",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Extract the total and date from this receipt."},
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}},
        ],
    }],
)

Only some models accept image input — see the Models page's Vision column for the current list. Sending an image to a text-only model returns an upstream error rather than silently ignoring it.

PDF input

Send a PDF as a file part, the same way you send an image. It works on every chat model — but by one of two routes, and the difference is worth knowing because it decides what the model can actually see.

import base64

with open("contract.pdf", "rb") as f:
    data = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="viro/clean",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Summarize this contract and list any dates that need action."},
            {"type": "file", "file": {
                "filename": "contract.pdf",
                "file_data": f"data:application/pdf;base64,{data}",
            }},
        ],
    }],
)
print(response.choices[0].message.content)

Models that read the file directly — anthropic/*, openai/* and google/* vision models — get the PDF itself and look at the pages. Layout, tables, stamps, signatures and scanned pages all survive. The Models page's PDF column has the current list.

Every other model gets the document's text, extracted server-side and passed through as text. That is exact for any PDF generated digitally — from Word, LaTeX, a browser, an invoicing system — and it costs a fraction of the tokens that page images would, so an open-weight model on viro/clean can read a contract. What it does not carry is layout: column order and table structure are lost, and a scanned PDF has no text to extract at all. A scan returns a clear error naming the models that can read it rather than sending the model an empty document.

Documents are capped at 20 MB each. On the extracted path there is no page limit, and the text is capped at 400,000 characters (truncated with a visible marker, not dropped). Models that read the file directly have their own page ceilings — anthropic/* accepts at most 100 pages per document, and a longer one is refused before it is uploaded, with the page count in the error.file_data must be a base64 data URL — Viro does not fetch remote URLs on the chat path.

Structured outputs

Pass response_format to force JSON output on models that support it (OpenAI, open-weight viro/* models, xAI, Google). Anthropic models have no structured-output mode upstream, so response_format is silently ignored on anthropic/* — use tool calling instead if you need reliable JSON from Claude.

response = client.chat.completions.create(
    model="openai/gpt-5.6-terra",
    messages=[{"role": "user", "content": "Give me a JSON object with name and age for a fictional person."}],
    response_format={"type": "json_object"},
)
print(response.choices[0].message.content)
Parameters
FieldNotes
modelRequired. A model slug from GET /v1/models, e.g. viro/gpt-oss-120b or anthropic/claude-sonnet-5.
messagesRequired. Standard OpenAI messages array (system/user/assistant/tool/developer roles).
streamSet true for a server-sent events stream instead of a single JSON response.
max_tokensCaps output length. Also used to size the pre-flight credit reservation.
temperature, top_p, stop, seed, presence_penalty, frequency_penaltyPassed straight through to the upstream model.
tools, tool_choiceOpenAI-style function calling. See below.
response_formatStructured outputs (json_object / json_schema). Not supported on anthropic/* models — silently ignored.