Viro API

Model Routing

Three virtual models — viro/clean, viro/optimized, and viro/frontier — that pick a real model per request instead of you naming one. Same endpoint, same request and response shape as always; the only difference is what you put in model.

The three routers
RouterOptimizes for
viro/cleanRecommended default. Best available model restricted to providers verified on renewable energy. This is a hard constraint, not a preference — if no clean-verified model can serve the request, the API returns an error instead of silently falling back to unverified infrastructure. Reach for this one unless a request specifically needs a model outside that pool.
viro/optimizedThe most efficient model that still clears the quality bar the task demands, across every provider. A rewrite or a quick summary goes to a small, cheap model; a hard debugging or reasoning task escalates automatically to something stronger. Prefers clean-energy models on close calls, but will use non-clean infrastructure when a task genuinely needs it.
viro/frontierMaximum capability for the task, regardless of cost, speed, or energy source. Cost and latency are only tiebreakers. Use it when you want the strongest possible answer and price isn't the concern.
Using a router

Pass the router slug exactly where a model slug would go. Everything else about the request — streaming, tool calling, structured output, multimodal input — works normally.

from openai import OpenAI

client = OpenAI(api_key="viro_sk_live_...", base_url="https://api.viro.app/v1")

response = client.chat.completions.create(
    model="viro/clean",
    messages=[{"role": "user", "content": "Rewrite this sentence professionally: thanks for the email"}],
)
print(response.choices[0].message.content)

Naming a concrete model (anthropic/claude-sonnet-5, viro/gpt-oss-120b, etc.) bypasses routing entirely — that path is unchanged by any of this.

How selection works

Two stages, in order, both driven by data rather than hard-coded rules:

1. Eligibility. Every model that can't actually serve the request is removed outright — not scored lower, removed. If your request includes tools, only tool-capable models remain. Images restrict to vision-capable models. A long conversation restricts to models whose context window fits it. viro/clean restricts to renewable-verified providers at this same stage — which is what makes its guarantee absolute rather than a soft preference: a non-clean model is never in the running to begin with.

2. Scoring. What's left is ranked on expected quality for the specific task, cost, and efficiency, weighted differently per router. viro/optimized also applies a minimum quality threshold based on how demanding the request looks — this is what stops it from simply always picking the cheapest model. A trivial rewrite has a low bar and a small model clears it easily; a request that reads as a hard debugging or reasoning problem raises the bar high enough that only strong models clear it, so the router escalates automatically.

Task difficulty is assessed from the request itself — length, code content, mathematical notation, explicit difficulty language, tool use, conversation length — with no extra model call in the loop. Routing overhead is well under a millisecond.

Seeing what was picked

The response body stays strictly OpenAI-compatible — routing metadata never changes its shape. Which model actually ran comes back in response headers instead:

x-viro-router: clean
x-viro-resolved-model: viro/gpt-oss-120b
x-viro-provider: nscale
x-viro-clean-energy: true

You can also try each router interactively in the Playground — it shows the resolved model above the response.

Billing and reliability

You're billed for the model that ran, at its normal rate. Routing itself costs nothing extra. Usage on the Usage page is attributed to the resolved model, not the router slug, so per-model spend reporting means exactly what it always has.

Fallback. If the router's top choice fails for an infrastructure reason (provider error, timeout, rate limit), the request automatically retries the next-ranked eligible model rather than failing outright — all still within the same router's constraints, so a viro/clean request can never fall back to non-clean infrastructure. If every eligible candidate fails, the request returns an error rather than retrying forever.

Current limitations

Per-task quality scores are informed estimates based on published model positioning, not benchmarks Viro has run itself. They determine routing quality and will be refined over time — the ordering between models is meaningful; the exact numbers are provisional.

viro/clean currently routes only to Nscale- and TensorX-hosted models — the open-weight viro/* catalog. This is a real eligibility constraint enforced before scoring, not a soft preference, and it will expand as more providers are verified.