Viro API

Pricing

Transparent, per-token billing. Only pay for what you use — no hidden fees, no monthly minimums.

Per-model rates, context windows, and modalities live on the Models page. This page covers how routing and costs actually work.

Model Routing

Pass one of these three instead of a specific model and Viro picks the best underlying model for each request — same request/response shape, no code changes beyond the model field. A router never has its own price: it bills at whichever model it resolves to, at that model's rate on the Models page.

Recommended
viro/clean
Renewable-verified infrastructure only

Best model restricted to providers verified on renewable energy. Never falls back to unverified infrastructure.

viro/optimized
Cheapest model that still gets it right

Most efficient model that still clears the quality bar for the task.

viro/frontier
Maximum intelligence, cost is secondary

Strongest available model for the task, regardless of cost or energy source.

Server tools

web_search

Brave Search API — ranked results, snippets, time-aware filters

$0.01/search

per call, any number of results

web_fetch

Retrieve readable text from public web pages — SSRF-protected, truncated at 120k chars

Free

only tokens

datetime

Current date and time in any IANA timezone — pure computation, no outbound call

Free

only tokens

Important

  • • Failed tool calls (vendor errors, rate limits, bad credentials) are not billed
  • • A search that legitimately finds no results is billable — the vendor ran it
  • • Tool spend is tracked separately from token spend in usage records
  • • Up to 5 upstream rounds and 10 tool calls per request
How billing works

Token costs

Input tokens are charged at the model's input rate per 1M tokens. Output tokens at the output rate. Both summed across all rounds of a multi-turn tool loop into a single usage event.

Prompt caching

Where the upstream provider caches prompts, Viro passes the discount straight through — nothing to enable, and no cache_control markers to add to your request. Repeated input that the provider serves from its cache is billed at the model's cache-read rate instead of its full input rate: roughly a tenth of the input rate on openai/* models, roughly a quarter on the TensorX-hosted models (viro/deepseek-v4-*, viro/glm-5.2, viro/minimax-m3). OpenAI models also charge a one-time cache-write premium on the request that populates the cache.

Caching does not apply everywhere. Nscale-hosted open-weight models — most of the viro/* catalog — and anthropic/*, google/* and xai/* models have no prompt cache through Viro today, so every input token bills at the normal input rate no matter how much of the prompt repeats. If you're building an agent around a long fixed system prompt, that difference is worth choosing a model on. Cached and fresh token counts are broken out per request under Usage.

Tool costs

web_search: $0.01 per search call, however many results come back (only if the vendor runs it and returns results or legitimately finds nothing). web_fetch and datetime are free — you only pay for the tokens their output becomes.

Reservations

Requests hold an estimated cost from your account until they complete (covers streaming disconnects). The reservation is reconciled to actual usage on completion. Stale reservations are swept after 5 minutes.

Rate limiting

Per-key RPM limits are configurable. Hitting a limit returns 429 with a Retry-After header. Per-account spend limits can be set in the console.

Example costs

Summarize a web page with viro/gpt-oss-120b

  • • web_fetch: free
  • • ~500 input tokens (prompt + page): $0.0005
  • • ~150 output tokens (summary): $0.0002
  • • Total: ~$0.0007

Search + fetch + synthesize with anthropic/claude-sonnet-5

  • • web_search (1): $0.01
  • • web_fetch (2 pages): free
  • • ~2000 input tokens (prompt + search + fetched pages): $0.003
  • • ~300 output tokens (synthesis): $0.0015
  • • Total: ~$0.0145