Pricing
Transparent, per-token billing. Only pay for what you use — no hidden fees, no monthly minimums.
Per-model rates, context windows, and modalities live on the Models page. This page covers how routing and costs actually work.
Pass one of these three instead of a specific model and Viro picks the best underlying model for each request — same request/response shape, no code changes beyond the model field. A router never has its own price: it bills at whichever model it resolves to, at that model's rate on the Models page.
Best model restricted to providers verified on renewable energy. Never falls back to unverified infrastructure.
Most efficient model that still clears the quality bar for the task.
Strongest available model for the task, regardless of cost or energy source.
web_search
Brave Search API — ranked results, snippets, time-aware filters
$0.01/search
per call, any number of results
web_fetch
Retrieve readable text from public web pages — SSRF-protected, truncated at 120k chars
Free
only tokens
datetime
Current date and time in any IANA timezone — pure computation, no outbound call
Free
only tokens
Important
- • Failed tool calls (vendor errors, rate limits, bad credentials) are not billed
- • A search that legitimately finds no results is billable — the vendor ran it
- • Tool spend is tracked separately from token spend in usage records
- • Up to 5 upstream rounds and 10 tool calls per request
Token costs
Input tokens are charged at the model's input rate per 1M tokens. Output tokens at the output rate. Both summed across all rounds of a multi-turn tool loop into a single usage event.
Prompt caching
Where the upstream provider caches prompts, Viro passes the discount straight through — nothing to enable, and no cache_control markers to add to your request. Repeated input that the provider serves from its cache is billed at the model's cache-read rate instead of its full input rate: roughly a tenth of the input rate on openai/* models, roughly a quarter on the TensorX-hosted models (viro/deepseek-v4-*, viro/glm-5.2, viro/minimax-m3). OpenAI models also charge a one-time cache-write premium on the request that populates the cache.
Caching does not apply everywhere. Nscale-hosted open-weight models — most of the viro/* catalog — and anthropic/*, google/* and xai/* models have no prompt cache through Viro today, so every input token bills at the normal input rate no matter how much of the prompt repeats. If you're building an agent around a long fixed system prompt, that difference is worth choosing a model on. Cached and fresh token counts are broken out per request under Usage.
Tool costs
web_search: $0.01 per search call, however many results come back (only if the vendor runs it and returns results or legitimately finds nothing). web_fetch and datetime are free — you only pay for the tokens their output becomes.
Reservations
Requests hold an estimated cost from your account until they complete (covers streaming disconnects). The reservation is reconciled to actual usage on completion. Stale reservations are swept after 5 minutes.
Rate limiting
Per-key RPM limits are configurable. Hitting a limit returns 429 with a Retry-After header. Per-account spend limits can be set in the console.
Summarize a web page with viro/gpt-oss-120b
- • web_fetch: free
- • ~500 input tokens (prompt + page): $0.0005
- • ~150 output tokens (summary): $0.0002
- • Total: ~$0.0007
Search + fetch + synthesize with anthropic/claude-sonnet-5
- • web_search (1): $0.01
- • web_fetch (2 pages): free
- • ~2000 input tokens (prompt + search + fetched pages): $0.003
- • ~300 output tokens (synthesis): $0.0015
- • Total: ~$0.0145