Server tools
Tools that Viro executes during your request, rather than handing back for you to run. Declare one and the model can use it mid-completion — no tool-calling loop in your code, and it works the same on every model, including open-weight viro/* models that have no built-in web access.
Retrieves a public web page and hands its readable text to the model. Free — you only pay for the tokens the fetched content adds.
from openai import OpenAI
client = OpenAI(api_key="viro_sk_live_...", base_url="https://api.viro.app/v1")
response = client.chat.completions.create(
model="viro/gpt-oss-120b",
messages=[{"role": "user", "content": "Summarize https://example.com"}],
tools=[{"type": "viro:web_fetch"}],
)
print(response.choices[0].message.content)The model may fetch several pages before answering. You get one final response — the round trips happen server-side.
Searches the web and hands ranked results — title, URL, snippet — to the model. It can search more than once to refine or compare before answering. 1¢ per search, on top of token cost.
response = client.chat.completions.create(
model="viro/gpt-oss-120b",
messages=[{"role": "user", "content": "What happened in AI policy this week?"}],
tools=[{"type": "viro:web_search"}],
)Results are snippets, not full pages — that keeps searches cheap. Enable viro:web_fetch alongside it and the model can pull the full text of any result it wants, for free:
tools=[{"type": "viro:web_search"}, {"type": "viro:web_fetch"}]The model can pass a freshness filter (pd past day, pw week, pm month, py year) when a question is time-sensitive.
Gives the model the current date and time. Free, and unlike the other two tools it never makes an outbound call — it's pure computation, so it can't fail. Useful for scheduling, deadlines, and any "how long ago" question a model would otherwise have to guess at.
response = client.chat.completions.create(
model="viro/gpt-oss-120b",
messages=[{"role": "user", "content": "What day of the week is it?"}],
tools=[{"type": "viro:datetime"}],
)Optionally pass a timezone (IANA name, e.g. America/New_York) — the model decides when to include it based on the question. Defaults to UTC. An unrecognized timezone falls back to UTC rather than erroring.
web_search — 1¢ per search. web_fetch and datetime — free.
Failed calls aren't billed. If a search fails on our side — vendor outage, our credential, a rate limit — you aren't charged for it. A search that runs and legitimately finds nothing is billed, because it ran.
The bigger cost is usually tokens, not the search fee: every result snippet and fetched page becomes input tokens on the model's next round. A request that searches twice and fetches a long article can cost several times the 2¢ in search fees.
Tool spend appears in your usage record separately from token spend, so you can see the split.
web_fetch works reliably on every model we've tested, including the gpt-oss family — it's specifically web_search that has known issues on a few models.
viro/gpt-oss-120b can occasionally produce an invalid response internally when a request triggers a second search — the model returns leftover internal data instead of an answer. We detect it automatically and retry with a corrective nudge, which resolves it most of the time. If it still can't recover: a routed request (e.g. viro/optimized) transparently falls back to another model, so you get a real answer either way; a request naming viro/gpt-oss-120b directly returns a normal 503 error — safe to retry — instead of a silently wrong-looking answer.
viro/gpt-oss-20b, viro/gpt-oss-120b, and google/gemini-3.6-flash also sometimes call web_fetch with a guessed URL instead of using web_search, when only web_search is declared. This fails safely — you get back an ordinary unresolved tool call rather than a broken answer — but it means search isn't reliably invoked on those three models specifically.
Verified working (2026-08-04): both tools on viro/kimi-k2.5, viro/qwen3-235b-instruct, viro/llama-4-scout, viro/deepseek-v4-pro, and the frontier models (openai/*, anthropic/*, xai/*).
Works exactly like any other streamed completion — standard OpenAI SSE, standard SDKs, no special handling required.
stream = client.chat.completions.create(
model="viro/qwen3-235b-instruct",
messages=[{"role": "user", "content": "What's the weather in Boston?"}],
tools=[{"type": "viro:web_search"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Optionally, you can show what the model is doing while it works. Alongside the normal chunks we emit additive progress frames with an empty choices array, so every standard client ignores them safely:
{"choices": [], "viro": {"tool_status": {
"name": "web_search", "phase": "started", "call_id": "call_abc"
}}}phase is started or completed; completed frames also carry ok. Read them to render a "Searching the web…" indicator, or ignore them entirely.
One thing to expect: time-to-first-token includes however long the tools take, because the model usually searches before it has anything to say. Some models (Claude especially) narrate first — that text streams immediately, before the search runs.
Server tools sit alongside your own function tools. We execute ours; if the model calls yours, the response comes back to you with tool_calls exactly as it does today.
tools=[
{"type": "viro:web_fetch"},
{"type": "function", "function": {"name": "get_weather", ...}},
]web_fetch, web_search, and datetime are reserved function names. Declaring your own function with any of them returns 400 invalid_request rather than silently shadowing one or the other.
Streaming is supported. Set stream: true alongside any viro:* tool and tokens arrive as they're generated, including any text the model writes before it calls a tool. Tool round trips still happen server-side — you get one continuous stream.
Up to 5 upstream rounds and 10 tool calls per request. Both caps exist because a model can issue several tool calls at once — with search billed per call, rounds alone wouldn't bound what a single request can spend. Past the ceiling the model is told to answer with what it has.
Only public pages. Private, loopback, and internal addresses are refused, including via redirects. Pages over ~120k characters are truncated; non-text content (images, PDFs, binaries) is rejected.
Failures are handed to the model, not to you. A page that 404s or times out becomes an error string the model reads and can react to — it won't fail your whole request.
Billing. Tokens across every round are summed into one usage record. Fetched page content becomes input tokens on the next round, so a large page costs real money — that's the main cost driver to watch.