Prompt caching
When you send the same prefix twice — a long system prompt, a set of tool definitions, a document you're asking several questions about — providers that support caching serve the repeated part from their cache instead of reprocessing it, and charge less for it. Viro passes that discount straight through. There is nothing to switch on and no cache_control marker to add to your request.
It does not apply everywhere, and where it doesn't there is no warning — the request simply costs the normal rate.
The Models page has a Cache $/1M column: a rate means that model caches, and a dash means it doesn't. That column reads from the same catalog the billing code prices against, so it can't drift from what you're actually charged — which is why this page doesn't repeat the list.
The shape of it, as of this writing:
- •
openai/*models cache, at roughly a tenth of their input rate. - • The TensorX-hosted open-weight models cache, at roughly a quarter of their input rate.
- • Nscale-hosted models — most of the
viro/*catalog — do not cache at all, and neither doanthropic/*,google/*orxai/*.
Some models also charge a one-time cache write premium on the request that populates the cache — it costs slightly more than a normal request so that every later one costs much less. Where that applies, the Models page shows it under the read rate.
Caching matches on an exact prefix. Everything up to the first byte that differs can be reused; everything after it cannot. So the order of your messages decides whether you get the discount at all:
- • Put the stable parts first — system prompt, tool definitions, fixed examples, the document.
- • Put the varying parts last — the user's actual question, a timestamp, a request ID.
- • Keep the stable part byte-identical. Injecting the current time into a system prompt invalidates the cache on every single request, which is the most common way to accidentally never cache anything.
Providers only cache prefixes above a minimum length, and the threshold varies between them. Short prompts won't cache no matter how often you repeat them — this is worth having on long system prompts and agent loops, and not worth restructuring a chatbot for.
The response's usage block breaks out how much of your input was served from the cache:
{
"usage": {
"prompt_tokens": 2890,
"completion_tokens": 84,
"total_tokens": 2974,
"prompt_tokens_details": { "cached_tokens": 2816 }
}
}cached_tokens is the portion billed at the cache-read rate; the rest is billed as fresh input. Absent means the provider reported no cache accounting for that request. The same split appears per request on the Usage page — open a request to see cached and cache-write tokens alongside the charge.
What that's worth in practice: an identical repeated request on viro/deepseek-v4-pro, measured 2026-09-02, cost 0.127¢ against 0.429¢ uncached — about a third of the price for the same answer.
This is a real gap rather than an oversight, and it's the one most likely to cost you money unexpectedly.
Anthropic only caches blocks a request explicitly marks with cache_control. Viro's Anthropic adapter rebuilds request content from an allowlist of fields and flattens the system prompt to a string, so that marker cannot survive the translation. If you send it, the request still works — the marker is silently dropped and you pay the ordinary input rate on every token, every time.
Nothing is mispriced by this: you are charged exactly the rate the Models page shows. But if your workload is an agent loop with a long fixed system prompt, an openai/* or TensorX-hosted model will cost materially less than an anthropic/* one for reasons that have nothing to do with the listed per-token price.
A prompt cache is provider-side state, which puts it outside the zero-retention posture that applies to ordinary inference. Where a model is marked ZDR on the Models page, that covers the request and response content; cached prefixes live under the provider's own cache retention policy. See Data Privacy & Retention.