Viro API

Energy & emissions methodology

Viro reports estimated energy consumption and greenhouse gas emissions for each API request. These are modeled estimates, not measurements. This page documents exactly how they are produced, what they assume, and where they are wrong.

In one paragraph

Energy and emissions are estimated using the EcoLogits methodology (v0.11.1), an open life-cycle assessment method maintained by the GenAI Impact non-profit. Viro supplies measured per-request latency and token counts; the remaining parameters use the methodology's published defaults, including a world-average grid intensity where the serving region is not disclosed by the upstream provider. Every substituted value is flagged on the individual request. Impact is not estimated at all for models whose parameter counts are unpublished.

What is calculated

Total impact is the sum of usage (electricity consumed serving the request) and embodied (a share of the manufacturing footprint of the hardware, amortized over a three-year life).

gpuEnergyPerToken = α · e^(β·batch) · activeParamsB + γ
gpuEnergy         = outputTokens · gpuEnergyPerToken / 1000        → kWh
latency           = min(measuredLatency, ttft + outputTokens / tps)
serverEnergy      = (latency/3600) · 1.2kW · (gpusRequired/8) · (1/batch)
itEnergy          = serverEnergy + gpusRequired · gpuEnergy
requestEnergy     = PUE · itEnergy

gwp   = requestEnergy · gridIntensity + embodiedGwp
water = itEnergy · (dcWUE + PUE · gridWUE)

memoryRequired = 1.2 · totalParamsB · 16 / 8
gpusRequired   = 2^ceil(log2(ceil(memoryRequired / 80GB)))
embodiedGwp    = latency/(batch · 3yr) · (gpusRequired·273 + 5700·gpusRequired/8)

α = 1.1665e-6, β = -1.1206e-2, γ = 4.0529e-5, fitted by EcoLogits against the ML.ENERGY Leaderboard.

Assumptions

Only latency and token counts are measured. Everything else is a published default:

  • Grid intensity — world average, 0.458 kgCO₂e/kWh. No upstream provider currently discloses which facility served a given request.
  • Hardware — NVIDIA H100 80GB, eight per server, 1.2 kW non-GPU draw. Applied uniformly; real fleets differ.
  • PUE — 1.2. Batch size — 64 concurrent requests.
  • Precision — 16-bit weights. Hardware life — 3 years.
Where these numbers are wrong
Serving region is unknown — up to ~9x

Grid carbon intensity varies enormously by geography: Norway is 0.028 and Ireland 0.257 kgCO₂e/kWh. Because no provider discloses the serving facility per request, we substitute the world average on every request and flag it. This is the single largest source of error in the estimate, and it is not resolvable by us alone.

Estimated parameter counts — up to ~3x spread

Closed models do not publish their architecture, so their parameter count is itself a modeled range. The console shows the midpoint of that range as a single number for readability, but the underlying spread is real: a model whose estimated active parameters span 200–600B, for example, could genuinely be anywhere in that interval — the displayed figure is the middle of it, not a measurement of it.

Input-heavy requests are understated

The model is driven entirely by output tokens; the prefill phase is not modeled. A request that generates little or no output — a refusal, a failed completion, long-context retrieval with a terse answer — reports far less impact than it actually caused. Requests generating zero output tokens report zero.

Training is excluded

These figures cover inference only. Independent analysis suggests amortized model training can represent a third to a half of total AI emissions. Amortizing it requires assumptions about a model's total lifetime token volume that we cannot currently defend, so it is omitted rather than guessed at.

Embeddings and image generation are not estimated

The methodology models token-by-token decoding, which describes neither. Those requests show “—” rather than a number derived from a model that does not apply.

Why some requests show “—”

Estimating impact requires a model's parameter count. Where that is not published on an official model card, Viro does not estimate — a plausible-looking figure derived from an invented architecture is worse than no figure.

Account totals state how many requests were excluded for this reason, so a total is never quietly computed over a subset.

Renewable claims are reported separately

The energy and CO₂e figures here are location-based physical accounting: the electricity actually drawn, at the average intensity of the grid supplying it.

Viro's renewable indicators — “100% Clean Energy” and “Energy Matched” — are market-based claims about generation and procurement. The GHG Protocol requires both be disclosed, and disclosed separately. Viro does not net one against the other, and a renewable claim never reduces the physical figures on this page. See Models for per-model energy provenance.

Standards alignment
  • Method: EcoLogits v0.11.1, life-cycle based, consistent with ITU-T L.1801 / ETSI ES 204 135.
  • Accounting frame: GHG Protocol, incl. the ICT Sector Guidance for cloud and data-centre services. For customers, Viro usage is Scope 3 Category 1.
  • Disclosure shape follows Watershed's AI Emissions Framework. Current calculation tier: Activity Tier (simplified) — we hold the token split and model family, but not the serving region.

Viro does not claim conformance with ISO/IEC TR 20226:2025 (a Technical Report, with nothing to conform to) or IEEE P7100 (not yet published).

Versioning

Every figure is stamped with the methodology version in force when it was produced — currently viro-ecologits-0.11.1-v1.

Historical records are never recomputed. A figure supplied for a reporting period stays reproducible under the method used to produce it; improvements get a new version and apply going forward.

EcoLogits is developed by GenAI Impact and licensed under MPL-2.0. Grid intensity factors derive from Our World in Data and ADEME Base Empreinte; embodied hardware factors from Boavizta.