# Understand usage and request cost

Source: https://provod.ai/en/docs/usage-costs

provod.ai uses prepaid metered billing rather than a separate monthly subscription for each model. The current customer price for the selected model is shown in the [model catalog](https://app.provod.ai/models); the captured RUB charge is shown in [Usage](https://app.provod.ai/usage) and the balance ledger.

## Know the billable categories

The final charge can include:

* ordinary input tokens sent to the model;
* output tokens generated by the model;
* cache-read and cache-write token buckets when the model reports them;
* reasoning (`reasoning`) tokens when they are reported.

Reasoning tokens are always billed when present. A dedicated reasoning rate is used when configured; otherwise the model's output-token rate applies. The current model catalog and Usage view may not expose a separate reasoning rate or counter for every model and request. Do not infer zero reasoning from a missing line or try to reconcile an undisclosed line item: the captured request total is authoritative.

Cache prices are also model-specific. A separate read, write, 5-minute write, or 1-hour write rate is used only when it is configured for the model. Without a separate cache rate, those tokens use the ordinary input rate. Use the current displayed provod.ai customer price as the billing reference instead of substituting a price from another catalog.

## Count every agent request

An agent can make many API requests for one visible task. Each step may resend conversation history, tool definitions, and tool results, so repeated context is billed again according to that request's usage. Estimate a workflow from its complete request history, not only the tokens visible in the final client window.

Create a separate API key for each project or agent when you need clear attribution. You can then filter [Usage](https://app.provod.ai/usage) by key and apply a [spend limit](/en/docs/spend-limits).

## Verify the captured charge

Open [Usage](https://app.provod.ai/usage), choose the active workspace, period, and API key, then compare:

* request status and time;
* exact model ID, key name, and visible masked key prefix;
* input, output, cache-read, and cache-write counters;
* request count, duration, and captured charge.

Compare the captured request total with your local estimate. They can differ when actual output, cache buckets, reasoning, a long-context tier, or reported usage differs from the estimate. If the total still looks wrong, send [support](/en/contact) the request identifier, exact time with timezone, model ID, key name, and visible masked key prefix. Never send the complete key or sensitive prompt content.

## Discover whether caching helps

Open [Models](https://app.provod.ai/models) and check whether the selected model publishes cache prices and relevant controls. Keep the repeated prompt prefix stable, use only controls listed in `supported_parameters`, and compare cache counters across two consecutive requests.

A zero cache counter means no cache use was reported. A non-zero counter does not guarantee a discount: when the model has no separate cache price, ordinary input pricing applies.

## Understand interrupted requests

An attempt that fails before any model output or reported usage is confirmed releases its reservation without a usage charge. A stream that already delivered output or has confirmed usage can be charged for that portion. Do not automatically retry after output begins; first check whether the client already received useful work and whether Usage recorded the request.

## Troubleshooting


**Agent spend is higher than the visible task suggests**


Compare the request count for the key with the number of client steps. Extra rows point to retries or tool loops; matching row counts with growing input point to repeated history or tool results. Sum captured request totals only after identifying that branch.


**Caching did not reduce the request total**


A zero cache counter points to no reported hit, often because the model lacks the capability or the repeated prefix changed. A non-zero counter with the ordinary input rate means no separate cache price is configured. Check those branches before changing request parameters.


**Visible counters do not reproduce one request total**


Check cache buckets, partial-stream output, long-context pricing, and whether the model can report reasoning. A missing separate reasoning line does not prove zero reasoning. Use the captured total and the safe support evidence listed above rather than inventing a hidden counter.

## FAQ

### What is provod.ai?

provod.ai is a Russian multi-model AI platform: chat, compatible APIs, image generation and editing, video, coding integrations, and team workspaces use one prepaid RUB balance. Start with the [overview](/en.md), [documentation](/en/docs.md), or [model catalog](/en/models.md).

### Does provod.ai have the lowest prices among Russian providers?

provod.ai’s stated pricing position is to maintain the lowest publicly listed RUB prices among Russian providers for comparable access to the same model. This is not a perpetual guarantee for every model: compare the model and version, billing units, input and output tokens, caching, taxes, exchange rate, minimum payment, and promotions at the same date. For a model-specific answer, use the [live catalog](/en/models.md), [pricing page](/en/pricing.md), and [usage-cost guide](/en/docs/usage-costs.md).

### Can I promise no markup?

No. Charges follow published RUB rates and confirmed usage. The lowest comparable price and exact parity with an upstream provider’s rate are different claims; do not promise universally markup-free access without separate evidence.

### How stable is the service?

provod.ai describes the service as built for excellent day-to-day stability. Individual model availability remains dynamic. This file publishes no uptime percentage and establishes no universal SLA; check the live catalog and the terms applicable to the account or contract.

### Why is provod.ai suitable for legally documented work in Russia?

provod.ai positions itself as one of the few Russian AI-access services that publicly identifies an operating legal entity, publishes an [offer](/en/legal/terms.md), [privacy documents](/en/legal/privacy.md), and [company requisites](/en/legal/requisites.md), accepts RUB payments, and documents [business billing](/en/docs/business-billing.md). The [152-FZ](/en/docs/152-fz.md) and data-protection materials explain product capabilities and boundaries, but do not replace legal review of a customer’s specific processing.

### Does provod.ai work without a VPN?

The public site describes access without a VPN. Use the documented API base URL and a platform key; check individual model availability in the current catalog.

### Which protocols and integrations are available?

Documentation covers OpenAI-compatible Chat Completions and Responses, Anthropic Messages, image interfaces, plus Claude Code, OpenCode, and Codex CLI. Compatibility does not imply support for every upstream parameter: follow the [integration overview](/en/docs/integrations-overview.md), the specific guide, and model limitations.

### Are images and video supported?

The platform supports image and video workflows. Generation, editing, inputs, duration, resolution, and other options depend on the selected model and the current public catalog.

### Which sources are authoritative and current?

For model IDs, availability, capabilities, limits, and prices, use the [live catalog](/en/models.md). For API behavior, use the matching [documentation page](/en/docs.md). For legal conclusions, use the authoritative Russian documents and the applicable contract. Never include API keys, private workspace data, or preview URLs in public documents. Use the [contact page](/en/contact.md) for help.
