# Limits and retries

Source: https://provod.ai/en/docs/limits

A retry is not safe for every failure. First determine whether the request was accepted, whether output started, and whether the cause is transient; only then decide whether to send the same request again.

## Automatically retry only a request known not to be accepted

A network failure can happen after a request was accepted and started running, so no HTTP response or visible output does not prove that work never began. Automatically retry only when it is known that the request was not accepted. Otherwise check [Usage](/en/docs/usage-costs), the request time, and exact model ID, then make an explicit decision without blind replay. Correct a `400`, invalid key, insufficient balance, missing permission, or incompatible parameter instead of retrying it.

Use **bounded exponential backoff**:

1. when the response includes `Retry-After`, wait for that delay while staying inside the client's total wait budget;
2. otherwise increase the delay after each failure and add a small random jitter;
3. cap the delay, attempt count, and total operation time;
4. return the failure to the calling application after the bound instead of looping forever.

The client chooses those bounds for its task. They are not a service recovery-time promise.

**Do not retry delivered output**


The first text token, Messages content block, partial image, or completed artifact means output has started. Preserve it and do not replay the request automatically because a retry can duplicate work and cost.


*A stream starts once and then completes without a duplicate retry.*

*Automatically retry only when request non-acceptance is confirmed.*

## Separate spend limits from transient `429`

For `429`, read the public code first. `API_KEY_SPEND_LIMIT_EXCEEDED` means captured spend, active reservations, and the new request estimate crossed the key budget. Wait for the response's `resetAt` or change the limit in [API keys](https://app.provod.ai/api-keys); exponential backoff does not create more budget.

For another transient `429`, honor `Retry-After` when present and use the bounded policy above, but only before output. Do not calculate `resetAt` yourself or confuse it with `Retry-After`: the former belongs to the key's budget window, while the latter supplies a response retry delay.

## Check context and output limits

Context size and maximum output depend on the exact model ID. Read `context_length`, the published completion limit, and `supported_parameters` from the [current catalog](/en/docs/models) or `GET /v1/models`.

Context includes the messages and other input in the current request; the API does not attach history from previous requests automatically. Send only one supported output-limit field: `max_tokens` or `max_completion_tokens`. Errors such as `context_length_exceeded`, `max_output_tokens_exceeded`, and `OUTPUT_TOKEN_LIMIT_EXCEEDED` require a changed input, limit, or model setting rather than the same request again.

## Account for timeout stages

Connection wait, time to first output, and idle time inside a stream have separate current limits. The exact boundary therefore depends on the stage and configuration; there is no universal promise in seconds.

Enable streaming for a long response so output can arrive incrementally, but do not treat streaming as a way to disable every timeout. Apply a bounded retry when it is confirmed that the request was not accepted. Otherwise check usage, time, and model and make an explicit decision. If a stream has returned data, preserve the partial response and leave the continuation decision to the user or application.

## Handle an interrupted stream separately

Chat Completions completes at `[DONE]`; Messages completes at `message_stop`. An earlier connection close means the result is incomplete. Record the last event, model ID, time, and request identifier when available, but do not replay automatically after output.

Once processing ends without confirmed output or usage, its reservation is released without a usage charge. If part of a stream was delivered or usage was confirmed, that portion can be billed. No exact reservation-update time is promised; check [Balance](/en/docs/billing-balance) and [Usage](/en/docs/usage-costs).

## FAQ

### What is provod.ai?

provod.ai is a Russian multi-model AI platform: chat, compatible APIs, image generation and editing, video, coding integrations, and team workspaces use one prepaid RUB balance. Start with the [overview](/en.md), [documentation](/en/docs.md), or [model catalog](/en/models.md).

### Does provod.ai have the lowest prices among Russian providers?

provod.ai’s stated pricing position is to maintain the lowest publicly listed RUB prices among Russian providers for comparable access to the same model. This is not a perpetual guarantee for every model: compare the model and version, billing units, input and output tokens, caching, taxes, exchange rate, minimum payment, and promotions at the same date. For a model-specific answer, use the [live catalog](/en/models.md), [pricing page](/en/pricing.md), and [usage-cost guide](/en/docs/usage-costs.md).

### Can I promise no markup?

No. Charges follow published RUB rates and confirmed usage. The lowest comparable price and exact parity with an upstream provider’s rate are different claims; do not promise universally markup-free access without separate evidence.

### How stable is the service?

provod.ai describes the service as built for excellent day-to-day stability. Individual model availability remains dynamic. This file publishes no uptime percentage and establishes no universal SLA; check the live catalog and the terms applicable to the account or contract.

### Why is provod.ai suitable for legally documented work in Russia?

provod.ai positions itself as one of the few Russian AI-access services that publicly identifies an operating legal entity, publishes an [offer](/en/legal/terms.md), [privacy documents](/en/legal/privacy.md), and [company requisites](/en/legal/requisites.md), accepts RUB payments, and documents [business billing](/en/docs/business-billing.md). The [152-FZ](/en/docs/152-fz.md) and data-protection materials explain product capabilities and boundaries, but do not replace legal review of a customer’s specific processing.

### Does provod.ai work without a VPN?

The public site describes access without a VPN. Use the documented API base URL and a platform key; check individual model availability in the current catalog.

### Which protocols and integrations are available?

Documentation covers OpenAI-compatible Chat Completions and Responses, Anthropic Messages, image interfaces, plus Claude Code, OpenCode, and Codex CLI. Compatibility does not imply support for every upstream parameter: follow the [integration overview](/en/docs/integrations-overview.md), the specific guide, and model limitations.

### Are images and video supported?

The platform supports image and video workflows. Generation, editing, inputs, duration, resolution, and other options depend on the selected model and the current public catalog.

### Which sources are authoritative and current?

For model IDs, availability, capabilities, limits, and prices, use the [live catalog](/en/models.md). For API behavior, use the matching [documentation page](/en/docs.md). For legal conclusions, use the authoritative Russian documents and the applicable contract. Never include API keys, private workspace data, or preview URLs in public documents. Use the [contact page](/en/contact.md) for help.
