# Public model pricing in rubles

Public RUB token prices for chat models in the catalog.

Prices are shown per 1M tokens in RUB. Cache columns are shown only when a published rate exists.

## Video generation starting prices

Input / 1M

Output / 1M

### Why can an agent cost more than its last context?

An agent can make several requests, each with history and tool output. Check request history rather than only the final visible context.

### Why is part of my balance reserved?

In-flight requests can hold an estimated amount. The available balance reflects active reservations until the request is settled.

### Is an interrupted stream billed?

A failure before confirmed output or usage releases its hold. A stream that already delivered output can be charged for confirmed usage.

### What if a payment is still pending?

Check the operation status and refresh the balance. If a confirmed payment is missing, send support its payment ID, time, and receipt before paying again.

## Related pages

- [/en/models](/en/models.md)

Source: public catalog; checked 2026-09-30T22:02:16.046Z.

## Anthropic: Claude Fable 5

- ID: claude-fable-5

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

[Markdown](/en/models/anthropic/claude-5-fable-20260609.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 835.59

- completionPricePer1mTokens: 4,177.94

- cacheReadPricePer1mTokens: 83.56

- cacheWritePricePer1mTokens: 1,044.49

- cacheWrite5mPricePer1mTokens: 1,044.49

- cacheWrite1hPricePer1mTokens: 1,671.18

---

## Claude Fable 5.1

- ID: claude-fable-5.1

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

[Markdown](/en/models/anthropic/claude-fable-5.1-20260831.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, tool\_choice, tools, verbosity

- Currency: RUB

- promptPricePer1mTokens: 835.59

- completionPricePer1mTokens: 4,177.94

- cacheReadPricePer1mTokens: 20.89

- cacheWritePricePer1mTokens: 1,044.49

- cacheWrite5mPricePer1mTokens: 1,044.49

- cacheWrite1hPricePer1mTokens: 1,671.18

---

## Claude Haiku 4.5

- ID: claude-haiku-4.5

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

[Markdown](/en/models/anthropic/claude-4.5-haiku-20251001.md)

- Context: 200000

- Max output tokens: 64000

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, reasoning, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 83.56

- completionPricePer1mTokens: 417.79

- cacheReadPricePer1mTokens: 8.36

- cacheWritePricePer1mTokens: 104.45

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: 167.12

---

## Anthropic: Claude Opus 4.6

- ID: claude-opus-4.6

- Publisher: Anthropic

- Kind: text

- Available: true

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

[Markdown](/en/models/anthropic/claude-4.6-opus-20260205.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,088.97

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: 522.24

- cacheWrite5mPricePer1mTokens: 522.24

- cacheWrite1hPricePer1mTokens: 835.59

---

## Anthropic: Claude Opus 4.7

- ID: claude-opus-4.7

- Publisher: Anthropic

- Kind: text

- Available: true

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

[Markdown](/en/models/anthropic/claude-4.7-opus-20260416.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,088.97

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: 522.24

- cacheWrite5mPricePer1mTokens: 522.24

- cacheWrite1hPricePer1mTokens: 835.59

---

## Anthropic: Claude Opus 4.8

- ID: claude-opus-4.8

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

[Markdown](/en/models/anthropic/claude-4.8-opus-20260528.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,088.97

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: 522.24

- cacheWrite5mPricePer1mTokens: 522.24

- cacheWrite1hPricePer1mTokens: 835.59

---

## Claude Opus 5

- ID: claude-opus-5

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

[Markdown](/en/models/anthropic/claude-opus-5-20260723.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,088.97

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: 522.24

- cacheWrite5mPricePer1mTokens: 522.24

- cacheWrite1hPricePer1mTokens: 835.59

---

## Claude Opus 5.5

- ID: claude-opus-5.5

- Publisher: Anthropic

- Kind: text

- Available: true

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

[Markdown](/en/models/anthropic/claude-opus-5.5-20260921.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, verbosity, top\_p

- Currency: RUB

- promptPricePer1mTokens: 334.24

- completionPricePer1mTokens: 1,671.18

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: 417.79

- cacheWrite5mPricePer1mTokens: 417.79

- cacheWrite1hPricePer1mTokens: 668.47

---

## Anthropic: Claude Sonnet 4.6

- ID: claude-sonnet-4.6

- Publisher: Anthropic

- Kind: text

- Available: true

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

[Markdown](/en/models/anthropic/claude-4.6-sonnet-20260217.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 250.68

- completionPricePer1mTokens: 1,253.38

- cacheReadPricePer1mTokens: 25.07

- cacheWritePricePer1mTokens: 313.35

- cacheWrite5mPricePer1mTokens: 313.35

- cacheWrite1hPricePer1mTokens: 501.35

---

## Anthropic: Claude Sonnet 5

- ID: claude-sonnet-5

- Publisher: Anthropic

- Kind: text

- Available: true

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

[Markdown](/en/models/anthropic/claude-sonnet-5-20260630.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 835.59

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: 208.90

- cacheWrite5mPricePer1mTokens: 208.90

- cacheWrite1hPricePer1mTokens: 334.24

---

## DeepSeek: DeepSeek V4 Flash 0423

- ID: deepseek-v4-flash

- Publisher: DeepSeek

- Kind: text

- Available: true

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

[Markdown](/en/models/deepseek/deepseek-v4-flash-20260423.md)

- Context: 1048576

- Max output tokens: 393216

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, thinking, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 55.15

- completionPricePer1mTokens: 165.45

- cacheReadPricePer1mTokens: 1.75

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## DeepSeek: DeepSeek V4 Flash 0731

- ID: deepseek-v4-flash-0731

- Publisher: DeepSeek

- Kind: text

- Available: true

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

[Markdown](/en/models/deepseek/deepseek-v4-flash-20260731.md)

- Context: 1048576

- Max output tokens: 393216

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, thinking, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 55.15

- completionPricePer1mTokens: 165.45

- cacheReadPricePer1mTokens: 1.75

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## DeepSeek: DeepSeek V4 Pro

- ID: deepseek-v4-pro

- Publisher: DeepSeek

- Kind: text

- Available: true

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

[Markdown](/en/models/deepseek/deepseek-v4-pro-20260423.md)

- Context: 1048576

- Max output tokens: 384000

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, thinking, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 165.45

- completionPricePer1mTokens: 496.34

- cacheReadPricePer1mTokens: 5.51

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## DeepSeek V4 Pro 0813

- ID: deepseek-v4-pro-0813

- Publisher: DeepSeek

- Kind: text

- Available: true

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

[Markdown](/en/models/deepseek/deepseek-v4-pro-20260813.md)

- Context: 1024000

- Max output tokens: 384000

- Input: text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 77.21

- completionPricePer1mTokens: 231.62

- cacheReadPricePer1mTokens: 2.57

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## DeepSeek V4.1 Flash

- ID: deepseek-v4.1-flash

- Publisher: DeepSeek

- Kind: text

- Available: true

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

[Markdown](/en/models/deepseek/deepseek-v4.1-flash-20260910.md)

- Context: 1048576

- Max output tokens: 393216

- Input: text, image

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logprobs, max\_tokens, presence\_penalty, reasoning, reasoning\_effort, response\_format, stop, temperature, tool\_choice, tools, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 100.27

- cacheReadPricePer1mTokens: 0.50

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 2.5 Flash

- ID: gemini-2.5-flash

- Publisher: Google

- Kind: text

- Available: true

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

[Markdown](/en/models/google/gemini-2.5-flash.md)

- Context: 1048576

- Max output tokens: 65535

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 208.90

- cacheReadPricePer1mTokens: 2.51

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 2.5 Flash Lite

- ID: gemini-2.5-flash-lite

- Publisher: Google

- Kind: text

- Available: true

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

[Markdown](/en/models/google/gemini-2.5-flash-lite.md)

- Context: 1048576

- Max output tokens: 65535

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 8.36

- completionPricePer1mTokens: 33.42

- cacheReadPricePer1mTokens: 0.84

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Gemini 2.5 Pro

- ID: gemini-2.5-pro

- Publisher: Google

- Kind: text

- Available: true

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

[Markdown](/en/models/google/gemini-2.5-pro.md)

- Context: 1048576

- Max output tokens: 65536

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 104.45

- completionPricePer1mTokens: 835.59

- cacheReadPricePer1mTokens: 10.44

- cacheWritePricePer1mTokens: 31.33

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 3 Flash Preview

- ID: gemini-3-flash-preview

- Publisher: Google

- Kind: text

- Available: true

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

[Markdown](/en/models/google/gemini-3-flash-preview-20251217.md)

- Context: 1048576

- Max output tokens: 65535

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 41.78

- completionPricePer1mTokens: 250.68

- cacheReadPricePer1mTokens: 4.18

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 3.1 Flash Lite

- ID: gemini-3.1-flash-lite

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

[Markdown](/en/models/google/gemini-3.1-flash-lite-20260507.md)

- Context: 1048576

- Max output tokens: 65536

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 20.89

- completionPricePer1mTokens: 125.34

- cacheReadPricePer1mTokens: 2.09

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 3.1 Pro Preview

- ID: gemini-3.1-pro-preview

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

[Markdown](/en/models/google/gemini-3.1-pro-preview-20260219.md)

- Context: 1048576

- Max output tokens: 65536

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 1,002.71

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 3.5 Flash

- ID: gemini-3.5-flash

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

[Markdown](/en/models/google/gemini-3.5-flash-20260519.md)

- Context: 1048576

- Max output tokens: 65536

- Input: text, image, file, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 125.34

- completionPricePer1mTokens: 752.03

- cacheReadPricePer1mTokens: 12.53

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Gemini 3.5 Flash Lite

- ID: gemini-3.5-flash-lite

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

[Markdown](/en/models/google/gemini-3.5-flash-lite-20260721.md)

- Context: 1048576

- Max output tokens: 65536

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 208.90

- cacheReadPricePer1mTokens: 2.51

- cacheWritePricePer1mTokens: 6.96

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Gemini 3.6 Flash

- ID: gemini-3.6-flash

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

[Markdown](/en/models/google/gemini-3.6-flash-20260721.md)

- Context: 1048576

- Max output tokens: 65536

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 62.67

- completionPricePer1mTokens: 313.35

- cacheReadPricePer1mTokens: 6.27

- cacheWritePricePer1mTokens: 3.48

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Gemini 3.7 Flash

- ID: gemini-3.7-flash

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

[Markdown](/en/models/google/gemini-3.7-flash-20260813.md)

- Context: 1048576

- Max output tokens: 65536

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 62.67

- completionPricePer1mTokens: 313.35

- cacheReadPricePer1mTokens: 6.27

- cacheWritePricePer1mTokens: 3.48

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Gemini 3.8 Flash

- ID: gemini-3.8-flash

- Publisher: Google

- Kind: text

- Available: true

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

[Markdown](/en/models/google/gemini-3.8-flash-20260902.md)

- Context: 1048576

- Max output tokens: 65536

- Input: text, image, video, file, audio

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, reasoning\_effort, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 62.67

- completionPricePer1mTokens: 313.35

- cacheReadPricePer1mTokens: 6.27

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.5

- ID: glm-4.5

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

[Markdown](/en/models/z-ai/glm-4.5.md)

- Context: 131072

- Max output tokens: 98304

- Input: text

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, response\_format, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 50.14

- completionPricePer1mTokens: 183.83

- cacheReadPricePer1mTokens: 9.19

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.5 Air

- ID: glm-4.5-air

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

[Markdown](/en/models/z-ai/glm-4.5-air.md)

- Context: 131072

- Max output tokens: 98304

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, max\_tokens, presence\_penalty, reasoning, repetition\_penalty, seed, stop, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 16.71

- completionPricePer1mTokens: 91.91

- cacheReadPricePer1mTokens: 2.51

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.5V

- ID: glm-4.5v

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

[Markdown](/en/models/z-ai/glm-4.5v.md)

- Context: 65536

- Max output tokens: 16384

- Input: text, image

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, max\_tokens, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 50.14

- completionPricePer1mTokens: 150.41

- cacheReadPricePer1mTokens: 9.19

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.6

- ID: glm-4.6

- Publisher: Z.ai

- Kind: text

- Available: true

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

[Markdown](/en/models/z-ai/glm-4.6.md)

- Context: 204800

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 50.14

- completionPricePer1mTokens: 183.83

- cacheReadPricePer1mTokens: 9.19

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.6V

- ID: glm-4.6v

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

[Markdown](/en/models/z-ai/glm-4.6-20251208.md)

- Context: 131072

- Max output tokens: 32768

- Input: image, text, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, max\_tokens, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 75.20

- cacheReadPricePer1mTokens: 4.18

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 4.7

- ID: glm-4.7

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

[Markdown](/en/models/z-ai/glm-4.7-20251222.md)

- Context: 204800

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 50.14

- completionPricePer1mTokens: 183.83

- cacheReadPricePer1mTokens: 9.19

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5

- ID: glm-5

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

[Markdown](/en/models/z-ai/glm-5-20260211.md)

- Context: 204800

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 83.56

- completionPricePer1mTokens: 267.39

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5 Turbo

- ID: glm-5-turbo

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

[Markdown](/en/models/z-ai/glm-5-turbo-20260315.md)

- Context: 202752

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, response\_format, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 100.27

- completionPricePer1mTokens: 334.24

- cacheReadPricePer1mTokens: 20.05

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5.1

- ID: glm-5.1

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

[Markdown](/en/models/z-ai/glm-5.1-20260406.md)

- Context: 204800

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 116.98

- completionPricePer1mTokens: 367.66

- cacheReadPricePer1mTokens: 21.73

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5.2

- ID: glm-5.2

- Publisher: Z.ai

- Kind: text

- Available: true

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

[Markdown](/en/models/z-ai/glm-5.2-20260616.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, parallel\_tool\_calls, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 80.72

- completionPricePer1mTokens: 253.68

- cacheReadPricePer1mTokens: 16.14

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5.3

- ID: glm-5.3

- Publisher: Z.ai

- Kind: text

- Available: true

GLM 5.3 is Z.ai’s reasoning model for long-context text and agent workflows.

[Markdown](/en/models/z-ai/glm-5.3.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, parallel\_tool\_calls, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 116.98

- completionPricePer1mTokens: 367.66

- cacheReadPricePer1mTokens: 21.73

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5.3 Flash

- ID: glm-5.3-flash

- Publisher: Z.ai

- Kind: text

- Available: true

GLM 5.3 Flash is Z.ai’s efficient multimodal reasoning model for long-context and agent workflows.

[Markdown](/en/models/z-ai/glm-5.3-flash.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text, image, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_tokens, min\_p, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 12.53

- completionPricePer1mTokens: 41.78

- cacheReadPricePer1mTokens: 4.18

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GLM 5.3 FlashX

- ID: glm-5.3-flashx

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

[Markdown](/en/models/z-ai/glm-5.3-flashx-20260918.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text, image, video

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, reasoning\_effort, response\_format, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 30.92

- completionPricePer1mTokens: 104.45

- cacheReadPricePer1mTokens: 6.27

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Z.ai: GLM 5V Turbo

- ID: glm-5v-turbo

- Publisher: Z.ai

- Kind: text

- Available: true

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

[Markdown](/en/models/z-ai/glm-5v-turbo-20260401.md)

- Context: 202752

- Max output tokens: 131072

- Input: image, text, video

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, response\_format, temperature, tool\_choice, tools, top\_k, top\_p

- Currency: RUB

- promptPricePer1mTokens: 100.27

- completionPricePer1mTokens: 334.24

- cacheReadPricePer1mTokens: 20.05

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-4.1

- ID: gpt-4.1

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

[Markdown](/en/models/openai/gpt-4.1-2025-04-14.md)

- Context: 1047576

- Max output tokens: 32768

- Input: image, text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 334.24

- completionPricePer1mTokens: 1,336.94

- cacheReadPricePer1mTokens: 83.56

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-4.1 Mini

- ID: gpt-4.1-mini

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

[Markdown](/en/models/openai/gpt-4.1-mini-2025-04-14.md)

- Context: 1047576

- Max output tokens: 32768

- Input: text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 66.85

- completionPricePer1mTokens: 267.39

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-4.1 Nano

- ID: gpt-4.1-nano

- Publisher: OpenAI

- Kind: text

- Available: true

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

[Markdown](/en/models/openai/gpt-4.1-nano-2025-04-14.md)

- Context: 1047576

- Max output tokens: 32768

- Input: text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 16.71

- completionPricePer1mTokens: 66.85

- cacheReadPricePer1mTokens: 4.18

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-4o-mini

- ID: gpt-4o-mini

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-4o mini is OpenAI's newest model after \[GPT-4 Omni\](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

[Markdown](/en/models/openai/gpt-4o-mini.md)

- Context: 128000

- Max output tokens: 16384

- Input: image, text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 100.27

- cacheReadPricePer1mTokens: 12.53

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.4

- ID: gpt-5.4

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

[Markdown](/en/models/openai/gpt-5.4-20260305.md)

- Context: 1050000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_key, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 208.90

- completionPricePer1mTokens: 1,253.38

- cacheReadPricePer1mTokens: 20.89

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.4 Mini

- ID: gpt-5.4-mini

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

[Markdown](/en/models/openai/gpt-5.4-mini-20260317.md)

- Context: 400000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_key, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 62.67

- completionPricePer1mTokens: 376.01

- cacheReadPricePer1mTokens: 6.27

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.4 Nano

- ID: gpt-5.4-nano

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

[Markdown](/en/models/openai/gpt-5.4-nano-20260317.md)

- Context: 400000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_key, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 16.71

- completionPricePer1mTokens: 104.45

- cacheReadPricePer1mTokens: 1.67

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.5

- ID: gpt-5.5

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

[Markdown](/en/models/openai/gpt-5.5-20260423.md)

- Context: 1050000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_key, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,506.76

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.6 Luna

- ID: gpt-5.6-luna

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

[Markdown](/en/models/openai/gpt-5.6-luna-20260709.md)

- Context: 1050000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_breakpoint, prompt\_cache\_key, prompt\_cache\_options, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 16.71

- completionPricePer1mTokens: 100.27

- cacheReadPricePer1mTokens: 1.67

- cacheWritePricePer1mTokens: 20.89

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.6 Sol

- ID: gpt-5.6-sol

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

[Markdown](/en/models/openai/gpt-5.6-sol-20260709.md)

- Context: 1050000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, prediction, presence\_penalty, prompt\_cache\_breakpoint, prompt\_cache\_key, prompt\_cache\_options, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 417.79

- completionPricePer1mTokens: 2,506.76

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: 522.24

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## OpenAI: GPT-5.6 Terra

- ID: gpt-5.6-terra

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

[Markdown](/en/models/openai/gpt-5.6-terra-20260709.md)

- Context: 1050000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, prompt\_cache\_breakpoint, prompt\_cache\_key, prompt\_cache\_options, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 208.90

- completionPricePer1mTokens: 1,253.38

- cacheReadPricePer1mTokens: 20.89

- cacheWritePricePer1mTokens: 261.12

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-6 Astra

- ID: gpt-6-astra

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

[Markdown](/en/models/openai/gpt-6-astra-20260903.md)

- Context: 1050000

- Max output tokens: 128000

- Input: file, image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, seed, structured\_outputs, tool\_choice, tools, parallel\_tool\_calls

- Currency: RUB

- promptPricePer1mTokens: 835.59

- completionPricePer1mTokens: 4,177.94

- cacheReadPricePer1mTokens: 83.56

- cacheWritePricePer1mTokens: 1,044.49

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-6 Luna

- ID: gpt-6-luna

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

[Markdown](/en/models/openai/gpt-6-luna-20260922.md)

- Context: 1050000

- Max output tokens: 128000

- Input: file, image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, seed, structured\_outputs, tool\_choice, tools, parallel\_tool\_calls

- Currency: RUB

- promptPricePer1mTokens: 8.36

- completionPricePer1mTokens: 41.78

- cacheReadPricePer1mTokens: 0.84

- cacheWritePricePer1mTokens: 10.44

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## GPT-6 Sol

- ID: gpt-6-sol

- Publisher: OpenAI

- Kind: text

- Available: true

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

[Markdown](/en/models/openai/gpt-6-sol-20260922.md)

- Context: 1050000

- Max output tokens: 128000

- Input: file, image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, seed, structured\_outputs, tool\_choice, tools, parallel\_tool\_calls

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 835.59

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: 208.90

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## xAI: Grok 4.3

- ID: grok-4.3

- Publisher: xAI

- Kind: text

- Available: true

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

[Markdown](/en/models/x-ai/grok-4.3-20260430.md)

- Context: 1000000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 104.45

- completionPricePer1mTokens: 208.90

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## xAI: Grok 4.5

- ID: grok-4.5

- Publisher: xAI

- Kind: text

- Available: true

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

[Markdown](/en/models/x-ai/grok-4.5-20260708.md)

- Context: 500000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 501.35

- cacheReadPricePer1mTokens: not published

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Grok 4.6

- ID: grok-4.6

- Publisher: xAI

- Kind: text

- Available: true

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by \[Grok 4.7\](/x-ai/grok-4.7).

[Markdown](/en/models/x-ai/grok-4.6-20260810.md)

- Context: 500000

- Max output tokens: 450000

- Input: image, text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, reasoning, reasoning\_effort, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 501.35

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Grok 4.7

- ID: grok-4.7

- Publisher: xAI

- Kind: text

- Available: true

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

[Markdown](/en/models/x-ai/grok-4.7-20260916.md)

- Context: 500000

- Max output tokens: 450000

- Input: text, image, file

- Output: text

- Parameters: include\_reasoning, max\_tokens, reasoning, reasoning\_effort, response\_format, seed, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 501.35

- cacheReadPricePer1mTokens: 41.78

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## xAI: Grok Build 0.1

- ID: grok-build-0.1

- Publisher: xAI

- Kind: text

- Available: true

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

[Markdown](/en/models/x-ai/grok-build-0.1-20260520.md)

- Context: 256000

- Max output tokens: 128000

- Input: text, image, file

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 104.45

- completionPricePer1mTokens: 208.90

- cacheReadPricePer1mTokens: 16.71

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Hy4 preview

- ID: hy4-preview

- Publisher: tencent

- Kind: text

- Available: true

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

[Markdown](/en/models/tencent/hy4-preview-20260827.md)

- Context: 1048576

- Max output tokens: 64000

- Input: text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 139.38

- completionPricePer1mTokens: 417.96

- cacheReadPricePer1mTokens: 7.02

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MoonshotAI: Kimi K2.6

- ID: kimi-k2.6

- Publisher: Moonshot AI

- Kind: text

- Available: true

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

[Markdown](/en/models/moonshotai/kimi-k2.6-20260420.md)

- Context: 262144

- Max output tokens: 262144

- Input: text, image

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 79.38

- completionPricePer1mTokens: 334.24

- cacheReadPricePer1mTokens: 13.37

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MoonshotAI: Kimi K2.7 Code

- ID: kimi-k2.7-code

- Publisher: Moonshot AI

- Kind: text

- Available: true

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

[Markdown](/en/models/moonshotai/kimi-k2.7-code-20260612.md)

- Context: 262144

- Max output tokens: 262144

- Input: text, image

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 79.38

- completionPricePer1mTokens: 334.24

- cacheReadPricePer1mTokens: 15.88

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MoonshotAI: Kimi K3

- ID: kimi-k3

- Publisher: Moonshot AI

- Kind: text

- Available: true

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

[Markdown](/en/models/moonshotai/kimi-k3-20260715.md)

- Context: 1048576

- Max output tokens: 1048576

- Input: text, image

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 250.68

- completionPricePer1mTokens: 1,253.38

- cacheReadPricePer1mTokens: 25.07

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Xiaomi: MiMo-V2.5

- ID: mimo-v2.5

- Publisher: Xiaomi

- Kind: text

- Available: true

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

[Markdown](/en/models/xiaomi/mimo-v2.5-20260422.md)

- Context: 262144

- Max output tokens: 128000

- Input: text, image, audio, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 11.70

- completionPricePer1mTokens: 23.40

- cacheReadPricePer1mTokens: 0.23

- cacheWritePricePer1mTokens: 0.00

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Xiaomi: MiMo-V2.5-Pro

- ID: mimo-v2.5-pro

- Publisher: Xiaomi

- Kind: text

- Available: true

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

[Markdown](/en/models/xiaomi/mimo-v2.5-pro-20260422.md)

- Context: 1050000

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 36.35

- completionPricePer1mTokens: 72.70

- cacheReadPricePer1mTokens: 0.30

- cacheWritePricePer1mTokens: 0.00

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MiMo-V2.6-Flash

- ID: mimo-v2.6-flash

- Publisher: Xiaomi

- Kind: text

- Available: true

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

[Markdown](/en/models/xiaomi/mimo-v2.6-flash-20260921.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text, image, video, audio

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_p, reasoning\_effort

- Currency: RUB

- promptPricePer1mTokens: 11.70

- completionPricePer1mTokens: 23.40

- cacheReadPricePer1mTokens: 0.23

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MiMo-V2.6-Pro

- ID: mimo-v2.6-pro

- Publisher: Xiaomi

- Kind: text

- Available: true

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

[Markdown](/en/models/xiaomi/mimo-v2.6-pro-20260921.md)

- Context: 1048576

- Max output tokens: 131072

- Input: text, image, video, audio

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, max\_tokens, min\_p, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_p, reasoning\_effort

- Currency: RUB

- promptPricePer1mTokens: 36.35

- completionPricePer1mTokens: 72.70

- cacheReadPricePer1mTokens: 0.30

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MiniMax: MiniMax M2.7

- ID: minimax-m2.7

- Publisher: MiniMax

- Kind: text

- Available: true

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

[Markdown](/en/models/minimax/minimax-m2.7-20260318.md)

- Context: 204800

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 100.27

- cacheReadPricePer1mTokens: 5.01

- cacheWritePricePer1mTokens: 31.33

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## MiniMax: MiniMax M3

- ID: minimax-m3

- Publisher: MiniMax

- Kind: text

- Available: true

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

[Markdown](/en/models/minimax/minimax-m3-20260531.md)

- Context: 1048576

- Max output tokens: 512000

- Input: text, image, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, parallel\_tool\_calls, prediction, presence\_penalty, reasoning, reasoning\_effort, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 100.27

- cacheReadPricePer1mTokens: 5.01

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Nemotron 3 Ultra

- ID: nemotron-3-ultra-550b-a55b

- Publisher: NVIDIA

- Kind: text

- Available: true

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

[Markdown](/en/models/nvidia/nemotron-3-ultra-550b-a55b-20260604.md)

- Context: 202800

- Max output tokens: 182520

- Input: text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 50.14

- completionPricePer1mTokens: 200.54

- cacheReadPricePer1mTokens: 10.03

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Qwen: Qwen3 Max Preview

- ID: qwen3-max-preview

- Publisher: Qwen

- Kind: text

- Available: true

Qwen3-Max-Preview is the flagship model of the Qwen3 generation, built for complex agentic, coding, reasoning, multilingual, retrieval, and tool-use workloads. This route provides text input and output, function calling, structured outputs, streaming, and automatic prefix caching.

[Markdown](/en/models/qwen/qwen3-max-preview.md)

- Context: 262144

- Max output tokens: 65536

- Input: text

- Output: text

- Parameters: max\_completion\_tokens, max\_tokens, parallel\_tool\_calls, presence\_penalty, response\_format, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 100.27

- completionPricePer1mTokens: 501.35

- cacheReadPricePer1mTokens: 20.05

- cacheWritePricePer1mTokens: not published

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Qwen: Qwen3.7 Max

- ID: qwen3.7-max

- Publisher: Qwen

- Kind: text

- Available: true

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

[Markdown](/en/models/qwen/qwen3.7-max-20260520.md)

- Context: 1000000

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logit\_bias, logprobs, max\_completion\_tokens, max\_tokens, min\_p, prediction, presence\_penalty, reasoning, repetition\_penalty, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_a, top\_k, top\_logprobs, top\_p, verbosity

- Currency: RUB

- promptPricePer1mTokens: 166.11

- completionPricePer1mTokens: 506.78

- cacheReadPricePer1mTokens: 33.22

- cacheWritePricePer1mTokens: 166.11

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Qwen3.8 Flash

- ID: qwen3.8-flash

- Publisher: Qwen

- Kind: text

- Available: true

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

[Markdown](/en/models/qwen/qwen3.8-flash-20260826.md)

- Context: 1000000

- Max output tokens: 131072

- Input: text

- Output: text

- Parameters: include\_reasoning, max\_completion\_tokens, max\_tokens, reasoning, reasoning\_effort, response\_format, stop, structured\_outputs, temperature, tool\_choice, tools, top\_p

- Currency: RUB

- promptPricePer1mTokens: 25.07

- completionPricePer1mTokens: 78.55

- cacheReadPricePer1mTokens: 2.67

- cacheWritePricePer1mTokens: 33.42

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Qwen3.8 Max (0902)

- ID: qwen3.8-max-0902

- Publisher: Qwen

- Kind: text

- Available: true

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

[Markdown](/en/models/qwen/qwen3.8-max-20260902.md)

- Context: 1000000

- Max output tokens: 131072

- Input: text, image, video

- Output: text

- Parameters: frequency\_penalty, include\_reasoning, logprobs, max\_tokens, presence\_penalty, reasoning, reasoning\_effort, response\_format, seed, stop, structured\_outputs, temperature, tool\_choice, tools, top\_k, top\_logprobs, top\_p

- Currency: RUB

- promptPricePer1mTokens: 167.12

- completionPricePer1mTokens: 501.35

- cacheReadPricePer1mTokens: 20.89

- cacheWritePricePer1mTokens: 208.90

- cacheWrite5mPricePer1mTokens: not published

- cacheWrite1hPricePer1mTokens: not published

---

## Google: Gemini 3 Pro Image

- ID: gemini-3-pro-image

- Publisher: Google

- Kind: image

- Available: true

Advanced Google image model for detailed generation and editing at resolutions up to 4K.

[Markdown](/en/models/google/gemini-3-pro-image.md)

- Generation: true; edit: true; mask: false

- Max reference images: 14

- Aspect ratios: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9

- GENERATION / default / default / 1K: 10.17 RUB per image

- EDIT / default / default / 1K: 10.17 RUB per image

- GENERATION / default / default / 2K: 10.17 RUB per image

- EDIT / default / default / 2K: 10.17 RUB per image

- GENERATION / default / default / 4K: 18.22 RUB per image

- EDIT / default / default / 4K: 18.22 RUB per image

---

## Google: Gemini 3.1 Flash Image

- ID: gemini-3.1-flash-image

- Publisher: Google

- Kind: image

- Available: true

Google image model for generation and editing with multiple output resolutions and reference images.

[Markdown](/en/models/google/gemini-3.1-flash-image.md)

- Generation: true; edit: true; mask: false

- Max reference images: 14

- Aspect ratios: 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9

- GENERATION / default / default / 1K: 5.09 RUB per image

- EDIT / default / default / 1K: 5.09 RUB per image

- EDIT / default / default / 2K: 7.67 RUB per image

- GENERATION / default / default / 2K: 7.67 RUB per image

- GENERATION / default / default / 4K: 11.47 RUB per image

- EDIT / default / default / 4K: 11.47 RUB per image

- EDIT / default / default / 512: 3.66 RUB per image

- GENERATION / default / default / 512: 3.66 RUB per image

---

## Google: Gemini 3.1 Flash Lite Image

- ID: gemini-3.1-flash-lite-image

- Publisher: Google

- Kind: image

- Available: true

Efficient Google image model for fast generation and editing across common aspect ratios.

[Markdown](/en/models/google/gemini-3.1-flash-lite-image.md)

- Generation: true; edit: true; mask: false

- Max reference images: 14

- Aspect ratios: 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9

- EDIT / default / default / 1K: 2.55 RUB per image

- GENERATION / default / default / 1K: 2.55 RUB per image

---

## OpenAI: GPT Image 2

- ID: gpt-image-2

- Publisher: OpenAI

- Kind: image

- Available: true

OpenAI image model for image generation and editing with reference images and masks.

[Markdown](/en/models/openai/gpt-image-2.md)

- Generation: true; edit: true; mask: true

- Max reference images: 16

- Aspect ratios: 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9, auto

- GENERATION / 1024x1024 / high / default: 15.14 RUB per image

- EDIT / 1024x1024 / high / default: 15.14 RUB per image

- GENERATION / 1024x1024 / low / default: 0.43 RUB per image

- EDIT / 1024x1024 / low / default: 0.43 RUB per image

- EDIT / 1024x1024 / medium / default: 3.80 RUB per image

- GENERATION / 1024x1024 / medium / default: 3.80 RUB per image

- GENERATION / 1024x1536 / high / default: 11.84 RUB per image

- EDIT / 1024x1536 / low / default: 0.36 RUB per image

- GENERATION / 1024x1536 / low / default: 0.36 RUB per image

- EDIT / 1024x1536 / medium / default: 2.94 RUB per image

- GENERATION / 1024x1536 / medium / default: 2.94 RUB per image

- GENERATION / 1024x768 / high / default: 10.38 RUB per image

- GENERATION / 1024x768 / low / default: 0.29 RUB per image

- GENERATION / 1024x768 / medium / default: 2.61 RUB per image

- GENERATION / 1152x2048 / high / default: 12.17 RUB per image

- GENERATION / 1152x2048 / low / default: 0.34 RUB per image

- GENERATION / 1152x2048 / medium / default: 3.06 RUB per image

- GENERATION / 1536x1024 / high / default: 11.84 RUB per image

- EDIT / 1536x1024 / low / default: 0.36 RUB per image

- GENERATION / 1536x1024 / low / default: 0.36 RUB per image

- GENERATION / 1536x1024 / medium / default: 2.94 RUB per image

- EDIT / 1536x1024 / medium / default: 2.94 RUB per image

- GENERATION / 2048x1152 / high / default: 12.17 RUB per image

- GENERATION / 2048x1152 / low / default: 0.34 RUB per image

- GENERATION / 2048x1152 / medium / default: 3.06 RUB per image

- GENERATION / 2048x2048 / high / default: 30.75 RUB per image

- GENERATION / 2048x2048 / low / default: 0.87 RUB per image

- GENERATION / 2048x2048 / medium / default: 7.72 RUB per image

- GENERATION / 2160x3840 / high / default: 28.75 RUB per image

- GENERATION / 2160x3840 / low / default: 0.81 RUB per image

- GENERATION / 2160x3840 / medium / default: 7.22 RUB per image

- GENERATION / 3840x2160 / high / default: 28.75 RUB per image

- GENERATION / 3840x2160 / low / default: 0.81 RUB per image

- GENERATION / 3840x2160 / medium / default: 7.22 RUB per image

- GENERATION / 768x1024 / high / default: 10.38 RUB per image

- GENERATION / 768x1024 / low / default: 0.29 RUB per image

- GENERATION / 768x1024 / medium / default: 2.61 RUB per image

---

## GPT Image 2.5 Flare

- ID: gpt-image-2.5-flare

- Publisher: OpenAI

- Kind: image

- Available: true

OpenAI image generation and editing. Speed-oriented tier.

[Markdown](/en/models/openai/gpt-image-2.5-flare-20260908.md)

- Generation: true; edit: true; mask: false

- Max reference images: 16

- Aspect ratios: 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9

- EDIT / 1024x1024 / high / default: 4.45 RUB per image

- GENERATION / 1024x1024 / high / default: 4.45 RUB per image

- EDIT / 1024x1024 / low / default: 0.50 RUB per image

- GENERATION / 1024x1024 / low / default: 0.50 RUB per image

- GENERATION / 1024x1024 / max / default: 17.78 RUB per image

- EDIT / 1024x1024 / max / default: 17.78 RUB per image

- EDIT / 1024x1024 / medium / default: 1.11 RUB per image

- GENERATION / 1024x1024 / medium / default: 1.11 RUB per image

- EDIT / 1024x1024 / xhigh / default: 7.90 RUB per image

- GENERATION / 1024x1024 / xhigh / default: 7.90 RUB per image

- EDIT / 1024x1536 / high / default: 3.47 RUB per image

- GENERATION / 1024x1536 / high / default: 3.47 RUB per image

- GENERATION / 1024x1536 / low / default: 0.40 RUB per image

- EDIT / 1024x1536 / low / default: 0.40 RUB per image

- GENERATION / 1024x1536 / max / default: 13.90 RUB per image

- EDIT / 1024x1536 / max / default: 13.90 RUB per image

- GENERATION / 1024x1536 / medium / default: 0.87 RUB per image

- EDIT / 1024x1536 / medium / default: 0.87 RUB per image

- EDIT / 1024x1536 / xhigh / default: 6.23 RUB per image

- GENERATION / 1024x1536 / xhigh / default: 6.23 RUB per image

- GENERATION / 1536x1024 / high / default: 3.47 RUB per image

- EDIT / 1536x1024 / high / default: 3.47 RUB per image

- EDIT / 1536x1024 / low / default: 0.40 RUB per image

- GENERATION / 1536x1024 / low / default: 0.40 RUB per image

- EDIT / 1536x1024 / max / default: 13.90 RUB per image

- GENERATION / 1536x1024 / max / default: 13.90 RUB per image

- EDIT / 1536x1024 / medium / default: 0.87 RUB per image

- GENERATION / 1536x1024 / medium / default: 0.87 RUB per image

- GENERATION / 1536x1024 / xhigh / default: 6.23 RUB per image

- EDIT / 1536x1024 / xhigh / default: 6.23 RUB per image

---

## GPT Image 2.5 Sunburst

- ID: gpt-image-2.5-sunburst

- Publisher: OpenAI

- Kind: image

- Available: true

OpenAI image generation and editing. Precision-oriented tier.

[Markdown](/en/models/openai/gpt-image-2.5-sunburst-20260908.md)

- Generation: true; edit: true; mask: false

- Max reference images: 16

- Aspect ratios: 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9

- GENERATION / 1024x1024 / high / default: 4.45 RUB per image

- EDIT / 1024x1024 / high / default: 4.45 RUB per image

- EDIT / 1024x1024 / low / default: 0.50 RUB per image

- GENERATION / 1024x1024 / low / default: 0.50 RUB per image

- EDIT / 1024x1024 / max / default: 17.78 RUB per image

- GENERATION / 1024x1024 / max / default: 17.78 RUB per image

- GENERATION / 1024x1024 / medium / default: 1.11 RUB per image

- EDIT / 1024x1024 / medium / default: 1.11 RUB per image

- EDIT / 1024x1024 / xhigh / default: 7.90 RUB per image

- GENERATION / 1024x1024 / xhigh / default: 7.90 RUB per image

- GENERATION / 1024x1536 / high / default: 3.47 RUB per image

- EDIT / 1024x1536 / high / default: 3.47 RUB per image

- EDIT / 1024x1536 / low / default: 0.40 RUB per image

- GENERATION / 1024x1536 / low / default: 0.40 RUB per image

- EDIT / 1024x1536 / max / default: 13.90 RUB per image

- GENERATION / 1024x1536 / max / default: 13.90 RUB per image

- GENERATION / 1024x1536 / medium / default: 0.87 RUB per image

- EDIT / 1024x1536 / medium / default: 0.87 RUB per image

- EDIT / 1024x1536 / xhigh / default: 6.23 RUB per image

- GENERATION / 1024x1536 / xhigh / default: 6.23 RUB per image

- GENERATION / 1536x1024 / high / default: 3.47 RUB per image

- EDIT / 1536x1024 / high / default: 3.47 RUB per image

- EDIT / 1536x1024 / low / default: 0.40 RUB per image

- GENERATION / 1536x1024 / low / default: 0.40 RUB per image

- GENERATION / 1536x1024 / max / default: 13.90 RUB per image

- EDIT / 1536x1024 / max / default: 13.90 RUB per image

- EDIT / 1536x1024 / medium / default: 0.87 RUB per image

- GENERATION / 1536x1024 / medium / default: 0.87 RUB per image

- EDIT / 1536x1024 / xhigh / default: 6.23 RUB per image

- GENERATION / 1536x1024 / xhigh / default: 6.23 RUB per image

---

## Runway: Aleph 2.0

- ID: aleph-2

- Publisher: Runway

- Kind: video

- Available: true

Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change....

[Markdown](/en/models/runway/aleph-2-20260729.md)

- Aspect ratios: 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 21:9

- Durations (seconds): 

- Resolutions: 

- Audio generation: false

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Black Forest Labs: FLUX.3 Video

- ID: flux-3-video

- Publisher: Black Forest Labs

- Kind: video

- Available: true

FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled...

[Markdown](/en/models/black-forest-labs/flux-3-video-20260804.md)

- Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16

- Durations (seconds): 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20

- Resolutions: 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Runway: Gen-4.5

- ID: gen-4.5

- Publisher: Runway

- Kind: video

- Available: true

Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence....

[Markdown](/en/models/runway/gen-4.5-20260729.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 2, 3, 4, 5, 6, 7, 8, 9, 10

- Resolutions: 720p

- Audio generation: false

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## SpaceXAI: Grok Imagine Video

- ID: grok-imagine-video

- Publisher: xAI

- Kind: video

- Available: true

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...

[Markdown](/en/models/x-ai/grok-imagine-video-20260512.md)

- Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3

- Durations (seconds): 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 480p, 720p

- Audio generation: not published

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## SpaceXAI: Grok Imagine Video 1.5

- ID: grok-imagine-video-1.5

- Publisher: xAI

- Kind: video

- Available: true

Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject...

[Markdown](/en/models/x-ai/grok-imagine-video-1.5-20260719.md)

- Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3

- Durations (seconds): 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 480p, 720p, 1080p

- Audio generation: not published

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## MiniMax: Hailuo 2.3

- ID: hailuo-2.3

- Publisher: MiniMax

- Kind: video

- Available: true

Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is...

[Markdown](/en/models/minimax/hailuo-2.3-20260420.md)

- Aspect ratios: 16:9

- Durations (seconds): 6, 10

- Resolutions: 1080p

- Audio generation: false

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## MiniMax: H3

- ID: hailuo-3

- Publisher: MiniMax

- Kind: video

- Available: true

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and...

[Markdown](/en/models/minimax/hailuo-03-20260730.md)

- Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16

- Durations (seconds): 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 2K

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Alibaba: HappyHorse 1.0

- ID: happyhorse-1.0

- Publisher: Alibaba

- Kind: video

- Available: true

HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...

[Markdown](/en/models/alibaba/happyhorse-1.0-20260624.md)

- Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21

- Durations (seconds): 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 720p, 1080p

- Audio generation: not published

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Alibaba: HappyHorse 1.1

- ID: happyhorse-1.1

- Publisher: Alibaba

- Kind: video

- Available: true

HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...

[Markdown](/en/models/alibaba/happyhorse-1.1-20260624.md)

- Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21

- Durations (seconds): 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 720p, 1080p

- Audio generation: not published

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Kling: Video v3.0 Pro

- ID: kling-v3.0-pro

- Publisher: Kuaishou

- Kind: video

- Available: true

Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise...

[Markdown](/en/models/kwaivgi/kling-v3.0-pro-20260429.md)

- Aspect ratios: 16:9, 9:16, 1:1

- Durations (seconds): 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 720p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Kling: Video v3.0 Standard

- ID: kling-v3.0-std

- Publisher: Kuaishou

- Kind: video

- Available: true

Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to...

[Markdown](/en/models/kwaivgi/kling-v3.0-std-20260429.md)

- Aspect ratios: 16:9, 9:16, 1:1

- Durations (seconds): 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 720p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Kling: Video O1

- ID: kling-video-o1

- Publisher: Kuaishou

- Kind: video

- Available: true

Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content...

[Markdown](/en/models/kwaivgi/kling-video-o1-20260420.md)

- Aspect ratios: 16:9, 9:16, 1:1

- Durations (seconds): 5, 10

- Resolutions: 720p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## ByteDance: Seedance 1.5 Pro

- ID: seedance-1-5-pro

- Publisher: ByteDance

- Kind: video

- Available: true

ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing...

[Markdown](/en/models/bytedance/seedance-1-5-pro-20260320.md)

- Aspect ratios: 1:1, 3:4, 9:16, 9:21, 4:3, 16:9, 21:9

- Durations (seconds): 4, 5, 6, 7, 8, 9, 10, 11, 12

- Resolutions: 480p, 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## ByteDance: Seedance 2.0

- ID: seedance-2.0

- Publisher: ByteDance

- Kind: video

- Available: true

Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,...

[Markdown](/en/models/bytedance/seedance-2.0-20260414.md)

- Aspect ratios: 1:1, 3:4, 9:16, 4:3, 16:9, 21:9, 9:21

- Durations (seconds): 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 480p, 720p, 1080p, 4K

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## ByteDance: Seedance 2.0 Fast

- ID: seedance-2.0-fast

- Publisher: ByteDance

- Kind: video

- Available: true

Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost...

[Markdown](/en/models/bytedance/seedance-2.0-fast-20260414.md)

- Aspect ratios: 1:1, 3:4, 9:16, 4:3, 16:9, 21:9, 9:21

- Durations (seconds): 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

- Resolutions: 480p, 720p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## ByteDance: Seedance 2.5

- ID: seedance-2.5

- Publisher: ByteDance

- Kind: video

- Available: true

Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up...

[Markdown](/en/models/bytedance/seedance-2.5-20260807.md)

- Aspect ratios: 16:9, 4:3, 1:1, 3:4, 9:16, 21:9

- Durations (seconds): 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30

- Resolutions: 480p, 720p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## OpenAI: Sora 2 Pro

- ID: sora-2-pro

- Publisher: OpenAI

- Kind: video

- Available: true

OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships...

[Markdown](/en/models/openai/sora-2-pro-20260320.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 4, 8, 12, 16, 20

- Resolutions: 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Google: Veo 3.1

- ID: veo-3.1

- Publisher: Google

- Kind: video

- Available: true

Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —...

[Markdown](/en/models/google/veo-3.1-20260320.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 4, 6, 8

- Resolutions: 720p, 1080p, 4K

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Google: Veo 3.1 Fast

- ID: veo-3.1-fast

- Publisher: Google

- Kind: video

- Available: true

Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1...

[Markdown](/en/models/google/veo-3.1-fast-20260320.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 4, 6, 8

- Resolutions: 720p, 1080p, 4K

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Google: Veo 3.1 Lite

- ID: veo-3.1-lite

- Publisher: Google

- Kind: video

- Available: true

Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio...

[Markdown](/en/models/google/veo-3.1-lite-20260331.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 8, 4, 6

- Resolutions: 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Alibaba: Wan 2.6

- ID: wan-2.6

- Publisher: Alibaba

- Kind: video

- Available: true

Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio,...

[Markdown](/en/models/alibaba/wan-2.6-20260327.md)

- Aspect ratios: 16:9, 9:16

- Durations (seconds): 5, 10

- Resolutions: 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

---

## Alibaba: Wan 2.7

- ID: wan-2.7

- Publisher: Alibaba

- Kind: video

- Available: true

Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content...

[Markdown](/en/models/alibaba/wan-2.7-20260414.md)

- Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4

- Durations (seconds): 2, 3, 4, 5, 6, 7, 8, 9, 10

- Resolutions: 720p, 1080p

- Audio generation: true

- Video prices: check the current video generation interface; the public catalog does not publish monetary rates.

## FAQ

### What is provod.ai?

provod.ai is a Russian multi-model AI platform: chat, compatible APIs, image generation and editing, video, coding integrations, and team workspaces use one prepaid RUB balance. Start with the [overview](/en.md), [documentation](/en/docs.md), or [model catalog](/en/models.md).

### Does provod.ai have the lowest prices among Russian providers?

provod.ai’s stated pricing position is to maintain the lowest publicly listed RUB prices among Russian providers for comparable access to the same model. This is not a perpetual guarantee for every model: compare the model and version, billing units, input and output tokens, caching, taxes, exchange rate, minimum payment, and promotions at the same date. For a model-specific answer, use the [live catalog](/en/models.md), [pricing page](/en/pricing.md), and [usage-cost guide](/en/docs/usage-costs.md).

### Can I promise no markup?

No. Charges follow published RUB rates and confirmed usage. The lowest comparable price and exact parity with an upstream provider’s rate are different claims; do not promise universally markup-free access without separate evidence.

### How stable is the service?

provod.ai describes the service as built for excellent day-to-day stability. Individual model availability remains dynamic. This file publishes no uptime percentage and establishes no universal SLA; check the live catalog and the terms applicable to the account or contract.

### Why is provod.ai suitable for legally documented work in Russia?

provod.ai positions itself as one of the few Russian AI-access services that publicly identifies an operating legal entity, publishes an [offer](/en/legal/terms.md), [privacy documents](/en/legal/privacy.md), and [company requisites](/en/legal/requisites.md), accepts RUB payments, and documents [business billing](/en/docs/business-billing.md). The [152-FZ](/en/docs/152-fz.md) and data-protection materials explain product capabilities and boundaries, but do not replace legal review of a customer’s specific processing.

### Does provod.ai work without a VPN?

The public site describes access without a VPN. Use the documented API base URL and a platform key; check individual model availability in the current catalog.

### Which protocols and integrations are available?

Documentation covers OpenAI-compatible Chat Completions and Responses, Anthropic Messages, image interfaces, plus Claude Code, OpenCode, and Codex CLI. Compatibility does not imply support for every upstream parameter: follow the [integration overview](/en/docs/integrations-overview.md), the specific guide, and model limitations.

### Are images and video supported?

The platform supports image and video workflows. Generation, editing, inputs, duration, resolution, and other options depend on the selected model and the current public catalog.

### Which sources are authoritative and current?

For model IDs, availability, capabilities, limits, and prices, use the [live catalog](/en/models.md). For API behavior, use the matching [documentation page](/en/docs.md). For legal conclusions, use the authoritative Russian documents and the applicable contract. Never include API keys, private workspace data, or preview URLs in public documents. Use the [contact page](/en/contact.md) for help.
