Yandex Cloud AI Studio pricing policy
Note
Currency of Service rates (prices) depends on the company you made a contract with:
- Prices in US dollars are applicable to customers of Iron Hive doo Beograd (Serbia) or Direct Cursus Technology L.L.C. (Dubai).
- Prices in Russian roubles are applicable to customers of Yandex.Cloud LLC.
All prices in RUB and KZT are inclusive of VAT; in USD, net of VAT.
Subscriptions
Payment for AI models and services via subscription.
|
Service |
Price per subscription per user per month,without VAT |
|
Generative Models. Go Subscription |
$20.41 |
|
Generative Models. Go X5 Subscription |
$45.00 |
|
Generative Models. Go X15 Subscription |
$94.18 |
Model Gallery
The cost of using the models depends on the operating mode and the number of tokens for different consumption types:
- Input query tokens.
- Output model response tokens.
- Cached tokens, if certain information is re-used without additional computation, such as instructions for a model.
- Tool tokens provided to the model as a result of invoking any tool.
Caching is enabled automatically where possible and applicable. Caching is not guaranteed and does not apply to output tokens.
Tool tokens include all uncached tokens stored in the message history at the time the tool's results were transmitted. Tool tokens are calculated only for AI Studio built-in tools and do not apply to the results of custom functions. Use of tools is charged separately.
Synchronous mode
|
Model |
Price per 1,000 input tokens,without VAT |
Price per 1,000 cached tokens,without VAT |
Price per 1,000 tool tokens,without VAT |
Price per 1,000 output tokens,without VAT |
|
Alice AI LLM |
$0.00409836 |
$0.00409836 |
$0.0010655736 |
$0.009836064 |
|
YandexGPT Pro 5.1 |
$0.006557376 |
$0.006557376 |
$0.001639344 |
$0.006557376 |
|
YandexGPT Pro 5 |
$0.009836064 |
$0.009836064 |
$0.009836064 |
$0.009836064 |
|
YandexGPT Lite |
$0.001639344 |
$0.001639344 |
$0.001639344 |
$0.001639344 |
|
Alice AI LLM Flash |
$0.000819672 |
$0.000204918 |
$0.000204918 |
$0.001639344 |
|
DeepSeek-V4.1-Flash |
$0.002459016 |
$0.000614754 |
$0.000614754 |
$0.00409836 |
|
DeepSeek V4 Flash |
$0.002459016 |
$0.000614754 |
$0.000614754 |
$0.00409836 |
|
DeepSeek-V4.1-Flash |
$0.002459016 |
$0.000614754 |
$0.000614754 |
$0.00409836 |
|
Qwen3 235B |
$0.00409836 |
$0.00409836 |
$0.00409836 |
$0.00409836 |
|
gpt-oss-120b |
$0.002459016 |
$0.002459016 |
$0.002459016 |
$0.002459016 |
|
gpt-oss-20b |
$0.000819672 |
$0.000819672 |
$0.000819672 |
$0.000819672 |
|
Qwen3.6 35B |
$0.001639344 |
$0.000409836 |
$0.000409836 |
$0.002459016 |
|
Speech Realtime 260528 |
$0.000819672 |
$0.000204918 |
$0.000204918 |
$0.001639344 |
|
Speech Realtime 250923 |
$0.006557376 |
$0.001639344 |
$0.001639344 |
$0.006557376 |
|
Speech Realtime DeepSeek V4 Flash |
$0.002459016 |
$0.000614754 |
$0.000614754 |
$0.00409836 |
Asynchronous mode
|
Model |
Price per 1,000 input tokens,without VAT |
Price per 1,000 output tokens,without VAT |
|
Alice AI LLM |
$0.00204918 |
$0.0083606544 |
|
YandexGPT Pro 5.1 |
$0.0033606552 |
$0.0033606552 |
|
YandexGPT Pro 5 |
$0.0049999992 |
$0.0049999992 |
|
YandexGPT Lite |
$0.000819672 |
$0.000819672 |
Batch mode
With models in batch mode, the minimum cost per run is 200,000 tokens.
|
Model |
Price per 1,000 input tokens, without VAT |
Price per 1,000 output tokens, without VAT |
|
Qwen2.5 7B Instruct |
$0.000819672 |
$0.000819672 |
|
Qwen2.5 72B Instruct |
$0.0049999992 |
$0.0049999992 |
|
QwQ 32B Instruct |
$0.0033606552 |
$0.0033606552 |
|
Llama-3.3-70B-Instruct |
$0.0049999992 |
$0.0049999992 |
|
Llama-3.1-70B-Instruct |
$0.0049999992 |
$0.0049999992 |
|
DeepSeek-R1-Distill-Llama-70B |
$0.0049999992 |
$0.0049999992 |
|
Qwen2.5 32B Instruct |
$0.0033606552 |
$0.0033606552 |
|
DeepSeek-R1-Distill-Qwen-32B |
$0.0033606552 |
$0.0033606552 |
|
phi-4 |
$0.001639344 |
$0.001639344 |
|
Qwen2 VL 7B |
$0.000819672 |
$0.000819672 |
|
Qwen2.5 VL 7B |
$0.000819672 |
$0.000819672 |
|
DeepSeek 2 VL |
$0.0033606552 |
$0.0033606552 |
|
DeepSeek 2 VL Tiny |
$0.000819672 |
$0.000819672 |
|
Gemma3 1B it |
$0.000819672 |
$0.000819672 |
|
Gemma3 4B it |
$0.000819672 |
$0.000819672 |
|
Gemma3 12B it |
$0.001639344 |
$0.001639344 |
|
Gemma3 27B it |
$0.0033606552 |
$0.0033606552 |
|
Qwen 2.5 VL 32B Instruct |
$0.0033606552 |
$0.0033606552 |
|
Qwen3-0.6B |
$0.000819672 |
$0.000819672 |
|
Qwen3-1.7B |
$0.000819672 |
$0.000819672 |
|
Qwen3-4B |
$0.000819672 |
$0.000819672 |
|
Qwen3-8B |
$0.000819672 |
$0.000819672 |
|
Qwen3-14B |
$0.001639344 |
$0.001639344 |
|
Qwen3-32B |
$0.0033606552 |
$0.0033606552 |
|
Qwen3-30B-A3B |
$0.0033606552 |
$0.0033606552 |
|
Qwen3-235B-A22B |
$0.049999992 |
$0.049999992 |
Dedicated instances
The cost of operation of a dedicated instance depends on the model and selected configuration. Dedicated instances are charged per second with rounding up to a billing unit. However, there is no charge for hardware maintenance and model deployment time.
Prices are shown for 1 hour of use. Billing occurs per second.
The price per 1 unit for a dedicated instance is $0.0083327856 without VAT.
| Model | Price per 1 hour,S configuration, without VAT |
Price per 1 hour,M configuration, without VAT |
Price per 1 hourL configuration, without VAT |
|---|---|---|---|
| Qwen 2.5 VL 32B Instruct | $6.70 | $13.40 | $20.10 |
| Qwen 2.5 7B Instruct | $6.70 | $13.40 | $20.10 |
| Gemma 3 4B it | $3.35 | $6.70 | $10.05 |
| Gemma 3 12B it | $3.35 | $6.70 | $10.05 |
| T-pro-it-2.0-FP8 | $6.20 | $12.40 | $18.60 |
Model fine-tuning
At the Preview stage, you can fine-tune models free of charge. A fine-tuned YandexGPT Lite model will cost the same as the basic YandexGPT Lite model.
Text tokenization
The use of tokenizer (TokenizerService calls and Tokenizer methods) is free of charge.
Text vectorization
The cost of text vectorization (getting text embeddings) depends on the size of the text submitted for vectorization. Yandex Cloud Billing breaks down the creation of embeddings in vectorization units. One unit equals one token.
| Model | Price per 1,000 tokens, without VAT |
|---|---|
| Embeddings | $0.0000827869 |
Example of cost calculation for text vectorization
The cost of vectorizing a text of 2,000 tokens will be:
- $0.0000827869: Cost of processing 1,000 tokens.
- $0.0000827869 / 1,000: Cost of processing one token.
2,000 × ($0.0000827869 / 1,000) = $0.0001655738
Total: $0.0001655738.
Text classifications
The cost of text classification depends on the classification model you use and the number of tokens you provide.
- When classifying with YandexGPT Lite, a billing unit is a request of up to 1,000 tokens.
- When classifying with YandexGPT Pro and fine-tuned classifiers, a billing unit is a request of up to 250 tokens.
Requests with less than one billing unit are rounded up to the next integer. Large texts are billed as multiple requests with rounding up.
For example, classifying a text of 770 tokens with YandexGPT Lite will be billed as a single request, i.e., as one billing unit.
The same 770-token text classified with YandexGPT Pro or a fine-tuned classifier will be billed as four requests.
| Service | Price, without VAT |
|---|---|
| 1 request (1,000 tokens) to classifier based on YandexGPT Lite | $0.0012499998 |
| 1 request (250 tokens) to classifier based on YandexGPT Pro | $0.0012499998 |
| 1 request (250 tokens) to tuned classifier | $0.0012499998 |
Image generation
The use of image generation models is charged for each generation request. Requests are not idempotent; therefore, two requests with the same settings and generation prompt are considered as two separate requests.
| Service | Price, without VAT |
|---|---|
| 1 request for image generation | $0.0182786856 |
Image recognition
Each successful request for image recognition performed using any recognition model is charged as a single pricing unit:
- If your request includes multiple images or a PDF file consisting of multiple pages, each image or page will be charged separately.
- If you send two requests to recognize text on the same image, you will be charged two billing units. This makes sense when a text is written in languages from different language models (for example, Arabic and Hebrew).
- Only successful analysis attempts are chargeable. You will not be charged if the server returned an error or the request configuration was incorrect.
| Service | Price per unit, without VAT |
|---|---|
| Printed text recognition | $0.0010827867 |
| Table recognition | $0.0099999984 |
| Document recognition (passport) | $0.0058196712 |
| Document recognition (driving license) | $0.0058196712 |
| Document recognition (vehicle registration certificate) | $0.0058196712 |
| Handwriting recognition | $0.0124590144 |
| License plate recognition | $0.0010827867 |
Agent Atelier
Voice agents
The cost of using voice agents consists of the following:
- Cost of speech recognition (incoming audio).
- Cost of speech synthesis (outgoing audio).
- Cost of text generation using the Speech Realtime model.
- Cost of tool invocation.
| Service | Price per unit of tariffing, without VAT |
|---|---|
| Incoming audio, per 1 second | $0.0002163934 |
| Outgoing audio, per 1 second | $0.0001663934 |
Example of cost calculation for a voice agent
-
Voice agent use case: support bot with knowledge base search.
Dialog with support bot:
Client:
Hi there!Bot (answers as per the instruction without access to the knowledge base):
Hi there! What can I do for you?Client:
I can't log into my personal account.Bot (answers as per the instruction without access to the knowledge base):
I see. Just to clarify: is there an error message, or does login just not work?Client:
It says the password is incorrect.Bot (answers as per the instruction without access to the knowledge base):
Alright. In this case, first try to recover your password using the Forgot your password? button on the login page.Client:
I tried, but I didn't receive the email. What should I do?Bot (answer generated based on the knowledge base):
I will now instruct you how to proceed in this situation.If you are not getting the email, follow these steps:
- Check your Spam folder.
- Make sure you are entering the email address linked to your account.
- Wait for 5 to 10 minutes and try again.
- If the email has not arrived anyway, contact support via the form on the website.
-
Session duration: 60 seconds
-
Incoming audio (bot listens to client): 60 seconds
-
Outgoing audio (bot responds to client): 20 seconds
-
Calling the file search tool: 1 request
-
Tokens spent:
Prompt / instructions
1,500
Answers without knowledge base search
1,500
Answer with knowledge base search
1,000
Document from the knowledge base
3,000
Total
7,000
Speech recognition cost calculation:
$0.0002163934 × 60 = $0.012983604
Total: $0.012983604: Cost per 60 seconds of speech recognition.
Where:
- $0.0002163934: Cost of processing 1 second of incoming audio.
Speech synthesis cost calculation:
$0.0001663934 × 20 = $0.003327868
Total: $0.003327868: Cost per 20 seconds of speech synthesis.
Where:
- $0.0001663934: Cost of synthesizing 1 second of outgoing audio.
Text processing cost calculation:
((1,500 + 3,000) / 1,000 × $0.006557376) + ((1,500 + 1,000) / 1,000 × $0.006557376) = $0.045901632
Total: $0.045901632: Text processing cost.
Where:
- 1,500: Tokens spent on reading the prompt.
- 3,000: Tokens spent on reading instructions from the knowledge base.
- $0.006557376: Cost per 1,000 incoming tokens.
- 1,500: Tokens spent on answers without knowledge base search.
- 1,000: Tokens spent on answers with knowledge base search.
- $0.006557376: Cost per 1,000 outgoing tokens.
Calculating the cost of calling the file search tool:
$2.459016 / 1,000 × 1 = $0.002459016
Total: $0.002459016: Cost per 1 call of the file search tool.
Where:
- $2.459016: Cost per 1,000 calls of the file search tool.
Total cost calculation:
$0.012983604 + $0.003327868 + $0.045901632 + $0.002459016 = $0.06467212
Total: $0.06467212: Cost per 60 seconds of dialog with the support bot.
Where:
- $0.012983604: Cost per 60 seconds of speech recognition.
- $0.003327868: Cost per 20 seconds of speech synthesis.
- $0.045901632: Cost of processing 7,000 tokens.
- $0.002459016: Cost per 1 call of the file search tool.
Text-based agents
The cost of using text-based agents consists of the following:
- Consumption of tokens as per the pricing plans of the Model Gallery models.
- Cost of tool invocation.
Invoking tools in agents
| Service | Price per 1,000 requests, without VAT |
|---|---|
| Web Search tool | $7.4999988 |
| File Search tool | $2.459016 |
| Code Interpreter tool | Free of charge |
| MCP tool | Free |
| Image Generation tool | $18.2786856 |
AI Search
The search index size is rounded up to the nearest whole gigabyte.
Searching within a search index without using an agent is charged as calling the file search tool.
| Service | Price per day per 1 GB, without VAT |
|---|---|
| Search index storage | $0.086885232 |
| AI Studio file storage | Free of charge |
MCP Hub
Storing MCP servers is free of charge. However, you may still be charged for tools created in MCP servers, such as Yandex Cloud Functions invocations.
When using external APIs, such as Kontur.Focus or amoCRM, you are charged directly by our respective partner.
Workflows
At the Preview stage, Workflows is free of charge.
Internal server errors
You are not charged for a request that failed due to an internal server error.