Yandex SpeechKit release notes: Speech synthesis

SpeechKit provides updates by model and version.

For more information about the voice models, see About the technology.

Release as of 21/01/2026

Added support for SpeechKit synthesis in Yandex Cloud ML SDK. For more information, see this ML SDK guide and the project's GitHub repository.

Release as of 28/05/25

Added the new option of creating a unique voice, SpeechKit Brand Voice Lite. For more information, see Yandex SpeechKit Brand Voice.

Release as of 30/04/25

Renamed the lola and lola_ru voices to zamira and zamira_ru.

Release as of 29/04/25

Added the new roles for the following voices: zamira, zamira_ru, and yulduz_ru.

Release as of 28/03/25

Added the Russian equivalents for the zamira and yulduz voices, zamira_ru and yulduz_ru, with multiple roles. Additional roles are now available for the following voices: saule, saule_ru, zhanar, zhanar_ru, and yulduz. For the full list of available voices, see List of voices.

Release as of 24/01/25

Added new voices. For synthesis in Kazakh, the zhanar female voice is now available. For synthesis in Uzbek, added the zamira and yulduz female voices.

Release as of 18/11/24

Fixed the pronunciation of tenge for synthesis in Russian. Now the model pronounces it with a soft (palatalized) t: [tʲɪnˈɡʲe].`

Release as of 10/10/24

  1. Added the new female voice for synthesis in Kazakh, saule, and its Russian equivalent, saule_ru.
  2. Renamed the madirus Russian voice to madi_ru. The voice is still available by its old name, but please use the new one where possible.

Release as of 20/09/24

  • Improved the quality for the filipp, ermil, and zahar voices.
  • Optimized the normalizer for Kazakh and Uzbek.

Release as of 09/09/24

Improved the interrogative intonation and general synthesis quality for all publicly available Russian voices.

Release as of 11/07/24

For synthesis in Russian:

  • Reduced the ambient noise level.
  • Fixed the accentuation in certain words.

Release as of 15/04/24

Fixed the bug of synthesizing speech at too fast a rate.

Release as of 09/04/24

In the API v1, marina is now the default voice.

Release as of 03/04/24

Changed the default voice in the API v3. All synthesis projects without an explicitly specified voice will now use the marina voice.

Release as of 20/02/24

Improved the quality for the masha, marina, anton, alexander, dasha, and julia voices.

Release as of 06/02/24

Added the REST API v3 support.

Release as of 10/01/24

  1. Added support for the cardinal number normalization in English. Normalization only works for positive integers. Ordinal numbers are not supported.
  2. Added the DurationHint to the API to use to specify the minimum and maximum time spent on synthesizing the text.
  3. Added the text_chunk, start_ms, and length_ms fields to the UtteranceSynthesisResponse message. These fields store text information, as well as the start and end time of the audio within the provided fragment.

Release as of 05/12/23

Improved the quality of speech synthesis for all languages except Russian.

Release as of 23/10/23

  1. A new voice, masha, is now available in three roles.
  2. Additional roles are now available for Russian-language voices.
  3. Optimized the normalizer for Kazakh.
  4. Improved the pronunciation of SMS for Kazakh and Uzbek.

Release as of 27/07/23

  1. Added the pitch_shift parameter to the API v3. You can use it to increase the pitch contour of an entire synthesized audio by a fixed value in Hz. Shifting the contour up makes a voice sound more lively.
  2. Seven new voices are now available for speech synthesis in Russian: dasha, julia, lera, marina, alexander, kirill, and anton.

Release as of 19/06/23

Improved the quality of pronunciation of car brands for Uzbek.

Release as of 08/06/23

  1. Added the normalization for cardinal numbers written in Arabic numerals for Uzbek.
  2. Improved the quality of speech synthesis for Uzbek, mainly affecting the synthesis of short texts.

Release as of 18/04/23

  1. Speech synthesis for Uzbek now supports phoneme-based transcription of texts. See the list of supported phonemes here. In addition, the Uzbek model can now automatically replace apostrophes. However, for efficient speech synthesis, you should only use the straight (ʼ) and reversed (ʻ) typographic apostrophes.
  2. For pattern-based synthesis, changed the default volume normalization. Now, if the normalization type is not set explicitly, the volume of variables is normalized based on the initial pattern.

Release as of 21/03/23

  1. Added the normalizer for Kazakh. Now the model can pronounce numbers written in Arabic numerals.

  2. Added two types of apostrophes for Uzbek: straight typographic apostrophe (ʼ) and reversed typographic apostrophe (ʻ). Now you can synthesize utterances in Uzbek written in the Latin script with these apostrophes.

    Yaʼni mana shu beret kiygan notanish odamni.
    Soʻng yana pastga qarab ketiladi.

    Warning

    Use only these two options for apostrophes. The model does not support auto replacement, and the synthesis quality strongly depends on the input quality.

Release as of 07/03/23

  1. Significantly revised the SpeechKit Brand Voice technology for creating custom voices.
  2. Added support for pauses in all languages in test mode when using the TTS markup. Please report any pausation errors by submitting a ticket to the support team. Your feedback will help us improve future releases.

Release as of 07/10/22

The general branch now has new voices and languages available for testing:

  • lea female voice: German.
  • madi male voice: Kazakh.
  • madirus male voice: Russian.
  • nigora female voice: Uzbek.

The general branch now has these new voices: amira and john.

Release as of 09/06/22

  1. Improved intonations and accentuation for all voices.

  2. Added new pausation features:

    • Fixed the error when pauses shorter than 1,200 milliseconds were disregarded in the SSML markup. Please note that pauses shorter than 700 milliseconds are considered a synthesis cue and cannot be used to control the precise duration of pauses between words.
    • SSML pauses with the x-weak, weak, and medium values have a greater impact on the synthesized text.
    • You can now place pauses when using the TTS markup. Use the <[small]> tag to set the pause length in the synthesized text, e.g., Hello, <[small]>. The possible pause lengths are tiny, small, medium, large, or huge.
  3. Discontinued support for filipp:deprecated. filipp:deprecated and filipp now sound the same.

Release as of 19/05/22

  1. Deprecated voices will no longer be supported starting May 31, 2022.

  2. The rc branch now has new voices and languages available for testing:

    • amira female voice: Kazakh.
    • john male voice: English.

    The voices are only available in the API v3 with the x-service-branch:rc header.

Release as of 30/03/22

  1. The standard voices are currently only available by the :deprecated tag and will be supported until May 31, 2022.

  2. Fixed the intonations and issues with rare artifacts in texts with many various numbers as per the CLOUDSUPPORT-138703 support ticket.

Release as of 17/03/22

  1. Added the feature of synthesizing audio files in MP3 format. It is available in the API v3 and when using premium voices in the API v1.

  2. New voices now support roles, i.e., extended emotional tones. See the emotion parameter in the API v1 and role in the API v3 for details. Different roles are available for different voices. For the full list of values, see List of voices. If you select an invalid role, the service will return an error.

  3. Fixed the quality regression in emphasis placement for alena and filipp. Improved the emphasis placement and subjective perception for all voices.

  4. Started a major update of standard voices: oksana, ermil, jane, omazh, and zahar will be replaced with oksana:rc, ermil:rc, jane:rc, omazh:rc, and zahar:rc, respectively. The update will not affect the cost of the regular voices. The existing oksana, ermil, jane, omazh, and zahar voices are available in the :deprecated branch.

Release as of 24/01/22

  1. Updated the generative model. Improved the pronunciation of numbers and abbreviations in the finance field.

  2. You can now emphasize using the markup: Are you **happy** to see me?

  3. Aligned the processing of SSML pauses and SIL tags to support integration with Yandex.Dialogs. In SSML or SIL notations, pauses in text are considered the end of an utterance, thus, the end-of-utterance intonation replaces the tag in the generated text. SSML pauses and SIL tags are supported when generating both short and long speech segments.

Release as of 16/12/21

  1. Increased the limits for API v3 requests: 250 characters or 24 seconds of audio for a synthesized utterance. Please note that request costs may increase in the future.

  2. The unsafe_mode option in the API v3 enables you to automatically split long segments of text submitted for synthesis into separate phrases.

  3. The silence after the last word is synthesized is much shorter now. An audio ends almost immediately after the final word is synthesized.

Release as of 18/11/21

  1. Introduced the fixes for stabilizing the synthesis of the alena premium voice. It now sounds consistent.
  2. Fixed the pronunciation errors for alena.
  3. Improved the pausation in the REST API.
  4. Added the new premium voices in test mode:
    • oksana:rc
    • ermil:rc
    • jane:rc
    • omazh:rc
    • zahar:rc