One workspace for multilingual audio creation
Text mode supports automatic and official voices.
Choose a target duration and the model will aim to match it
Your latest and previous audio will appear here
Choose a method, write a description, and start creating
Turn a description into a complete audio scene
Generate dialogue, performance, ambience, effects and music together.
Use visual mood or existing audio to guide a new result.
Let the model create a voice, or pick from the multilingual official voices.
Start with text, an image, or reference audio and create a complete scene with dialogue, ambience, sound effects, and background music.
Select text-to-audio, image-reference audio, or audio-reference generation. Text mode supports automatic and official voices.
Describe the dialogue language, performance, ambience, sound effects, and background music. The multilingual model detects the language and also supports natural-language timeline instructions.
Submit your request to generate the complete audio. Preview it online, download the result, and reopen previous generations from your creation history.