Scenario
← All Models
Audio

Seed Audio 1.0 Multilingual

Use it ↗

Multilingual text-to-speech in 20 languages with code-switching, timestamp control to hit exact durations, subtitle timing output, plus audio or image voice references.

Seed Audio 1.0 Multilingual by ByteDance turns text into expressive speech across 20 languages, auto-detecting the language and code-switching between them in a single generation. Define the voice in plain language, clone it from up to three audio clips, or derive it from a reference image. Precise time control sets per-sentence timestamps to land a read on an exact duration, and every generation can return per-word and per-sentence timing data for subtitles. Fine-tune speech rate, pitch, loudness, and sample rate.

More models from Bytedance