Seedance 2.5
Use it ↗Seedance 2.5 by ByteDance: one model for text-to-video, first/last frame, reference-to-video, editing, and extension, up to 720p and 4 to 30 seconds. The mode is inferred from your inputs and prompt, so you just describe the shot.
Seedance 2.5 is ByteDance's multimodal video model. Text-to-video, image-to-video (first frame, or first and last frame), reference-to-video, video editing, and video extension all live in one model. The mode is inferred from what you give it: a prompt alone drives text-to-video, a first frame switches to image-to-video, reference images or videos lock identity, product, world, or motion, and a source video enables editing or extension.
Outputs run up to 720p, from 4 to 30 seconds, in the aspect ratio you pick. Write the prompt as a compact director's brief: subject and action first, then scene, visual style, one camera move, and audio. Tag reference uploads in the prompt with @image1, @video1, and @audio1, and always state both what to use and what to ignore from each reference.
Audio: Generate Audio is off by default. For controlled, moderation-safe sound, generate or bring your own audio first and pass it in Reference Audio, instead of turning on Generate Audio (which sends the native track through moderation after the video is already made). Supplying the track yourself also lets you drive lip-sync and rhythm from audio you already trust.