Create connected videos from a prompt, a first and last frame, or a complete set of visual and audio references. The Seedance 2.5 AI video generator combines text to video AI, image to video AI, multi-shot direction, and synchronized sound in one production-focused workflow.
Seedance 2.5 is ByteDance's multimodal video model for longer, more controllable audiovisual creation. It supports text to video AI, image to video AI with optional first and last frames, and reference-led generation using images, video clips, and audio tracks.
One request can produce up to a 30-second AI video with several logically connected shots. This gives a scene room for an opening, development, transition, and payoff instead of holding on one composition for the entire clip.
The model announcement highlights native 4K output, while the generation controls currently available here provide 480p, 720p, and 1080p output. This distinction lets you choose only settings the connected API accepts while understanding the wider Seedance 2.5 model specification.
Combine as many as 30 images, 10 video clips, and 10 audio tracks in a multimodal video generation brief. Assign each asset a clear role for character identity, product details, visual style, scene design, camera language, motion, voice, or atmosphere.
Generate up to 30 seconds in one run and describe a deliberate sequence of shots. Longer timing supports product reveals, short narrative arcs, lifestyle scenes, trailers, and social ads with a clearer beginning and ending.
Turn a written brief into a structured sequence. Specify subjects, actions, camera position, lens behavior, lighting, pacing, transitions, dialogue, and sound so each shot serves the same creative idea.
Animate a first frame or use first and last frames together when the opening and ending composition matter. The frame-guided mode is separate from multimodal reference generation, making the intended type of control clear.
Generate video and sound together for tighter audiovisual timing. Prompts can describe dialogue, ambience, effects, and musical intent alongside visible action, while audio generation can also be switched off when a silent source is preferred.
Use product images, campaign references, and motion direction to plan product showcases, lifestyle scenes, seasonal campaigns, marketplace assets, and social creatives where clear details and consistent visual language matter.
Match the generation mode to the source material, give every reference a purpose, and describe the sequence as a series of intentional shots.
Use text to video AI for a prompt-only concept, image to video AI when the first or final composition must be anchored, or multimodal video generation when images, clips, and audio should guide different parts of the result.
Describe the subject, setting, action, shot order, camera movement, lighting, materials, dialogue, ambience, and ending. In reference mode, explain exactly what each uploaded asset should control.
Choose a 4–30 second duration, supported aspect ratio, 480p, 720p, or 1080p resolution, and whether to generate audio. Review continuity, product details, motion, sound, and transitions before refining the prompt.
Describe a shot sequence instead of one long list of visual details
Give every reference image, video, and audio file a single clear purpose
State which faces, clothing, products, logos, and colors must remain consistent
Separate camera direction, subject motion, lighting, dialogue, and sound in the prompt
Use first and last frames when exact opening and ending compositions matter most
Review generated text, product details, hands, faces, motion, and audio before publishing
Seedance 2.5 is ByteDance's multimodal AI video model for generating connected videos from text, first and last frames, or image, video, and audio references. It supports optional synchronized audio and up to 30 seconds of output through the connected workflow.
For Seedance 2.5 vs Seedance 2.0, the newer model expands video length, multimodal reference control, pre-generation planning, editing control, and the model's published output capabilities. These changes make Seedance 2.5 better suited to longer scenes, brand assets, and production-focused AI video workflows.
The reference workflow accepts up to 30 images, 10 video clips, and 10 audio tracks. Reference videos may total up to 30 seconds, and each asset should be assigned a clear role in the prompt.
Yes. The connected API supports video durations from 4 to 30 seconds. A detailed shot sequence helps the model use the extra time for connected action and transitions rather than a single lingering scene.
Yes. Audio can be generated together with the video. Describe dialogue, ambience, sound effects, or musical intent in the prompt, or disable audio generation when you need a silent clip for later post-production.
The Seedance 2.5 model announcement describes native 4K output. The current connected generator exposes 480p, 720p, and 1080p settings, so select from those available resolutions for requests made on this page.
It can support product showcases, lifestyle scenes, campaign concepts, marketplace assets, and social ads. Use clear product references, preserve important geometry and branding in the prompt, and review every output for accurate details before publishing.
Commercial use depends on the active service terms and your rights to every prompt, image, clip, audio file, logo, person, and character used as input. Review those rights and the final output before distribution.
Turn a prompt, a frame pair, or a multimodal reference set into a connected AI video with optional synchronized audio.
Plan the shots, assign each reference a purpose, choose the available output settings, and generate your first Seedance 2.5 video.
Create text-to-video, image-to-video, and multimodal AI video with Seedance 2.5.