Skip to main content
Seedance 2.0 is BytePlus / Doubao’s professional multimodal video creation model. It accepts images, videos, audio, and text as reference inputs, supports video editing and video extension, and can also generate videos from explicit start and end frames. The model is positioned around motion stability, visual detail retention, camera control, and character consistency across complex scenes. It is intended to preserve material, lighting, tone, and shot language more faithfully than a single-input workflow. Typical use cases include advertising production, film-style content creation, and social media campaign assets where controllability and production-ready output matter more than raw speed.

Model ID

Pass this exact public ID in the top-level model field. Do not substitute an upstream provider alias.

Supported modes

  • text_to_video
  • image_first_frame
  • image_first_last_frame
  • multimodal_reference
  • video_modify
  • video_extend

Inputs

  • Image
  • Video
  • Audio
  • Text

Outputs

  • Video

Capabilities

  • Multimodality-to-video
  • Video editing
  • Video extension

Technical details

Availability and final customer prices can change. Check the Dashboard model catalog before production use.
See the API reference overview for authentication, request conventions, and response schemas.