Write better Seedance prompts
Generate Media turns a text prompt plus optional reference images and videos into a 10-second vertical video with the Seedance 2.0 series. The model reads prompt and references together to produce consistent, on-brief results. This guide covers the essentials of writing prompts that give you stable, predictable output. For the full model capabilities, see the official BytePlus Seedance 2.0 prompt guide at https://docs.byteplus.com/en/docs/ModelArk/2222480.
- Treat the prompt as an instruction: who, where, doing what, camera, and order of events.
- Attach references in the order you refer to them so that Image 1, Video 1 labels match your prompt.
- Use one camera movement per shot and sequence multi-part videos as Shot 1 / Shot 2 / Shot 3.
- Close every prompt with quality, style, and constraint words to lock the look.
- Keep actions slow, specific, and body-part precise — avoid vague movements.
How Generate Media works
The Generate Media page gives you a prompt box that accepts up to 8000 characters and a reference tray where you attach images and video clips. The order in which you select or attach references maps directly to Image 1, Image 2, Video 1, and so on — use these labels inside your prompt to point Seedance at the right asset.
Three Seedance 2.0 model variants are available: Seedance 2.0 for best quality, Seedance 2.0 Fast for speed, and Seedance 2.0 Mini for economy. All runs produce a 10-second 9:16 vertical video at 720p with audio enabled by default.
Start with the basic formula
When your prompt should draw from a reference, use a clear phrase that names what to take and where to find it: "Reference the dance motion in Video 1 to generate a character performing the same routine in a neon-lit alley" or "Reference the face in Image 1 to generate a close-up of that person speaking to camera."
When you are editing or extending an existing video rather than generating from scratch, address the video directly: "Strictly edit Video 1 and change the background from a kitchen to a beach at sunset." Do not write "reference Video 1" for edits — that wording tells the model to treat the asset as a loose reference rather than the source to modify.
Use the advanced formula
Build prompts like engineering instructions with five layers: first, define the subject; second, describe what they do; third, set the scene and atmosphere; fourth, explain how to shoot it; fifth, tighten the result with quality, style, and constraint words.
Define each subject once and reuse the same label throughout. Instead of "a man" then "a guy" then "the person," stick with one name like "the police officer" or "the thief" so the model maintains consistent identity across the full 10-second output.
Sequence shots like a storyboard
For multi-part videos write Shot 1, Shot 2, Shot 3 in event order. For each shot describe the camera movement, the subject action or expression, the position in frame, and any audio cue. Use only one camera movement per shot — do not stack push, pull, and pan into the same shot description.
Avoid strict second-precise timing such as "0 to 3 seconds." Let the model pace the video naturally from the sequence of shots and actions you describe.
Describe actions precisely
Be specific about body parts: name hands, legs, head, eyes, and shoulders explicitly, then add range, speed, and force. "Slowly raise one hand to chest height" is clearer than "gesture." "Quickly turn the head to look left" is clearer than "react."
Prioritize slow, gentle, continuous movements over sprints, big jumps, or sudden motion — Seedance handles smooth pacing better than abrupt changes. Describe transitions between actions so the model knows how long to hold one pose before moving to the next. Express emotion physically rather than with adjectives: "the character drops their shoulders and looks down" instead of "very sad" or "extremely angry."
Add quality, style, and constraint words
Close your prompt with quality cues and style direction. Quality words such as "HD," "rich details," "cinematic texture," and "soft lighting" raise the visual bar. Style words such as "cyberpunk blue-purple tone" or "retro film grain" define the aesthetic. Constraint words prevent unwanted artifacts: "keep it subtitle-free," "do not generate a logo," and "do not generate a watermark" are the most common fixes.
Pick references that anchor the result
A strong reference set uses 4 to 5 assets total: one or two character images (a close-up face and a full-body shot), one scene or setting image, one motion or camera-reference video, and one audio clip. Do not max out the asset limit — fewer, higher-quality references give the model a cleaner signal.
For character identity, use a dedicated close-up face reference plus a full-body reference. Avoid multi-view character sheets that pack front, side, and back angles into one image — they worsen identity drift. Place your highest-precision references earlier in the prompt so the model weights them more heavily.
Fix common problems
If the output includes unexpected subtitles, add "keep it subtitle-free" to your prompt and prefer landscape input sources when available. If a logo or watermark appears, add "do not generate a logo" or "do not generate a watermark." If the visual style drifts across the 10 seconds, add explicit style words and consider adding a scene reference image. If character identity drifts, add a dedicated close-up face reference and check that your subject labels are consistent throughout the prompt.