ManyMotions
Media ForgeGenerate images & video, fastMedia StudioNewBuild images & video by chatting with AIWorkflowsAutomate pipelines at scaleClipperNewBreak a video down & cut the highlightsMarketing AgentDone-for-you UGC video ads
MCPPricingFAQ
Log inStart for $1
AI Models/Seedance 2.0/Prompting

How to prompt Seedance 2.0

Seedance 2.0 generates video and audio in a single pass, so a good prompt directs both the picture and the soundtrack. It responds best to prompts written like a shot list: concrete subject and action, an explicit camera instruction, lighting, then what should be heard — dialogue in quotes, ambient sound, foley. There is no negative prompt and no camera-lock switch: everything, including a static camera, is asked for in plain language inside the prompt.

Anatomy of a Seedance 2.0 prompt

  1. 1

    Subject & action

    Open with who or what and the specific action — "a barista steams milk", not "a coffee scene". Vague openers produce generic motion.

  2. 2

    Setting & time

    Where and when: "in a cramped Tokyo kissaten at night". Environment words drive texture, reflections and background detail.

  3. 3

    Camera

    One explicit camera instruction per shot: "slow dolly-in from medium to close-up", "locked-off static camera", "handheld tracking shot". Seedance 2.0 understands director-level grammar — dolly zoom, rack focus, POV switch.

  4. 4

    Lighting & look

    Light source, direction and mood: "single warm tungsten bulb overhead, deep shadows". Add a film-look reference if you want one ("shot on 35mm, shallow depth of field").

  5. 5

    Audio

    Say what should be heard. Dialogue goes in quotes ("she says: …") and gets lip-synced; ambient and foley are described ("rain on glass, low café chatter"). If you skip this block the model invents the soundtrack.

  6. 6

    Multi-shot plan

    Optional: number the shots with timing ("Shot 1 (0-4s): … Cut to Shot 2 (4-10s): …") and Seedance cuts them together in one output.

Template

[subject] [specific action], [setting and time of day], [camera instruction], [lighting and film look], [audio: dialogue in quotes / ambient / foley]

10 example prompts that work

UGC ad with spoken dialogue

9:16 · 10s · 1080p · audio on
A woman in her 30s in a bright bathroom holds a green glass serum bottle up to the camera, selfie framing, slight handheld sway, morning window light. She says with a casual smile: "Okay, one week with this and my skin is actually glowing — look at this." She tilts the bottle so the label catches the light. Ambient: quiet room tone, faint birdsong outside.

The quoted line is lip-synced natively — no separate TTS pass. Selfie framing + handheld sway reads as authentic UGC; naming the label moment forces a product beat instead of generic waving.

Product hero shot (static camera)

1:1 · 6s · 4K · audio on
A matte black wireless earbud case rotates slowly on a marble pedestal, studio product photography, locked-off static camera, single softbox from the left with a rim light from behind, dark seamless background, macro detail on the hinge as it opens. Audio: soft mechanical click when the lid opens, low ambient hum.

"Locked-off static camera" replaces the missing camera-lock switch — say it explicitly or the model will drift. One named micro-event (the lid opening) gives the clip a beat and a foley cue to sync.

Cinematic b-roll

21:9 · 8s · 4K · audio on
Neon-lit pizzeria on a rainy street corner at night, slow lateral tracking shot past the fogged window, a cook tosses dough inside, reflections smearing across wet asphalt, shot on 35mm, anamorphic lens flares, moody teal-and-red palette. Audio: rain, muffled Italian radio from inside, a passing car.

Tracking + a single interior action gives layered parallax. The palette and lens words are followed closely — this model rewards specific color direction over adjectives like "beautiful".

Multi-shot sequence in one generation

16:9 · 12s · 1080p · audio on
Shot 1 (0-4s): wide establishing, a hiker reaches a foggy ridge at dawn, drone orbit. Cut to Shot 2 (4-8s): close-up of her boots crunching frost, low tracking shot. Cut to Shot 3 (8-12s): she turns to camera and says "worth every step", handheld, golden light breaking through. Continuous ambient: wind, distant birds; boots foley in shot 2.

Numbered shots with timings produce real cuts inside a single clip — no external editing. Keep one subject across shots and restate her ("the hiker", "she") so identity holds.

Physics showcase

16:9 · 6s · 1080p · audio on
A ceramic mug slides off a wooden desk in slow motion and shatters on a concrete floor, fragments scattering with realistic weight, locked-off low-angle camera at floor level, hard window light casting long shadows. Audio: the slide, a beat of silence, then the sharp shatter with room echo.

Seedance 2.0 is strong on collisions and material behavior — name the materials (ceramic, concrete) and the physics you expect. The "beat of silence" line times the audio to the fall.

POV switch

16:9 · 8s · 1080p · audio on
First-person POV: gloved hands push open a heavy vault door, flashlight beam sweeping over dusty shelves. Mid-clip, switch to a reverse shot: the explorer framed in the doorway, backlit by the corridor. Audio: metal groan of the door, footsteps with stone echo, tense low drone.

"Switch to a reverse shot" is understood as an in-clip cut — one of the camera moves this model handles that most others refuse. Keep the location constant so the reverse shot matches.

Dialogue scene, two speakers

16:9 · 10s · 1080p · audio on
Two friends at a diner booth, over-the-shoulder framing that alternates with each line. The first, a man in a denim jacket, says: "You actually quit?" The second, a woman stirring coffee, smiles and answers: "Signed the papers this morning." Warm practical lamps, soft film grain. Audio: cutlery, low diner chatter under the dialogue.

Attribute each quoted line to a described speaker and the lip-sync lands on the right face. Alternating over-the-shoulder framing is a camera instruction, not an edit note — the model stages it.

Vertical hook for TikTok

9:16 · 6s · 1080p · audio on
Extreme close-up of hands cracking an egg one-handed over a sizzling pan, whip-pan up to a chef grinning at camera, kitchen towel over shoulder, punchy overhead light. He says: "Stop cracking eggs like a tourist." Audio: loud sizzle, the crack, a short record-scratch on the whip-pan.

Front-load the visual hook in the first second — the close-up action — then the whip-pan delivers the face and the line. Sound effects described at exact moments become the edit rhythm.

Image-to-video: animate a still

i2v · 16:9 · 6s · 1080p
Bring this photo to life: the model keeps her pose but turns her head slowly toward camera, hair moving in a light breeze, background bokeh shimmering, subtle parallax as the camera pushes in a few centimeters. Audio: soft wind, distant city ambience.

In image-to-video the prompt describes only the MOTION — composition, wardrobe and light come from your uploaded frame. Small camera moves ("a few centimeters") keep the source photo's fidelity.

Reference-to-video: consistent character

ref2video · 16:9 · 12s · 1080p
Using the reference images of the astronaut character: Shot 1 — she runs across a red desert plain, wide tracking shot, dust kicking up. Shot 2 — close-up inside her helmet, eyes scanning, HUD reflections. Shot 3 — she plants a flag, low heroic angle, sun flare. Same suit, same face throughout. Audio: breathing inside helmet, wind, a swelling synth note on the flag plant.

Reference-to-video accepts up to 9 reference images and holds the character across every shot — the fix for identity drift in multi-scene stories. Restate "same suit, same face" as cheap insurance.

Settings that matter

  • Duration

    4-15s. Use 4-6s for single actions, 10-12s for multi-shot sequences; past 12s reserve for stories that genuinely need it — pacing gets looser.

  • Resolution

    Draft at 480p to iterate on the prompt cheaply, then rerun the winning prompt at 1080p or 4K. The prompt transfers 1:1 across resolutions.

  • Aspect ratio

    9:16 for TikTok/Reels, 16:9 for YouTube, 21:9 for cinematic b-roll, 1:1 for product tiles. Pick it before writing the prompt — framing language should match.

  • Audio

    Leave on and direct it in the prompt. Turn it off only for silent loops or when you plan to score the clip yourself.

  • Mode

    Text-to-video for from-scratch shots, image-to-video when you have the exact frame, reference-to-video when a character must survive multiple shots.

Do and don't

Do

  • Write one camera instruction per shot, in film language — dolly, tracking, rack focus, whip-pan.
  • Put dialogue in quotes and describe the speaker; lip-sync follows the attribution.
  • Describe the soundtrack (ambient, foley, music) — silence about audio means the model improvises it.
  • Number shots with timings when you want in-clip cuts.
  • Name materials and weights for physics moments — "ceramic mug", "heavy vault door".
  • Iterate at 480p, finalize at 1080p/4K with the identical prompt.

Don't

  • Don't use negative-prompt syntax ("no text, no blur") — it isn't supported; describe what you DO want instead.
  • Don't expect a camera-lock parameter — ask for "locked-off static camera" in the prompt.
  • Don't stack five adjectives where one concrete noun works; "cramped Tokyo kissaten" beats "beautiful atmospheric detailed café".
  • Don't change subject mid-prompt without a shot cut — identity smears.
  • Don't write dialogue longer than the clip: roughly 2.5 words per second is what fits naturally.

Advanced techniques

Director-level camera grammar

Seedance 2.0 executes named camera techniques that most video models approximate or ignore. Use the exact term and, when it helps, describe the intent: "dolly zoom on his face as the realization hits" reads better than a bare "dolly zoom".

  • dolly-in / dolly-out — physical push toward or away from the subject
  • dolly zoom (Vertigo shot) — background stretches while the subject holds
  • rack focus — focus shifts between two named planes ("from the glass to her eyes")
  • tracking / lateral shot — camera travels with the subject
  • whip-pan — fast blur transition, great for hooks
  • POV and reverse-shot switches — in-clip perspective cuts
  • handheld — adds sway and realism to UGC-style footage

Writing audio like a sound designer

Treat the audio block as its own script. Layer it: room tone first, then point sounds tied to visible events, then dialogue. Timing words ("a beat of silence, then…") shift when sounds land. For music, describe genre, energy and where it should swell — "low synth pulse that swells on the last shot" — instead of naming artists.

Multi-shot stories that hold together

The reliable pattern for sequences: one protagonist described once, then referred back to ("the hiker", "she"); one location or a stated location change per cut; continuous ambient sound declared once so the cut feels intentional rather than like two stitched clips. For character-driven stories across many generations, switch to reference-to-video and reuse the same reference set every time.

Which variant to use

  • Seedance 2.0

    Default. Full quality text-to-video and image-to-video, 480p-4K, native audio.

  • Seedance 2.0 Fast

    Drafting and prompt iteration — quicker and cheaper text-to-video, same prompt grammar, then rerun on the full model.

  • Seedance 2.0 Ref2Video

    Up to 9 reference images, one coherent multi-shot scene — character consistency across cuts that plain t2v can't hold.

Frequently asked

Does Seedance 2.0 support negative prompts?

No. Describe the shot you want instead of listing what to avoid — "clean background, no on-screen text" becomes "seamless dark studio background". Positive, concrete description is the only steering mechanism.

How do I keep the camera completely still?

Ask for it in the prompt: "locked-off static camera" or "tripod shot, no camera movement". There is no parameter for it, but the instruction is followed reliably.

How long can spoken dialogue be?

Budget about 2.5 words per second of clip. A 10-second clip carries roughly 25 words of natural-paced dialogue; more than that gets rushed or truncated.

Text-to-video or image-to-video for product ads?

If you have a real product photo, image-to-video: the label, colors and proportions stay exact and the prompt only directs motion and audio. Text-to-video is for scenes where no exact asset must survive.

Why does my character look different in every clip?

Separate generations don't share identity. Use Seedance 2.0 Ref2Video with the same reference images for every scene — up to 9 refs, and the character holds across shots and generations.

Try Seedance 2.0 in Media Forge →About Seedance 2.0All models
ManyMotions

Tell me what you need and grab a coffee. I'll have it ready before you finish the cup.

Product

Marketing AgentWorkflow StudioMedia ForgeTemplatesAI Models

Open in app

Workflow StudioMedia ForgeProjectsAvatars

Company

ContactPrivacy PolicyTerms of ServiceContent Moderation PolicyLegal NoticeCancel subscriptionWithdraw from contract

© 2026 ManyMotions

support@manymotions.com