Heading #Chapter 1 — Quick Start: What Seedance Can Do
Seedance is ByteDance Seed team's multi-modal video generation model. Seedance 2.0 and 2.5 are fully integrated into viddo.ai. The capability matrix centers on four pillars: long-form narrative, strong reference, precise editing, and multilingual fluency. If you only remember one line about it:
Seedance turns natural language + multi-modal references into 4–30 second videos with full control over shot, sound, subtitles and storyboard, plus post-production precision editing.
1.1 Capability Snapshot

1.2 Typical Use Cases
Seedance is not a toy. In production it shines for:
- Brand TVC commercials (15–30s, single-take)
- Short drama / cinematic shorts (one-shot, multi-character micro-expressions)
- E-commerce SKU showcase videos (multi-angle, multi-shot)
- Education / explainers (step-by-step, biological growth, cultural heritage)
- 3D / gray-model rendering (architecture, product)
- Film pre-visualization (camera pre-viz, age progression, cross-style variants)
Heading ##1.3 Reader Mindset
Think of Seedance 2.5 as a visual content producer — you (the director) supply the script, it shoots it. Everything that follows is essentially teaching you to write the script so the model shoots it right.
Heading #Chapter 2 — Text Prompt Basics
Seedance understands natural language directly. There is no rigid template, but a battle-tested formula helps you organize your thoughts:
Subject + Action or Event + Scene & Environment (optional) + Visual Style (optional) + Camera & Editing (optional) + Sound (optional)

Heading ##2.2 Base Template
performs in .
Visuals: .
Camera: .
Sound: .
## Heading ##2.3 Worked Example: Potter Making a Cup
A potter in a sunlit morning studio finishes a pale-blue ceramic cup, lifting it off the wheel and placing it at the center of a wooden shelf.
Soft morning light streams through the window; the damp clay shows delicate sheen; the workbench stays tidy.
Camera starts with a medium shot of the wheel-throwing motion, then slowly pushes in to the cup surface texture, and finally cuts to a head-on shot of the shelf.
Keep the wheel's low hum, the clay-rubbing sound and a quiet indoor ambience.
## Heading ##2.4 Special Markers for Sound and Text
Prompts can use plain language. To explicitly separate music, SFX, dialogue and subtitles, use these markers:

## Heading ##2.5 Negative Controls (What NOT to Generate)
No background music. Keep only dialogue, ambience and action SFX.
No subtitles.
No sound at all.
## Heading ##2.6 Language Reinforcement for Dialogue
When the dialogue text is not the model's default (Chinese), or you need to force English/Japanese/Korean, strongly recommend specifying the language in front of the line. Recommended formula:
Dialogue language + regional variant or accent + delivery style + speaker + {line}
Dialogue language: American English. The girl says naturally and casually in American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles American English. The young man speaks with natural LA slang: {No way, you actually made it.}
The girl whispers softly in Japanese: {もう大丈夫です}
Heading #Chapter 3 — Multi-Modal References: Pinpointing @Image / @Video / @Audio
Heading ##3.1 Why "Precise Referencing" Is the Core Rule
Seedance accepts any combination of text + images + videos + audio. The more assets you supply, the more critical it becomes that the model knows exactly what each asset is for — a hard rule repeated across S1, S2 and S3:
Mapping between assets and on-screen elements must be written into the prompt. Do not rely on text labels baked into images, and never let the model guess which asset corresponds to which character, prop or scene.
Seedance 2.5 accepts up to 50 assets per task. Going beyond the recommended range is possible but stability drops:

Heading ##3.3 General Sentence Pattern
Reference / Extract / Combine / Follow + Image n + the referenced element, generate + the scene description, keep + the referenced element + consistent
Common verbs: reference / extract / combine / follow / lock / preserve / exclude
Heading ##3.4 Documenting Each Asset's Role
For every asset, do two things: (1) describe what attribute it provides, (2) when the asset risks being "carried in" wrongly, write down what to exclude. S3's recommended template:
@Image1 is used for 's <appearance, clothing, structure or material>.
@Video1 is used for <action, camera motion or rhythm>.
@Audio1 is used for 's .
performs
in .
Visuals: . Camera: .
## Heading ##3.5 Worked Example: Potter
@Image1 is used for the potter's facial features, hair and deep-green apron. Do not use its background.
@Image2 is used for the studio's wooden workbench, window placement and morning light. Do not use any person in this image.
@Video1 is used for the hands-on-wheel motion, lifting the cup and placing it. Do not use the person, clothing or scene from the video.
The potter finishes a pale-blue ceramic cup in the morning studio, lifting it off the wheel and placing it at the center of the wooden shelf.
Camera starts with a medium shot of the wheel-throwing motion, then slowly pushes in to the cup surface texture. Keep the wheel hum, the clay-rubbing sound and quiet indoor ambience.
## Heading ##3.6 Multi-View: Different Angles of the Same Subject
When multiple images show different angles of the same subject (person or product), explicitly label them as "the same X from different views" so the model doesn't treat them as distinct subjects.
@Image1 defines the front view of the same folding lamp.
@Image2 defines the left side structure of the same folding lamp.
@Image3 defines the right side structure of the same folding lamp.
@Image4 defines the rear structure of the same folding lamp.
All four images together define a single folding lamp; the output must contain only one folding lamp.
# Heading #Chapter 4 — On-Screen Text: Slogans, Subtitles and Speech Bubbles
Seedance 2.0 onwards supports on-screen text generation across T2V, I2V, R2V and V2V scenarios. The model auto-matches style and color, and also accepts explicit specifications for color, style, timing, placement and behavior.
Prefer common characters. Avoid rare characters and special symbols for best rendering.
## Heading ##4.1 Slogans (Opening or Closing Tag)
Formula: "text content" + "appearance timing" + "position" + "appearance behavior", "text traits (color, style)"
Hand-drawn comic style. Three people sit together eating fried chicken from Image 1 in a warm friendly atmosphere. The frame gradually blurs and the text "Happiness Lives in Seedance" appears in the center.
## Heading ##4.2 Subtitles (Voiceover or Dialogue)
Subtitles appear at the bottom of the frame. The text content is "...". Subtitles must stay perfectly in sync with the audio rhythm.
**Example 1: Voiceover**
Generate a video with voiceover: a deep, calm male voice says: "In the grand cosmos, our world is only a brief moment. Yet within it, life thrives against all odds."
The scene slowly transitions from night to dawn. Stars fade away, the sun rises from behind the mountains. Subtitles appear at the bottom of the frame in sync with the dialogue.
**Example 2: Dialogue**
The two people from Image 1 chat in an office. The woman speaks first: "You always arrive just on time. Do you enjoy that perfect timing?"
The man smiles and replies: "I have my own rhythm." Their conversation flows naturally; matching subtitles appear at the bottom of the frame.
## Heading ##4.3 Speech Bubbles
says: "...". As they speak, a speech bubble appears around them containing the dialogue.
**Example 1: Running Track Conversation**
The two people from Image 1, wearing sportswear, run on the school track. The girl looks at the boy and smiles confidently: "We can definitely do it!"
Cut to a close-up of the boy; he hesitates: "Are you sure?"
Cut back to the girl's medium close-up; she answers brightly: "Yes!" The mood is bright and decisive. Speech bubbles with the matching dialogue pop up around each speaker.
**Example 2: Strawberry Picking**
Using the girl from Image 1 and Image 2 as reference, the girl picks a strawberry in a strawberry garden, takes a bite and smiles: "This is the real deal!"
A speech bubble with the line appears around her.
# Heading #Chapter 5 — Keyframes, First/Last Frames, Gray Models and Storyboards
## Heading ##5.1 First/Last Frame Generation
In multi-modal reference mode, you can simply write `@Image1` as the first frame and `@Image2` as the last frame at the top of the prompt — no need to switch to a dedicated first/last-frame mode. The system locks the output aspect ratio to the first frame's ratio and lets you set the duration in the UI/API. First and last frames must share the same aspect ratio; otherwise the last frame may stretch.
Document each anchor image separately — do not collapse them into "Image 1 and Image 2 as first/last frames". Other reference images only supply the specified attributes; they do not replace the first/last frame composition.
@Image1 as the first frame, defining the starting composition, subject position, pose, prop state, scene and camera direction.
@Image2 as the last frame, defining the ending composition, subject position, pose, prop state, scene and camera direction.
@Image3 is used for 's . Do not alter the first-frame composition from @Image1, nor the last-frame composition from @Image2.
@Image4 is used for 's . Do not alter the first-frame composition from @Image1, nor the last-frame composition from @Image2.
.
The scene starts naturally from @Image1's first frame, flows through the continuous motion and arrives at @Image2's last frame.
Keep consistent between first and last.
**Example: Perfume Atelier**
@Image1 as the first frame, defining the starting composition, character position, pose, tabletop prop state and camera direction of the perfume atelier.
@Image2 as the last frame, defining the ending composition, character position, pose, tabletop prop state and camera direction of the perfume atelier.
@Image3 is used for the perfumer's face, hair and deep-green apron. Do not alter the first-frame composition from @Image1, nor the last-frame composition from @Image2.
@Image4 is used for the glass perfume bottle's shape, material and label placement. Do not alter the first-frame composition from @Image1, nor the last-frame composition from @Image2.
The perfumer starts from the first-frame pose, picks up a dropper and the glass perfume bottle, drips amber essence into the bottle, gently shakes it and caps it, then places the finished product at the center of the table, finally arriving at @Image2's last frame.
Keep the perfumer's identity and clothing, the bottle's count and structure, the wooden table layout, the warm side lighting and camera direction consistent throughout.
## Heading ##5.2 Multi-Keyframe Sequence Control
When multiple independent images define different stages of a process, open the prompt with "Using @Image1 through @ImageN in order as keyframes", then describe the state each image corresponds to. Independent keyframes usually align better than grid images, but they only control stage order and key states — not frame-by-frame replication.
Use @Image1 through @ImageN in order as keyframes.
@Image1 as the first frame, defining .
@Image2 defines the second keyframe: .
@Image3 defines the third keyframe: .
@ImageN as the last frame, defining .
The scene traverses the states defined by @Image1, @Image2, @Image3 through @ImageN in order, with natural continuous motion bridging each stage.
Keep consistent throughout.
**Example: Orange Paper Plane Flight (4 keyframes)**
Use @Image1 through @Image4 in order as keyframes.
@Image1 as the first frame, defining the orange paper plane resting on the left side of a classroom wooden desk, nose pointing right, in a locked medium shot.
@Image2 defines the second keyframe: the same paper plane is lifted by a hand from the desk, nose direction unchanged.
@Image3 defines the third keyframe: the same plane glides past the window, the curtain sways gently to the right.
@Image4 as the last frame: the same plane lands on the middle shelf on the right, nose still pointing right.
The scene traverses the states defined by @Image1, @Image2, @Image3 and @Image4, with continuous direction and speed.
Keep the plane's orange material, size and creases, the classroom layout, the afternoon side light and the camera axis consistent throughout.
## Heading ##5.3 Grid Storyboards and Multi-Grid References
Grid storyboards provide the overall story, shot order and rough composition. They are not meant for pixel-perfect replication. Keep them within 15 cells, use simple line art or clean diagrams with minimal text labels. Prompts should declare the reading order and then describe each cell's subject action, shot size or camera motion, plus final look and sound.
@Image1 provides an 's shot order and rough composition, read . Do not adopt .
@Image2 defines 's .
@Image3 defines 's .
Shot 1: .
Shot 2: .
...
Shot N: .
Final visuals use . Sound includes .
**Example: Pottery Four-Grid Storyboard**
@Image1 provides a four-cell pottery storyboard's shot order and rough composition, read left-to-right, top-to-bottom. Do not adopt the line-art style or text labels.
@Image2 defines the potter's face, short hair and dark-grey apron.
@Image3 defines the blue-glaze cup's body proportions, glaze color and curved handle.
Shot 1: wide shot of the quiet pottery studio; the potter sits at the wheel.
Shot 2: medium side shot of the potter's hands steadying the spinning wet clay as the cup forms.
Shot 3: close-up of fingers trimming the rim and handle joint; slip slowly slides down the fingertips.
Shot 4: medium close-up of the fired blue-glaze cup placed on the wooden shelf; the potter withdraws their hands.
Final visuals use documentary-realism texture. Keep the wheel hum, the wet clay sound and studio ambience.
# Heading #5.4 Gray-Model Reference and Rendering
Gray-model references split into coarse and fine-grained. Decide whether the gray model supplies the motion skeleton or a fully built structure, then pick the matching template.

### Heading ###5.4.1 Coarse Gray Model (Mobile Cart Example)
@Video1 is a coarse gray-model reference, only providing character walking path, cart movement direction, locked camera, one push-in and two cuts. Do not adopt the gray geometric look or empty scene.
The tall cylinder in @Video1 corresponds to the docent.
The cuboid in @Video1 corresponds to the mobile cart.
@Image1 defines the docent's face, blue uniform and badge.
@Image2 defines the mobile cart's white metal frame and transparent cover.
@Image3 defines the tech showroom's curved wall, grey floor and ceiling strip lights.
The docent pushes the mobile cart along the curved wall, stops at the central platform and opens the transparent cover.
Keep @Video1's walking path, subject blocking, push-in direction and cut positions.
Visuals use a bright realistic documentary style. Keep footsteps, cart wheels and the showroom ambience.
### Heading ###5.4.2 Fine-Grained Gray Model (Ring Installation Re-Render)
@Video1 is a fine-grained gray-model reference. Keep the ring installation's full structure, three-ring rotation, pedestal position, orbit camera and cuts. Do not adopt the original gray material or empty background.
@Image1 defines the outer ring's brushed brass material.
@Image2 defines the inner leaf layer's translucent blue glass material.
@Image3 defines the contemporary art gallery's white curved walls, deep grey floor and overhead soft light.
Re-render the ring installation in @Video1 as a dynamic sculpture of brass and blue glass; re-render the scene as a contemporary art gallery.
Keep @Video1's structure, rotation rhythm, orbit camera and cuts. Keep the installation's low rotation sound and quiet indoor ambience.
# Heading #Chapter 6 — Video Editing: Add, Remove, Replace, Subject & Background Swaps
Seedance 2.0 already supports video editing; Seedance 2.5 raises precision further. Editing tasks automatically lock the output aspect ratio to the input video's and lock the duration (which you cannot override). Output duration may drift by up to ~0.3s due to input frame handling, but overall content and event order remain essentially preserved.
# Heading #6.1 Universal Three-Part Template
【Edit Goal】
Edit @Video1 to across .
【Original Video Responsibility】
@Video1 is the sole editing master. It owns .
【Target Asset Responsibility】
@Image1 or @Audio1 supplies .
【Edit Scope】
Touch only