Seedance 2.5 will take a one-line description and give you something back. It just won't be the shot you had in mind. Learning how to prompt Seedance 2.5 well is mostly learning what the model needs to be told and what it will invent if you stay silent.
This guide covers what changed in 2.5, the blocks a prompt needs, the four ways people format them, and then ten complete prompts — each with the video it produced, ready to copy and adapt.
What changed in Seedance 2.5
Three changes matter for how you write, and one widely repeated claim is wrong.
Clips run to 30 seconds. Videos generate up to 30 seconds in a single pass, with the option to extend twice. That is long enough to hold a beginning, a middle and an end — which means your prompt now has to describe a sequence, not a moment.
Sound is generated with the picture. Seedance 2.5 uses a unified audio-video joint-generation architecture: dialogue, ambience and effects are produced in the same pass, not dubbed on afterwards. This is the change that catches people out. If you don't specify the audio, the model invents it — and what it invents is often a generic music bed you didn't want.
References got much wider. A single generation accepts up to 30 images, 10 video clips and 10 audio clips. That turns reference handling from "attach a face" into a casting and art-direction job.
The 4K claim is not real. ByteDance's own pages state no 4K output spec, and providers are more conservative still — fal lists text-to-video output at 480p and 720p. Write for HD delivery and treat 4K as an upscaling step afterwards.
The blocks a prompt needs
A working prompt answers six questions. Miss one and the model fills the gap with its own default:
| Block | What it decides | If you leave it out |
|---|---|---|
| Subject | Who or what is on screen, and what they do | Generic casting, vague action |
| Camera | Body, lens, framing, movement | A drifting, over-smooth default move |
| Location | Where, time of day, light source | Flat, evenly-lit nowhere |
| Style | Era, film stock, grade, texture | Clean digital video look |
| Audio | Dialogue, ambience, music or its absence | Invented music bed, sometimes narration |
| Constraints | What must not appear | Subtitles, watermarks, modern objects |
The last row is the one people skip. A quarter of the prompts further down this page spend real word count on what must not appear — that isn't padding, it's where the model drifts.
As a starting skeleton:
STYLE: [look, era, film stock, grade]
CAMERA: [body, lens, movement, framing]
SUBJECT: [who, what they wear, what they do]
LOCATION: [where, time of day, light source]
AUDIO: [dialogue, ambience, music or the absence of it]
CONSTRAINTS: [what must not appear]
Four ways to format a prompt
Those blocks can be written four ways, and they form a progression from plain writing toward structured data:
| Level | Format | What it looks like | Use when |
|---|---|---|---|
| 1 | Prose | Plain paragraphs, no headers | One continuous idea or mood |
| 2 | Labeled | CAMERA: / Style: colon headers | Two instructions start competing |
| 3 | Bracketed | [STYLE + CAMERA], [CHARACTERS], [TIMELINE] blocks | Sections grow past a few lines |
| 4 | JSON | A shots array with time, camera, dialogue keys | Directing distinct shots with their own dialogue |
Here is the thing worth knowing: most working prompts are plain prose. The elaborate bracketed skeletons that circulate in screenshots are the minority practice, not the norm. Formatting is not a sophistication ladder — it is a response to how many things you are trying to hold in place at once. Start at level 1 and escalate only when something breaks.
When to add a timeline
Most guides tell you a 30-second shot needs a timeline — ranges like 0-6s, 6-12s, each with its own action. fal.ai's guide opens with exactly that.
In practice, most creators don't use one. Of the prompts collected on this page, well under half have a timeline at all, and the ones that do have something in common: the order of events carries the meaning. A freeze-and-rewind effect. A single unbroken take crossing a party. A walk that degrades across five stages of public humiliation. Shuffle the beats and the piece is destroyed.
The rest are atmosphere pieces — a mood, a texture, one continuous emotional state. Nothing needs scheduling, so nothing is scheduled.
That gives you a test: does your shot have a plot? If reordering the middle would ruin it, write a timeline. If not, a timeline is overhead — and worse, it invites the model to manufacture events to fill ranges you hadn't thought through.
If you do use one, keep the ranges contiguous with no gaps and one clear action each. Treat a range as a time budget, not a cut point; actions land near boundaries, not on them.
The camera vocabulary that works
One word does more work than any other in Seedance 2.5 prompts: handheld. It appears more often than every other camera term combined — more than close-up, medium shot, over-the-shoulder, steadicam and tracking shot put together.
That is not an accident. Experienced users are not asking for smooth, stabilized, "cinematic" motion. They are asking for the camera to feel operated by a person. Left to itself the model produces a gliding, weightless move that reads as synthetic instantly. handheld is the cheapest single fix for that.
The second habit is naming real equipment. Not "cinematic look" but ARRI Alexa XT with Cooke anamorphic lenses. Terms like film grain, motion blur, shallow depth, 16mm and MiniDV recur constantly, because they carry a specific texture the model has actually learned.
What is missing from these prompts is just as telling: dolly, crane and rack focus appear zero times across the whole set, as does every term in the lighting department's vocabulary. Writing cinematic Seedance 2.5 prompts takes that measurement apart — what replaces the glossary, and how to write a camera move and a lighting setup without it.
Locking a character's identity
When a reference image defines your character, pin it in the first instruction, before scene or action, and enumerate the attributes rather than asking for consistency in the abstract:
Use the uploaded reference image as the exact character reference.
Preserve her facial identity, eye color, skin tone, hairstyle, makeup,
body proportions, and overall appearance throughout the video.
"Keep her consistent" does not work. The list does.
With up to 30 image references available, an image-to-video workflow can pin far more than a face. The stronger pattern is to give every reference a name and one job. One prompt further down carries fourteen named references — [Pocket Watch Reference], [Color Grading Reference], [Character Reference_01] — each stated to control exactly one thing. Naming beats numbering, because the name travels with the instruction.
Writing the sound
Since audio is generated jointly, silence in your prompt is an instruction to improvise. Say what you want heard:
- Attach each line to a speaker and a moment, not to the end of the prompt.
- Keep spoken lines short. Paragraph-length scripts degrade both delivery and mouth motion.
- State ambience separately from music, and say explicitly if you want no music.
- Add
no subtitlesunless you actively want text on screen.
That last one is worth dwelling on. ByteDance's own announcement says the model "minimizes uncontrolled occurrences in subtitles and background music" — the vendor acknowledges the failure mode. Experienced users still write the guard explicitly, because minimized is not eliminated. Burned-in subtitles and an unwanted score are the two audio artifacts worth naming every time.
How long should a prompt be?
The prompts on this page run from 837 to 5,815 characters. The shortest is a finished commercial shot. The longest is a single continuous emotional close-up. Neither is padded — they are solving different problems.
Length tracks how many independent things must be held in place at once, not ambition or quality. A one-beat idea with one subject and one light source lands under 1,000 characters. A multi-reference scene with staged action and dialogue runs past 3,000. Write until everything that must not drift has been pinned, then stop.
10 prompts and the videos they produced
Ten complete prompts, ordered shortest to longest, spanning all four formats above. Each is reproduced in full — collapse them if you are skimming — with credit and a link to the original post.
1. Heatwave Ice Water Cooldown Ad
The shortest one here. · 837 characters · format: Prose · timeline: no
No timeline, no labels, no reference image. One paragraph that names format, subject, action, and light, then stops. It is proof that a short Seedance 2.5 prompt can carry a finished ad shot when the idea is a single beat.
Full prompt (837 characters)
High quality video 16:9 aspect ratio 14 seconds duration Cinematic medium shot of a young woman with a short brown bob haircut wearing an open light blue button down shirt over a white crop top and denim skirt She stands on an outdoor patio with hazy warm lighting and out of focus string lights in the background A text graphic reading 111 F is overlaid at the top center She is visibly hot and fans her face with her right hand A hand enters the frame from the right holding a tall glass of ice water with a lemon and a straw She takes the glass with both hands and begins sipping through the straw As she drinks the text graphic rapidly counts down to 69 F She closes her eyes and looks completely refreshed Ambient hazy lighting Upbeat refreshing summer pop music playing Photorealistic 8k resolution highly detailed natural lighting
Prompt by @bmx_ai13.
2. Korean Laundry Home Video
Short prose, but it declares its own shot count. · 1,096 characters · format: Prose · timeline: no
15s handheld home-video vlog, 7-shot montage — the prompt sets a budget in its first line instead of a timeline. Cheaper to write than time ranges and it survived seven cuts.
Full prompt (1,096 characters)
15s handheld home-video vlog, 7-shot montage. Photorealistic phone footage with slight tilt, natural shake, window light, subtle film grain.
A woman (use Image1 only for facial identity and hairstyle) does laundry alone on a quiet morning. Outfit: oversized cream linen shirt with rolled sleeves, grey knit shorts, loose cotton apron. Cozy sunlit laundry nook with an open front-load washer, overflowing basket, wooden drying rack, clothespins, and warm sunlight. She is the only person in the video.
Sequence: untangles wet clothes → shakes out a shirt → checks a collar stain by the window → hangs it saying "Good enough." → finds a mismatched sock → struggles with a heavy bedsheet while laughing → finishes hanging it and quietly admires the sunlit laundry.
Dialogue is natural spoken Korean (except "Good enough"), reacting casually to each moment. Ambient sound only: washer winding down, wet fabric, clothespins, rustling clothes, soft laughter, breeze. No subtitles, text, logos, or watermarks. Do not recreate or copy the reference image—use it only for facial identity and hairstyle.
Prompt by @doctorwasif.
3. Antique Market Multi-Reference Test
Fourteen named references in one prompt. · 1,955 characters · format: Bracket · timeline: no
Every reference gets a bracket name and exactly one job: [Pocket Watch Reference], [Color Grading Reference], [Character Reference_01]. It is the clearest demonstration on this page of naming references instead of numbering them.
Full prompt (1,955 characters)
Use [Worn Leather Reference]
as a reference for WORN LEATHER Use
[Tarnished Frames Reference]
as a reference for TARNISHED FRAMES Use
[Pocket Watch Reference]
as a reference for POCKET WATCH Use
[Location Reference]
as the location reference (only use this as a location reference, not as a color or style reference) Use
[Layout Refernce]
as a reference for LAYOUT REFERENCE Use
[Drinking Glass Reference]
as a reference for DRINKING GLASS Use
[Cracked Porcelain Reference]
as a reference for CRACKED PORCELAIN Use
[Color Grading Reference]
as a color grading reference (only use this as a color grading reference, not a location reference) Use
[Character Reference_03]
as a character reference for ISAAC Use
[Character Reference_02]
as a character reference for SYDNEY Use
[Character Reference_01]
as a character reference for MARKUS Use
[Cased Taxidermy Reference]
as a reference for CASED TAXIDERMY Use
[Brass Sextant Reference]
as a reference for BRASS SEXTANT Use
[Brass Instruments Reference]
as a reference for BRASS INSTRUMENTS Use
[Three Objects Reference]
as a reference for THREE OBJECTS Thirty-second continuous take, an outdoor antique market at overcast midday. ISAAC moves slowly along a crowded stall, camera tracking with him at chest height. The stall is dense with the referenced objects — BRASS INSTRUMENTS, CRACKED PORCELAIN, TARNISHED FRAMES, WORN LEATHER, CASED TAXIDERMY — arranged exactly as in the LAYOUT REFERENCE. ISAAC lifts a BRASS SEXTANT, turns it once, sets it down. He lifts a POCKET WATCH, opens the case, closes it, sets it down. MARKUS (the stall/shop owner) watches him without speaking. ISAAC passes SYDNEY (who is browsing the far end). Then ISAAC stops. Camera pushes in as he reaches past THREE OBJECTS and lifts a green glass ornate DRINKING GLASS. ISAAC turns it to the light. His expression changes completely. Camera holds on his face. Photoreal, overcast diffuse light, muted palette, 35mm, natural handheld.
Prompt by @CuriousRefuge.
4. The Umbrella Escape
The only JSON-formatted prompt here. · 2,660 characters · format: Json · timeline: no
A shots array where each object carries time, type, action, camera, and dialogue. It is the timeline idea taken to its logical end — and the only prompt here that assigns dialogue per shot rather than in a lump.
Full prompt (2,660 characters)
{ "title": "The Umbrella Escape", "style": "3D Pixar family animation. Rainy city afternoon, vibrant reflections on wet pavement, action-comedy chase, warm and cool contrast, dynamic camera.", "shots": [ {"time":"00:00-00:03","type":"CLOSE-UP","action":"Kid in yellow raincoat opens a bright red umbrella. A gust of wind hits. The umbrella inverts and yanks free.","camera":"Macro, raindrops flying.","dialogue":"Umbrella: 'FREEDOM!' Kid: 'Hey!'"}, {"time":"00:03-00:06","type":"TRACKING","action":"Umbrella surfs down the sidewalk, opening and closing to dodge pedestrians. It hops over a puddle like a skipping stone.","camera":"Low tracking, fast.","dialogue":"Umbrella: 'Can't catch me! I'm born to fly!'"}, {"time":"00:06-00:09","type":"WIDE","action":"Umbrella weaves through a flock of pigeons. They scatter. One pigeon gets caught inside and spins out dizzy.","camera":"Chaotic wide.","dialogue":"Pigeon: 'What the-'"}, {"time":"00:09-00:12","type":"ACTION","action":"Umbrella surfs a puddle wave, then catches a crosswind and smacks into a lamppost, spinning around it like a tetherball.","camera":"Spinning camera with umbrella.","dialogue":"Umbrella: 'Wheee! Okay, that hurt.'"}, {"time":"00:12-00:15","type":"TRACKING","action":"A street sweeper approaches. The umbrella gets sucked into the brush, spins wildly, and launches out like a frisbee.","camera":"Following shot, fast.","dialogue":"Umbrella: 'I REGRET NOTHING!'"}, {"time":"00:15-00:18","type":"WIDE","action":"Umbrella glides toward a tree, sticks perfectly in the branches, and sighs. Kid arrives below, out of breath.","camera":"Wide, rain falling.","dialogue":"Kid: 'Got... you...'"}, {"time":"00:18-00:22","type":"TWO-SHOT","action":"Kid climbs up. Reaches for umbrella. Umbrella closes tight, refusing. Kid pouts. Umbrella opens one eye.","camera":"Close two-shot in the tree.","dialogue":"Umbrella: 'I'm not coming back. I tasted the wild.' Kid: 'I'll let you pick the movie.'"}, {"time":"00:22-00:25","type":"CLOSE-UP","action":"Umbrella pauses. Opens fully. Gently covers the kid from the rain, settling onto their shoulder.","camera":"Warm push-in.","dialogue":"Umbrella: 'Fine. But I'm driving next time.'"}, {"time":"00:25-00:28","type":"WIDE","action":"Kid walks home under the umbrella. The umbrella steers them left, then right, playfully. They splash through puddles together.","camera":"Wide, beautiful rainy street.","dialogue":"Kid: 'You're impossible.' Umbrella: 'Thank you.'"}, {"time":"00:28-00:30","type":"TITLE CARD","action":"Black screen. Title 'THE UMBRELLA ESCAPE' in raindrop letters with a red umbrella icon.","camera":"Static.","dialogue":"(soft rain)"} ] }
Prompt by @Dheepanratnam.
5. Brazil House Party Sequence Shot
A single unbroken take with a spoken script. · 2,466 characters · format: Prose · timeline: yes
Written by an audio person, and it shows: dialogue is attached to characters and moments rather than dumped in a block at the end.
Full prompt (2,466 characters)
SEQUENCE SHOT. NO CUT.
Single unbroken handheld take throughout, 30 seconds total.
Young Caucasian woman, chin-length auburn bob haircut, green eyes, wearing a yellow and green Brazil national football jersey and white denim shorts, sitting pensively on a sofa inside a large bright suburban Parisian house, summer daytime, sunlight through the windows, young adults dancing and laughing around her holding drinks, loud music implied by energetic crowd movement.
0-4s: She rises from the sofa empty-handed, walks toward a kitchen table lined with alcohol bottles and glasses. Handheld camera follows close behind her shoulder, slightly shaky.
4-8s: She pours herself a glass of champagne, sips it. Camera holds behind her, shaky, close distance.
8-13s: She weaves through the dancing, laughing crowd, some standing, some in groups laughing on sofas, toward a large glass bay window. Camera continues following close behind her, no cut.
13-16s: She steps outside into the garden, a big pool comes into view with young adults swimming and laughing inside it. Camera smoothly rotates from following behind her to facing her. She looks toward the pool with amusement, smiles, sets her champagne glass down on a small table beside a lounge chair.
16-20s: She removes her jersey, places it on the back of the lounge chair, then removes her white denim shorts, places them on the lounge chair, revealing a swimsuit underneath. Camera holds a medium shot facing her, slight handheld sway. 20-24s: She walks toward the pool, camera repositions behind her as she walks, then she breaks into a run. Handheld camera follows close, shaky, matching her pace.
24-27s: She dives into the pool shouting "take care", landing in the water among other laughing young adults. Camera dives in with her, briefly submerged underwater alongside the swimmers.
27-30s: She resurfaces laughing, joins the laughing adults nearby. Camera is up in the air, revealing a wide aerial view of the house, garden and pool.
Bright natural summer daylight, warm sunlit color palette, candid documentary energy, slight handheld lens wobble throughout.
Maintain her exact hairstyle, green eyes, Brazil jersey and white denim shorts until she undresses, then swimsuit consistency for the remainder. Background party guests and swimmers generic, unnamed. No text overlays, no flickering, no ghosting, no morphing artifacts, natural skin tones, stable continuous handheld motion, no hard cuts at any point.
Prompt by @matt_elevenlabs.
6. 1950s Diner Freeze and Rewind
Timeline used for a physical effect, not a story. · 2,525 characters · format: Prose · timeline: yes
Freeze-and-rewind needs the model to know the order of events. This is the case where time ranges are not stylistic — remove them and the effect cannot work.
Full prompt (2,525 characters)
Photorealistic cinematic 1950s American diner, chrome stools, red vinyl, neon glow and checkerboard floor, shot with modern lived-in realism and soft natural window light. Subtle handheld texture, warm practicals, rich period detail, heavy film grain.
0-4s: [Medium Wide] A striking young woman in her early 20s sits alone at the counter, calm and slightly amused, slowly sipping a tall thick milkshake through a straw. Behind her a young waitress in classic uniform approaches with a tray of eggs and bacon in one hand and a full glass coffee pot in the other. An older lady starts rising from a nearby booth.
4-8s: [Dynamic Tracking] The older lady collides hard into the waitress. Tray, plate, eggs, bacon and coffee pot explode upward in chaotic slow motion. Coffee erupts into long liquid ribbons and perfect suspended droplets. Camera immediately begins a smooth continuous orbit around the impact. Time locks completely at the peak of the spill. Every face freezes in pure shock. Only the girl at the counter keeps moving, completely unfazed.
8-17s: [Slow 360° Orbital] Camera glides in a full elegant orbit through the frozen diner. Coffee hangs in mid-air as glassy ribbons and spheres with perfect volume and surface tension. Bacon strips, eggs and the spinning tray float weightlessly. Patrons and waitress remain locked in startled expressions. The girl at the counter takes one slow, deliberate sip, eyes half-lidded, almost bored, while the entire frozen world (except her) begins an elegant reverse: every droplet, every piece of food and every person rewinds smoothly back to the exact starting positions.
17-24s: [Medium Shot] Rewind lands perfectly. Waitress stands balanced again with tray and coffee pot. The girl lifts her eyes, raises two fingers in a small casual gesture and softly calls the waitress by name. The waitress turns toward her just before the older lady begins to stand, completely avoiding the collision. A tiny private smile crosses the girl’s face.
24-30s: [Extreme Close-Up] Hard cut to her face as she takes one last slow sip. Soft knowing smile, eyes almost closed in quiet satisfaction, like she has done this a hundred times. Shallow depth of field, creamy bokeh of the neon diner behind her.
Photorealistic, ultra-detailed fluid physics, perfect motion blur only on moving elements, stable characters, cinematic lighting, heavy natural film grain, no artifacts, movie-level temporal coherence, high rewatch value.
30 seconds long is the new AI standard! hope you like it 🫡
Prompt by @techhalla.
7. K-pop Backstage MiniDV Vlog
Labeled sections carrying a texture spec. · 3,240 characters · format: Labeled · timeline: no
MiniDV camcorder look, first-person selfie POV. The colon-labeled format keeps the texture instructions from bleeding into the action description.
Full prompt (3,240 characters)
Prompt:
Handheld MiniDV camcorder aesthetic with a genuine first-person selfie POV. CHASE is always holding the camera herself—there is never a camera operator. Natural handheld movement with slight nervous shake, imperfect framing, autofocus hunting, exposure breathing, realistic motion blur, subtle tape grain, soft highlight bloom, authentic DV colors, and onboard camcorder microphone audio.
Create a 30-second ultra-realistic backstage vlog of CHASE, a stunning Korean K-pop idol moments before going on stage. The entire video should feel completely unscripted, as if she casually started recording herself while walking backstage.
The video opens with the camera accidentally turning on while she's already walking, only half of her face visible before she notices and laughs softly. She adjusts the camera and naturally says, "Oh... it's recording already? Okay... hi. I have about one minute before I go on stage."
She continues walking through a busy backstage hallway while staff members pass naturally in the background. Keeping her voice low so she doesn't disturb anyone, she smiles and says, "You know what's funny? I've performed this song so many times... but this moment always feels new."
She stops briefly near a mirror, adjusts one sparkling earring, notices her slightly trembling hand, smiles at herself and quietly laughs, saying, "My hands are always cold before a show... see? Still happening."
As she slowly approaches the stage entrance, the hallway becomes quieter. She suddenly pauses, hearing something beyond the curtain. She looks toward the stage, listens for a moment, then whispers, "Wait... can you hear that?" At the exact moment she says those words, realistic muffled audience cheers begin building naturally from behind the curtain. The cheering grows louder while genuine emotion appears on her face through subtle eye movement and a soft smile—no exaggerated acting.
She turns the camera toward the bright stage lights spilling through the curtain before bringing it back to herself and quietly says, "That's why I never get tired of this... Thank you for waiting for me."
A stage manager calls from off camera, "CHASE, standby!" She immediately takes a slow breath, smiles warmly into the camera, makes a tiny finger heart, gives a small reassuring nod as if sharing the moment with the viewer, softly whispers, "Okay... let's go," then lowers the camera slightly and jogs naturally toward the stage. The camera shakes realistically as she runs, her silhouette disappears into the bright stage lights, the audience erupts into loud cheers, and the recording cuts naturally at the exact moment she steps onto the stage.
Performance should rely on micro-expressions, natural blinking, realistic breathing, tiny pauses, genuine smiles, subtle nervousness, authentic eye contact, and understated acting. The voice must sound like a real Korean woman in her 20s with soft, warm, conversational delivery—not an AI narrator, influencer, or announcer. Prioritize perfect character consistency, realistic facial animation, natural lip sync, believable body language, authentic environmental audio, and cinematic realism that makes viewers genuinely question whether this was filmed in real life.
Prompt by @frametheory058.
8. Tokyo Boyfriend Vlog
Timeline plus an explicit single-take constraint. · 3,339 characters · format: Prose · timeline: yes
Single unbroken handheld boyfriend vlog take throughout, 30 seconds total opens the prompt, then the time ranges fill it in. The constraint comes first, the schedule second.
Full prompt (3,339 characters)
Single unbroken handheld boyfriend vlog take throughout, 30 seconds total. A realistic personal travel vlog filmed by a boyfriend following his girlfriend during a normal day in Tokyo. Use the woman from the reference image as the main character. Maintain her exact facial identity, hairstyle, facial features, body proportions, and overall appearance throughout the entire video. She must remain the same person in every shot. The camera feels like a real boyfriend holding a small mirrorless camera or phone, not a professional production. Natural handheld movement, imperfect framing, occasional camera shake, spontaneous reactions, authentic everyday moments. The woman does not pose for the camera. She behaves naturally, sometimes forgetting the camera is there. 0-5s: Morning at a small Tokyo apartment. The camera starts recording as the boyfriend casually walks into the room. Soft morning sunlight enters through the window. The woman is sitting near the bed, fixing her hair and preparing for the day. She notices the camera, smiles naturally, laughs, and playfully tells him to stop filming. The camera stays close, slightly shaky, capturing a private everyday moment. 5-10s: Walking through Tokyo neighborhood streets. The boyfriend follows behind her as they leave the apartment. She walks through a quiet Tokyo street, carrying a small bag. Morning shops are opening, bicycles pass by, locals walk along the street. She stops at a convenience store. The camera follows her inside. She looks at different drinks and snacks, turns around and asks the person behind the camera which one she should choose. Natural interaction, casual conversation, realistic body language. 10-18s: Local food experience. The camera follows her through a small Tokyo alley to a cozy local restaurant. She sits down and tries a bowl of ramen or a local dish. The camera captures close handheld moments: her picking up chopsticks, tasting the food, reacting naturally, laughing when the food is hotter than expected. The boyfriend laughs behind the camera. The moment feels unplanned and authentic. 18-25s: Tokyo afternoon exploration. The couple walks through a lively neighborhood. She browses small shops, looks at interesting objects, takes photos, and occasionally looks back at the camera. The camera moves naturally between her face, her hands, the street atmosphere, and small details of daily life. Crowds pass naturally around them. The city feels alive and real. 25-30s: Tokyo night ending. Night falls. The camera follows her through illuminated Tokyo streets. She walks slightly ahead, then turns back and smiles at the camera. They ride a train home. She sits beside the window, watching city lights pass outside. The camera slowly moves closer as she rests quietly, ending like a real personal memory. Visual style: Authentic boyfriend travel vlog footage. Realistic handheld camera movement. Natural lighting. Casual documentary realism. Unplanned everyday moments. Real human expressions and interactions. Slight motion blur, natural exposure changes, realistic camera autofocus adjustments. No commercial advertisement style. No dramatic posing. No perfect cinematic composition. No text overlays. No logos. No face changes. No identity changes. No artificial transitions. No CGI feeling. Stable character consistency throughout.
Prompt by @BubbleBrain.
9. Medieval Shame Walk
Bracket skeleton and timeline together. · 2,832 characters · format: Bracket · timeline: yes
The only prompt here using both. Five bracketed blocks, one of which is [TIMELINE] broken into five ranges — and it is also the most camera-specific prompt here (ARRI Alexa XT, Cooke anamorphic, Steadicam plus handheld).
Full prompt (2,832 characters)
[STYLE + CAMERA + ATMOSPHERE]
Gritty high-end medieval television production look. Shot on ARRI Alexa XT with Cooke anamorphic lenses, mix of Steadicam tracking and handheld inside the crowd. Natural overcast daylight, desaturated dirty palette, visible film grain, realistic crowd physics and fabric movement. No modern polish.
[CHARACTERS]
Central figure: proud middle-aged noblewoman with roughly cropped short blonde hair, wearing a plain rough grey woolen penitential robe that fully covers her, barefoot, pale skin, rigid upright posture that slowly cracks under public condemnation. Stern middle-aged woman in plain brown religious robes walking just behind her, continuously ringing a large heavy iron handbell and chanting in a loud flat voice. Dense crowd of dirty medieval city dwellers of every age and class in period clothing packed on both sides and leaning from windows.
[LOCATION]
Narrow winding cobblestone streets of a medieval coastal city, high stone walls, arched doorways, wooden shutters, mud on the ground.
[TIMELINE]
0-6s: [Steadicam tracking medium-wide from the side] The woman in the plain robe walks steadily forward with forced dignity. The robed woman stays half a step behind ringing the bell and chanting “Shame. Shame. Shame.” Crowd begins to notice, first heads turn, early shouts of “Shame!” rise.
6-12s: [Handheld inside the pack, pushing closer] Crowd presses tighter along the path. Faces show pure contempt. Children point and call out. The woman keeps her chin high but her eyes start to glaze. Bell never stops. Chant continues: “Shame. Shame. Shame.”
12-18s: [Low tracking shot moving with her feet then tilting up to face] Bare feet slap wet cobblestones. A woman leans from a window and shouts “Shame!” The central woman’s jaw tightens, first tears form but she does not break stride. Crowd noise becomes a continuous wall of overlapping “Shame!” mixed with the bell.
18-24s: [Medium close-up handheld, slight shake] Camera stays locked on her face as the controlled mask cracks. Tears finally fall. She stares straight ahead, breathing harder. Behind her the religious woman rings harder and keeps the flat chant. The plain robe shifts with every step under the weight of the stares.
24-30s: [Pull-back Steadicam wide tracking] The full street is visible: wall of bodies on both sides, continuous shouting of “Shame! Shame! Shame!” mixed with the bell. The woman continues walking, posture still upright but now visibly broken, tears streaming, until the frame holds on her isolated figure moving through the condemnation.
[STYLE & QUALITY BOOSTERS]
Exact period production texture of a major series, coherent physics of every body and fabric movement, stable character continuity, natural motion blur, no modern digital cleanliness, no artificial enhancement.
we are cooked 😮
Prompt by @techhalla.
10. Tear-Filled Gunpoint Confrontation
The longest one here. · 5,815 characters · format: Labeled · timeline: no
Nearly all of it is performance direction for one continuous emotional beat. Almost no plot. It is the clearest counter-example to 'long prompts are padded': the length buys nuance in a single face.
Full prompt (5,815 characters)
Use the uploaded character reference image as the only protagonist. Strictly preserve the protagonist’s facial features, hairstyle, skin tone, age, overall aura, and visual identity throughout the entire clip. Generate a 30-second emotionally intense dramatic scene. This is not an action scene or a gunfight. The focus is not on shooting itself, but on the protagonist’s extremely painful and conflicted emotional struggle in the moment before pulling the trigger. The entire clip should center on the character’s performance and facial expression.
The setting should be a quiet, oppressive, high-tension environment with a cinematic atmosphere. It can be a dim interior, an abandoned room, a dark hallway, an empty warehouse, or a nighttime indoor setting. The background should not be too visually busy or distracting. The lighting should feel restrained, dramatic, and moody, with clear contrast on the protagonist’s face. The eyes and the visible moisture of tears must be clearly readable. The overall atmosphere should feel silent, heavy, and emotionally suffocating.
The camera uses an over-the-shoulder shot from behind the other person. Do not show the other person’s face. The foreground should include a blurred shoulder, the back of the head, or the silhouette edge of the other person’s body, creating a clear over-the-shoulder composition. The camera faces the protagonist directly. Focus must stay locked on the protagonist’s facial expression and the gun pointed forward. The gun is aimed toward the camera, meaning it is aimed at “us,” the unseen other person. The camera style should feel like a stable cinematic handheld shot with a subtle sense of breathing. A very slight slow push-in is allowed, gradually tightening on the protagonist’s face to intensify the emotion, but avoid exaggerated camera movement.
The protagonist is holding a gun aimed at the person in front of him. His emotions are extremely complex: anger, hurt, resentment, and the forced determination to do something terrible, mixed with clear reluctance, sorrow, and an inability to truly go through with it. He is not a cold-blooded killer, and he is not simply afraid. This is the feeling of “I have to do this, but I cannot bring myself to do it.” His eyes are red and wet with tears. He is breathing hard. His lips tremble slightly. His jaw is tense. He keeps the gun raised, but his hand shakes subtly. He tries to stay hard, angry, and determined, but that determination keeps breaking under the weight of his pain and reluctance.
At the start of the shot, the protagonist already has the gun raised and aimed directly forward, but his expression is tightly restrained and full of tension. He stares at the person in front of him as if forcing himself not to soften. At first, his face holds down anger and emotional restraint. Then the emotion slowly rises: his eyes become wetter, his breathing grows heavier, his throat tightens, and it feels like there are many unspoken words stuck inside him. Several times, it looks like he is about to pull the trigger. His finger tightens slightly. But every time, he stops at the last possible moment. His eyes constantly shift between rage, heartbreak, pain, resentment, and deep unwillingness, as if he wants to hate the person in front of him, yet cannot truly harm them.
Several key details must be clearly visible:
The gun stays raised the whole time, but his arm and fingers tremble slightly.
He repeatedly tries to force himself to commit, breathing harder, as if trying to make himself do it.
Tears remain held back; he does not fully break down crying, but the tear-filled eyes must be very visible.
When he looks at the other person, the expression is not purely hateful — it contains pain, attachment, heartbreak, and reluctance.
His emotional layers are rich and mixed: anger, hurt, resentment, unwillingness, tenderness, pain, self-forced determination.
Throughout the clip, he never reaches a clean or easy decision. He remains trapped in an overwhelming emotional conflict.
The performance should be restrained, not exaggerated. No screaming, no melodramatic breakdown. What makes the scene powerful is not shouting, but the feeling of someone holding themselves together while being on the edge of collapse. Let the tears gather in his eyes. Let his breathing become heavier. Let his lips shake slightly. Let him look like he may completely fall apart in the next second, but is still forcing himself to hold on. The ending should remain in this most tense unresolved state: he is still holding the gun, still unable to pull the trigger, eyes full of tears, breathing uneven, as if he is still forcing himself to make the decision. The scene should end on this suspended, unresolved, high-pressure moment. Do not resolve it too easily.
The overall style should feel like a cinematic emotional close-up driven by performance, with emphasis on subtle facial acting, breathing, tear-filled eyes, the gun, trembling hands, and the over-the-shoulder composition. Do not turn it into a gunfight. Do not make it feel like an action movie. Do not cut too frequently. Do not add large physical movements. The scene should feel more like a tense dramatic film moment than an action sequence.
For sound design, use only real environmental sound. No background music. You may include faint ambient room tone, distant hollow reverb, the protagonist’s uneven breathing, soft clothing movement, and tiny sounds caused by the trembling hand and body tension. No subtitles. No text. No narration. No UI elements.
Style keywords:
cinematic, emotional confrontation, over-the-shoulder shot, gun aimed forward, tear-filled eyes, restrained breakdown, conflicted emotion, unable to shoot, internal struggle, high tension silence, character close-up, dramatic film scene.
Prompt by @BubbleBrain.
A five-point checklist
If you want a starting point for how to prompt Seedance 2.5, these five defaults carry most of the weight:
- Default to prose. Add structure only when instructions start competing. Most working prompts never needed more.
- Say
handheldunless you specifically want the camera to feel mechanical. - Pin identity first, enumerating attributes, and give every reference a name and one job.
- Write the audio, including what you do not want —
no subtitles,no background music. - Add a timeline only if reordering the middle would ruin the shot.
Then change one thing at a time. The first generation is rarely the good one, and the fastest way to learn what the model responds to is to keep your own record of the prompts that failed and the single edit that fixed them — that record will be worth more to you than any published collection.
One caveat before you copy anything: these prompts were run across different services with different caps. A prompt that produced a 30-second clip on one platform may be truncated on another, so check your platform's duration and resolution limits first. That caveat turns out to be bigger than it sounds — 18 of the 21 prompts measured for this series are longer than the documented 2,000-character prompt cap, and Seedance 2.5 prompt best practices works out what a hard cut at that limit would remove.
If you want to run them without setting up API access, you can try them in our Seedance 2.5 generator — it takes text, image, and audio references, and new accounts get free credits to burn on exactly the kind of one-variable testing described above. The credits and pricing page shows what a longer run costs.
