Why Your 30 Second Prompt Only Produces 12 Good Seconds
You write a long, careful paragraph. A woman walks into a laundromat, loads a machine, takes a phone call, argues with someone on the other end, then sits down alone. You get back a clip where she spends twenty seconds walking through the door, and the phone call never happens. Nothing in the output is broken. The model simply had no schedule, so it spent its whole budget on your first sentence.
This is the defining problem with long form video generation, and it is a writing problem before it is a model problem. Seedance 2.5 generates up to roughly 30 continuous seconds in one pass with synchronized sound, which means a full scene now fits inside a single prompt. What it will not do is guess how you wanted that scene paced.
Seedance 2.5 is available on Atlabs, so everything in this guide is something you can run today without a separate ByteDance account.
Try Seedance 2.5 on Atlabs
Why Prompt Structure Matters More at 30 Seconds
At four seconds, a prompt is a description. There is only one moment, so the model renders that moment and stops. Most prompting advice written in 2024 and 2025 assumes this, which is why it reads like a list of adjectives.
At 30 seconds, a prompt is a schedule. You are no longer describing one image, you are allocating time across several events, and every second you do not account for is a second the model fills on its own. The same shift applies to everything else in the frame. Across four seconds, a character's face barely has time to drift. Across thirty, identity, wardrobe, and set dressing all have room to wander unless something holds them in place.
That is what the reference system is for. Seedance 2.5 accepts up to 50 inputs across images, video, and audio, and the model weights them by the order you supply them, with earlier references carrying more influence. An untagged reference still gets used, it just gets used however the model decides, which is where most consistency complaints actually originate.
Three habits carry almost all of the quality difference: segment the clip into timed beats, tag every reference with the specific thing it controls, and describe the audio as deliberately as you describe the picture. Everything below is built on those three.
The Perfect Prompt Formula
ByteDance's own structure is a director's brief, in this order: subject, action or event, setting and environment, visual style, camera work, sound.
In template form it reads: [Subject] does [action] in [setting]. The image is [visual style]. The camera uses [camera work]. Sound includes [audio].
Broken apart, a working example looks like this. Subject and action: a ceramic artist finishes a pale blue cup on the wheel. Setting: a small studio at dawn with shelves of unfired work behind her. Visual style: soft morning light, wet clay sheen, tidy workbench. Camera: medium shot of the wheel, slow push in to the texture, then a cut to a frontal shelf view. Sound: the low hum of the wheel, clay friction, subtle indoor ambience.
Notice that the camera line names moves rather than moods. "Cinematic" is not an instruction. "Slow push in", "steadicam follow", "whip pan", and "rack focus" are, because each one describes something the frame physically does.
Past roughly ten or twelve seconds, wrap that formula in a schedule. Assign beats: what happens from 0s to 8s, from 8s to 18s, from 18s to 30s. Then, before the beats, give the model one line per reference stating exactly what that asset controls. Image one is identity and clothing. Video one is camera movement and pacing only. Audio one is dialogue and rhythm. Those two additions, beats and roles, fix more output problems than any amount of extra description.
Start building your scene on Atlabs
The Atlabs Workflow for This
Seedance 2.5 sits in the Atlabs model lineup alongside Kling, Veo, Hailuo, and Wan, which means you write the scene once and choose the model that suits it rather than rebuilding the same brief in a different tool.
The workflow that matches this kind of prompting most closely is Script to Video, because it turns the same discipline into a form you fill in rather than a paragraph you have to remember to write correctly.
Script to Video runs in three steps. In Step 1 you either add your script as free form text or add your screenplay. When Atlabs detects screenplay formatting it prompts Screenplay Detected and moves you into a dedicated section, where the left panel shows the AI analysed Narrator Voice, Locations, and Actors, and the right panel breaks the piece down scene by scene with Location Type, Location, and Time of Day. Inside each scene you get individual SHOT descriptions and DIALOGUE blocks, each with a character name, the spoken line, and a delivery tone. That shot list is the interface version of timed beats, and the cast panel is the interface version of reference tagging.
For music led pieces the same logic lives in the Music Video workflow, where the Creative Direction step generates six scene concepts from your track's detected tempo, mood, and genre.
Step by Step Walkthrough
1. Write your beats before you touch the tool. Draft the scene as three or four time blocks on paper first. One action per block, one camera move per block. If a block contains two unrelated events, split it. This is the step people skip, and it is the step that decides whether the output is usable.
2. Open Script to Video and choose your input mode. Use "Add your script" for narration led pieces and "Add your screenplay" when you already have shots and dialogue. If you paste anything with screenplay formatting, Atlabs surfaces Screenplay Detected and switches to the structured view. There is also an AI Script Writer button and a Suggested Scripts panel with three generated starting points if you want a base to edit rather than a blank field.
3. Build the scene breakdown. In screenplay mode, set Location Type as Exterior or Interior, then Location and Time of Day. Add your SHOT descriptions in the order they occur, one per beat, and add DIALOGUE blocks where characters speak, filling in the character name, the line, and the delivery tone. Shots and dialogue can be added or removed per scene, so this is where you tighten pacing rather than rewriting a paragraph and regenerating blind.
4. Set your style. Step 2 gives you Aspect Ratio as 9:16 for TikTok and Instagram, 16:9 for YouTube, or 1:1 for LinkedIn, Twitter, Facebook, and Pinterest. Video Style offers AI Video (Recommended), AI Storyboard, and Upload if you are bringing your own media. Pick the ratio before you finalise framing, because a push in that reads well at 16:9 crops differently at 9:16.
5. Finalize the cast. Step 3 lists the characters pulled out of your script, each with Click to Edit, plus an Add Character option and dropdowns for Country Accent and Narrator Voice. Defining a character here is the platform equivalent of putting your identity image first in a reference stack. It gives the model one fixed answer to who this person is, instead of a fresh guess every shot.
6. Generate on Seedance 2.5, then edit rather than restart. When one stretch is wrong, take the clip into Modify Video, upload it, reference it as @Video1 in the prompt, describe the final video you want, and attach up to four reference images for style guidance. Keep Original Audio has an on and off toggle so a picture fix does not cost you a good soundtrack.
Open Script to Video on Atlabs
Why This Approach Works Well on Atlabs
The first reason is model routing. Prompt discipline is portable, but the aesthetic is not, and the model that renders a stylized character closeup well is rarely the one that renders a wide coastal exterior well. Atlabs runs Seedance 2.5 for long continuous takes and reference heavy scenes, Seedance 2.0 for stylized content, anime, and character closeups, Kling 3.0 and Kling 2.6 for cinematic motion and smooth movement, Google Veo 3.1 for photorealism and establishing shots, Hailuo 2.3 for high motion and anime adjacent visuals, and Wan 2.6 for open source cinematic output. One workflow, a different model behind it, no separate API accounts to hold together.
The second is that the interface enforces the structure instead of trusting you to remember it. A shot list with SHOT and DIALOGUE blocks per scene cannot quietly collapse into one run on paragraph, which is the single most common failure in long form prompting.
The third is that motion and identity are separable. Motion Control takes a reference video between 3 and 30 seconds, applies its motion to a character image you upload, and keeps the motion entirely under the reference's control rather than the prompt's, with an optional prompt field for background and scene detail. That is the same idea as saying video one is for camera movement only, except the tool guarantees it. Lip Sync works the same way for performance, taking a character image or video plus an audio file between 2 and 120 seconds.
Example Prompts
1. The timed one take. A man in a charcoal three piece suit walks the length of a hotel lobby at night, marble floor, brass fittings, moody amber tungsten light. 0s to 8s: steadicam follow from behind as he crosses the floor, staff turning to watch him pass. 8s to 18s: he pushes through the brass doors into the street, whip pan to reveal a crowd of photographers, flashbulbs strobing. 18s to 30s: slow push in on his face, half smile, rack focus to the marquee behind him. Sound is crowd murmur, camera shutters, distant traffic, no music. (Best routed through Seedance 2.5) Try this prompt in Atlabs Script to Video
2. The candid documentary montage. Multi shot observational footage in a self service laundromat late at night, seven shots, handheld with imperfect framing, no single continuous take. A woman in an oversized cotton T shirt loads a machine, takes a phone call, and sits alone under buzzing fluorescent tubes. Wide from the doorway, then over the shoulder at the drum, then a tight profile at the window. Sound is raw production audio only, machines, traffic outside, no score. (Best routed through Seedance 2.5) Try this prompt in Atlabs Script to Video
3. The beat synced performance. A singer in a red satin jacket performs on a rain slick rooftop at blue hour, neon signage behind her, city haze catching the light. Camera orbits slowly for the verse then cuts harder on the chorus, cuts landing on the beat. Shallow depth of field, cool highlights against warm skin tones, cinematic grade. (Best routed through Kling 3.0) Try this prompt in Atlabs Music Video
4. The character consistency lock. A woodworker walks through a sawdust lit workshop and stops beside a finished walnut chair. The camera tracks from behind, moves into a side view, and ends on a close up of her hands on the armrest. Preserve the face, the clothing, the product shape, and the room layout exactly. No additional people, no subtitles, no added music, no camera cuts. (Best routed through Seedance 2.5) Try this prompt in Atlabs Modify Video
5. The physics test. A glass marble rolling fast along a chain reaction track built from dominoes, brass gears, and small wooden ramps on a workbench. One continuous smooth shot, macro depth of field, hard side light picking out the edges. The marble obeys real gravity and momentum throughout. Crisp clicks, rolling glass, no music, no cuts. Try this prompt in Atlabs Script to Video
6. The motion transfer. A dancer in a loose linen shirt performs in an empty concrete gallery, tall windows casting hard rectangles of afternoon light across the floor, dust visible in the beams. Motion comes entirely from the reference clip, background and lighting from the prompt. Wide static frame, no camera movement, natural room reverb. (Best routed through Hailuo 2.3) Try this prompt in Atlabs Motion Control
Run these prompts on Atlabs
Pro Tips
Put the face first. Reference order carries weight, and the model treats earlier inputs as more authoritative. Identity goes at position one, wardrobe next, location after that, style grade last. If your character keeps drifting, check the order before you rewrite the prompt.
Say what each reference is not for. "Video one is for camera movement and pacing only" prevents the model borrowing that clip's lighting and wardrobe as well. Exclusions are as useful as instructions, which is why "no subtitles, no added music, no extra people" belongs at the end of most reference heavy prompts.
Describe the sound or accept whatever arrives. Audio is generated in the same pass, so silence in the prompt is not a request for silence, it is a blank cheque. Name the ambience, name the language any dialogue is spoken in, and if you want nothing under the scene, write that explicitly as production sound only.
FAQ
Is Seedance 2.5 available on Atlabs? Yes. It runs inside the same workflows as the rest of the lineup, so you can build a scene in Script to Video or Music Video and choose Seedance 2.5 for the generation without setting up a separate account anywhere else.
How long can a Seedance 2.5 clip be? Native continuous generation runs to roughly 30 seconds in one pass. Stitching shorter clips together is no longer required, which is the main practical difference from earlier versions.
How many references can I use? Up to 50 in total across images, video, and audio, commonly split as around 30 images, 10 video clips, and 10 audio files depending on the platform. Order them by importance, since earlier references carry more weight.
What is the fastest single improvement to a long prompt? Add timestamps. Splitting a 30 second brief into three or four timed blocks fixes more pacing and continuity problems than any other change, and it takes about a minute.
Should I use Seedance 2.5 or Seedance 2.0? Choose 2.5 when the clip is long, when several references have to stay locked, or when the whole scene needs to hold together in one pass. Choose 2.0 for shorter stylized work, anime, and character closeups where you want a quick result and continuity is not being stretched across half a minute.
Can I fix part of a clip without regenerating everything? Yes. Modify Video takes an existing clip between 3 and 10 seconds, referenced as @Video1 in your prompt, along with a description of the final video and up to four reference images for style. Keep Original Audio can stay on so a picture change does not cost you the sound.
Final Verdict
Long form prompting rewards planning over description. Write the beats before you write the sentences, tell the model exactly what each reference controls, and name the audio with the same care you give the camera. Those three habits carry across every model you will use this year, which is why they are worth building now rather than relearning at each release.
The part that is harder to carry across is everything around the prompt: the shot list, the cast, the aspect ratio, the model choice, and the repair pass when one stretch comes back wrong. That is what a full workflow is for. Script to Video gives you the structure, Music Video gives you the track led version of it, and Modify Video and Motion Control handle the fixes without sending you back to the beginning, with Seedance 2.5 sitting behind all of them.










