Seedance 2.5 White Model Reference Prompt Guide
Seedance 2.5 white model reference prompt guide showing how to write text for 3D clay passes, handle camera moves, and avoid common generation errors.
The team behind Seadanse. We run the models we write about, and we publish the specs, prices and limits the marketing leaves out.

TL;DR: A white model reference pairs an untextured 3D layout clip with text. It's used to guide scene layout and camera travel across three scenes and seven prompt examples.
A white model reference is a simple animation made from plain 3D shapes. In 3D work, it's called a blockout or clay pass.
It shows where objects sit and how the camera moves through the space.
When you send this gray preview to Seedance 2.5, you don't have to worry about surface details right away.
The video sets the camera path and layout. Your text describes materials, lighting, and mood.
In our tests on September 10–11, 2026, we built three five-second gray scenes in Blender 5.2.1 LTS at 24 fps (120 frames).
We tested eight runs at 480p, including a first test, an accepted revision, and six prompt variations. Here's how to write prompts that work with your 3D layouts.
What a white model reference is, and what it'll guide
A white model reference uses basic gray geometry without realistic textures or reflections. It doesn't need complex materials.
Its main job is to show object positions and camera motion.
When you attach this gray clip, it guides layout, framing, and camera paths.
As explained in documented white-model guidance, the reference sets staging before it's given materials and light.
It doesn't lock your shapes into rigid limits. The AI interprets your layout rather than copying exact mesh surfaces.
In our tests, blocky shapes morphed into stylized objects while following the camera path from the clay clip.
You can see the difference between source layout and output in our baseline test below:
The gray source establishes the path. The generated video builds the final scene on top of it.
When a white model is the wrong input
Clay references give you structural control, but building 3D layouts takes extra time. For some shots, that step isn't needed.
Skip a white model reference when:
- You want an organic shot with loose, unpredictable camera drift.
- You want surreal visual effects that shift without fixed limits.
- You want a quick text-only concept test before planning exact blocking.
- You don't have access to 3D tools or keyframing software.
When you need natural human performance, writing descriptive text or using live-action video clips can fit your goals better than basic digital mockups.
The prompt formula: Subject, Action, Material, Light, Camera
When you write prompts with a clay pass attached, use a five-part formula: Subject, Action, Material, Light, and Camera.
| Formula element | What it does in your prompt | Practical example |
|---|---|---|
| Subject | Identifies what fills the frame | Twelve thick sculptural pieces across three depth layers |
| Action | Describes visible movement or stillness | Pieces remain completely still under balanced conditions |
| Material | Defines surface textures, finishes, and opacity | Smooth finish of opaque white glazed ceramic |
| Light | Directs illumination angle, softness, and shadows | Restrained afternoon illumination enters from the side |
| Camera | Sets framing, or stays omitted for clay motion | Leave out movement words to let the clay pass guide the view |
Describe visible details rather than emotional mood. When you create a Seedance video with a reference attached, keep words focused on physical elements.
Tell the tool what surfaces look like and where light hits them. Leave out vague praise like "photorealistic" or "hyper-detailed."
What changes when a clay pass is attached
In text-only prompts, your words must describe subjects, materials, lighting, and camera paths at the same time. If you leave out camera directions, the AI picks a path for you.
Once you attach a clay reference, the video carries the camera path and staging. You don't need long sentences describing pans, tilts, or dolly moves.
Writing camera instructions that fight your 3D video causes conflicts. Your text prompt should describe objects, surfaces, and lighting, while letting the reference clip steer the view.
You can see more examples of effective prompt phrasing in our Seedance prompt examples library.
The optional controls (reference roles, timing beats, continuity, audio, ending)
You can add extra controls to your prompt when a shot needs specific cues. Use them to solve real staging needs rather than filling out a template.
- Reference roles: Attach still images to direct colors, patterns, or costumes on top of your 3D layout.
- Timing beats: Single-event shots often need no timestamps. If an action happens mid-shot, describe it plainly without second counters.
- Visual continuity: Mention fixed landmarks like doorframes, floor joints, or street markings to anchor scale.
- Ending conditions: If a camera move must stop and hold, state that the subjects and viewpoint rest stationary at the end.
- Audio details: Describe scene sounds directly in your prompt text. While the interface lacks reference-audio uploads, the AI can still generate native audio based on your words.
Keep these extra lines brief so the generator stays focused on your main scene elements.
Camera language you still need (shot size, lens feel, movement)
Even when a clay pass guides your camera path, standard camera terms help you frame shots in Blender and write clear text. The RunDiffusion Seedance 2.5 prompt guide highlights how clear framing terms describe scenes.
Shot size and framing
- Wide shot: Shows the whole setting and where objects sit relative to each other.
- Medium shot: Focuses on main subjects while keeping nearby surroundings in view.
- Close-up and macro: Focuses on small details, textures, or miniature props.
- Overhead and POV: Provides top-down layouts or first-person viewpoints.
Lens feel and camera motion
Perspective depends on where you place your camera. Changing your focal length changes your field of view. When you step back and use a longer lens to match framing, the background appears compressed.
- Wide-angle feel: Gives a broader field of view, making nearby and distant objects feel further apart.
- Telephoto feel: Narrower field of view that appears to compress depth when shot from further away.
- Dolly and tracking: The camera physically pushes in, pulls back, or travels alongside a subject.
- Orbit, rise, and tilt: The camera turns around a central point, moves upward, or angles up and down.
- Locked shot: The camera stays completely still while action happens in front of it.
Understanding these terms helps when you set up your scene using the Seedance 2.5 Blender add-on.
Prompt 1. Name the material and the light, leave the camera alone
This baseline test uses Scene A, an untextured 3D scene where geometric shapes sit across three depth layers. In the source animation built with our Blender export integration, the camera glides sideways and doesn't move from frame 85 through frame 120, bringing the pieces into a butterfly alignment.
In our five-second test (Run 03-r1), the lateral glide reveals the butterfly silhouette and shows a visible ending hold. The early wire-like fragments inflate and morph into ceramic wings while following the camera path from the clay pass.
A quiet, human-scale art gallery measures roughly six by seven metres with clean rectangular doorframes and straight floor joints. Twelve thick, rounded sculptural pieces and one narrow central element rest stationary across three distinct depth layers in the room. Every sculptural element features a smooth finish of opaque white glazed ceramic, catching soft highlights along defined geometric contours. Restrained afternoon illumination enters from the side, casting clear, subtle shadows across the floor without heavy haze or obscuring blur. All sculptural components remain completely still in their fixed physical positions under even, balanced ambient daylight.Scene A baseline: Smooth white glazed ceramic shapes align into a butterfly silhouette as the camera glides sideways, with visible ending hold and shape morphing.
Prompt 2. Swap the material, keep everything else
To test surface control, we kept the exact Scene A clay reference and changed only the material phrase from white ceramic to polished copper metal.
In Run 04, the output applied polished copper surfaces across all pieces. The lateral camera movement remained intact, but shape inflation was still present, and the lighting appearance differed slightly from the baseline. This shows that the material phrase changed surface finish in this take, though geometry still morphed.
A quiet, human-scale art gallery measures roughly six by seven metres with clean rectangular doorframes and straight floor joints. Twelve thick, rounded sculptural pieces and one narrow central element rest stationary across three distinct depth layers in the room. Every sculptural element features a smooth finish of polished copper metal, catching soft highlights along defined geometric contours. Restrained afternoon illumination enters from the side, casting clear, subtle shadows across the floor without heavy haze or obscuring blur. All sculptural components remain completely still in their fixed physical positions under even, balanced ambient daylight.Run 04 result: Polished copper metal surfaces appear on the sculptures, with lateral camera tracking preserved alongside shape inflation and subtle lighting shifts.
Prompt 3. Ask for a camera the clay pass doesn't have
What happens when your text prompt contradicts the movement in your clay reference? In Run 05, we used the moving Scene A reference but added a final sentence telling the camera to stay fixed in one position.
The output followed the reference video's lateral tracking movement toward the butterfly view and ignored the text instruction. While this single conflict result doesn't establish a universal instruction priority, it demonstrates that the reference video's motion carried through despite the text request.
A quiet, human-scale art gallery measures roughly six by seven metres with clean rectangular doorframes and straight floor joints. Twelve thick, rounded sculptural pieces and one narrow central element rest stationary across three distinct depth layers in the room. Every sculptural element features a smooth finish of opaque white glazed ceramic, catching soft highlights along defined geometric contours. Restrained afternoon illumination enters from the side, casting clear, subtle shadows across the floor without heavy haze or obscuring blur. All sculptural components remain completely still in their fixed physical positions under even, balanced ambient daylight. The camera stays at one fixed position throughout the entire shot.Run 05 result: The camera continues its lateral track toward the butterfly view, following the Scene A clay clip rather than the prompt's fixed-camera clause.
Prompt 4. Let an appearance image carry the look (06)
In Scene B, an untextured layout shows a miniature rescue with cereal rings, a toy helicopter, and tiny figures in a bowl. Note that the upright floating ring already exists in the 3D clay source.
For Run 06, we attached both the clay video and a 2D appearance image showing blue wave patterns, red dots, and vehicle styling.

The generated clip picked up the blue waves, red dots, exterior patchwork, and red/blue helicopter styling. The explorer figure appeared on the upright ring. Because the prompt specified a neighbouring upright ring and wording also differed from Run 07, this test represents a combined image and text variation rather than an isolated image ablation.
A miniature rescue scene unfolds across an arrangement of ringed cereal pieces resting in still liquid inside a wide bowl. On the outer rim, a small transport helicopter rests parked with stationary rotor blades and grounded skids. Beside the parked aircraft, an eight-millimetre miniature pilot stands firmly in place, offering a slight, reassuring nod toward a companion. On a neighbouring upright cereal ring, a stranded miniature explorer pauses steadily and returns a subtle, clear hand wave. The quiet exchange conveys relief and calm focus across the tabletop setting without any sudden movement or physical relocation.Run 06 result: Blue patterns, red dots, and vehicle colors from the image transfer onto the 3D layout, with the character placed on the upright ring.
Prompt 5. Try a different geometry and path (07)
For Run 07, we used the same Scene B clay pass without an appearance image. We guided the look using descriptive text for whole milk and toasted cereal rings.
The camera backs out through the upright ring hole and rises to reveal the bowl and helicopter on the rim. The upright ring exists in the clay source rather than being invented by the AI. Ring position and size differ slightly from the clay reference at matched times.
A tabletop breakfast scene hosts an eighteen-centimetre plain glazed ceramic bowl containing still whole milk and six toasted cereal rings. One ring stands vertically in the pool, forming an open circular archway above the milky surface. Perched on the wide curved rim of the bowl, a three-centimetre painted helicopter sits grounded with distinct tail boom, landing skids, and unmoving rotor. Nearby, an eight-millimetre miniature pilot stands beside the aircraft, while an eight-millimetre stranded explorer waits on an adjacent cereal ring. Warm morning light illuminates the scene, catching linear grooves in the textured placemat below.Run 07 result: Warm morning lighting, cereal textures, and milk applied to Scene B clay geometry using text instructions alone without an appearance image.
Prompt 6. Test a compound camera move (08)
Scene C tests a complex camera path: a low push along the street followed by a lateral upward rise around a frozen bicycle courier suspended above a car.
In Run 08, the output rendered the urban street details, suspended delivery papers, and late-afternoon light while following the compound push-and-rise trajectory around the frozen cyclist.
A dramatic urban moment freezes completely over rough street asphalt marked by worn white pedestrian crossing stripes and three evenly spaced metal bollards. Suspended motionless above the hood of a stationary passenger car, a bicycle courier hangs mid-air alongside a lightweight bicycle, displaying one prominent foreground wheel. Six large delivery paper sheets remain locked at varying depths around the suspended cyclist, caught in static mid-flight. The courier wears a simple helmet, fitted jacket, and compact delivery bag. Grounded late-afternoon sunlight casts crisp directional shadows, preserving the intense physical tension while keeping all vehicles, debris, and figures strictly motionless.Run 08 result: Urban materials, frozen papers, and sunlight mapped onto a compound push-and-rise camera path around suspended 3D geometry.
Control. Repeat Prompt 1 with no reference (09)
To see what happens without a white model, Run 09 ran the exact baseline prompt with zero video or image references attached.
The prompt doesn't mention camera movement or butterfly shapes. Without the clay clip, the generator created rounded ceramic vessels in a gallery with a slow lateral drift and no butterfly alignment. Run 09 went through the KIE provider endpoint rather than the Apimart endpoint used in earlier runs, which adds an extra variable alongside the missing reference.
A quiet, human-scale art gallery measures roughly six by seven metres with clean rectangular doorframes and straight floor joints. Twelve thick, rounded sculptural pieces and one narrow central element rest stationary across three distinct depth layers in the room. Every sculptural element features a smooth finish of opaque white glazed ceramic, catching soft highlights along defined geometric contours. Restrained afternoon illumination enters from the side, casting clear, subtle shadows across the floor without heavy haze or obscuring blur. All sculptural components remain completely still in their fixed physical positions under even, balanced ambient daylight.Run 09 control: Without a clay reference, the prompt produces generic ceramic vessels with slow camera drift and no butterfly alignment.
How many references you can actually attach
Documentation sources sometimes state conflicting limits for reference inputs. Within the Seadanse add-on, the verified limits are:
- Reference video: Exactly one MP4 or MOV file (2 to 30 seconds, up to 100 MB).
- Appearance images: Up to 30 images (PNG, JPG, or WebP, up to 30 MB each).
- Audio references: 0 files (the interface doesn't support audio uploads).
Adding appearance images doesn't increase your generation quote. But the length of your reference video directly affects your final quote calculation.

Common failures and the fix for each
Working with 3D references introduces specific failure modes when prompts and video inputs interact. Here is what to check or try for each case.
1. Shape morphing and softening
- What you see: Sharp, boxy 3D blocks turn into soft, rounded objects during generation.
- What to check: The model interprets blockouts as layout guides rather than exact mesh boundaries.
- What to try: Try describing rigid surfaces, sharp seams, and mechanical materials in your prompt text to encourage crisper outlines.
2. Conflicting camera motion
- What you see: The camera jitters, drifts unexpectedly, or ignores text cues.
- What to check: Your text prompt might include camera motion phrases that contradict the clay video's movement.
- What to try: Remove camera movement words from your text prompt and let your clay video guide the viewpoint.
3. Missed alignment points
- What you see: Objects arranged in depth fail to line up cleanly at the intended viewpoint.
- What to check: The camera might move through the alignment zone without enough resting time.
- What to try: Extend the rest at the alignment position in your 3D animation (in our Scene A source, holding stationary from frame 85 to 120 allowed the alignment to register).
A testing order that doesn't burn credits
Generating AI video uses processing credits. Follow these steps to test your ideas efficiently before spending credits:
- Check local playback: Scrub your 3D layout in the Blender timeline to check camera motion and pacing.
- Make a local clay pass: Export your gray preview locally in Blender. Local export runs on your computer at no credit cost.
- Draft your prompt: Use the five-part formula to describe subjects, materials, and lighting.
- Run a standard-definition draft: Generate a five-second test at 480p to check layout and surface details.
- Inspect the results: Check object framing and motion before exporting higher-resolution passes at 720p or 1080p.



For more detail on path planning, check our guide to 3D camera paths in Seedance.
What still needs attention
While white model references give solid composition control, several technical limits remain:
- Shape inflation: Sharp geometric blocks often soften into rounded forms during generation.
- Complex rigging: Detailed mechanical linkages like landing gear or folding hinges can drift from their 3D paths.
- Contact points: Precise physical contact, like feet on uneven terrain or hands grasping handles, remains hard to lock completely.
Designing scenes with these limits in mind helps you get usable results. For practical shot concepts that fit these tools, explore ideas for Blender projects.
Start from your own clay pass (CTA)
You can build a simple blockout in Blender today and use it to direct your next video shot.
- Get the Blender add-on. Drag the ZIP file into Blender 4.2 LTS or newer and confirm installation.
- Open the Seadanse panel in the 3D Viewport sidebar (press N). Complete the first browser code authorization.
- Build your basic scene geometry and keyframe your shot.
- Export your clay pass locally or load an existing clip into the panel.
- Enter your prompt text, review your settings, and click Generate.

The Seadanse panel in Blender's sidebar gives you tools to export clay passes, edit prompts, and generate videos.
For a full walkthrough of setting up your 3D blockout scene, see our step-by-step blockout tutorial.
FAQ
What is a white model reference?
A white model reference is an untextured 3D animation (often called a clay pass or gray layout). It provides spatial blocking, object placement, and camera paths to the Seedance 2.5 video generator, letting you plan framing before generating final materials and lighting.
What 3D software or blockout tool should I use?
You can use Blender 4.2 LTS or newer with the download the Blender add-on integration to set up keyframes and render clay passes directly. You can also load existing MP4 or MOV blockouts made in any 3D software using the panel's file picker.
Are white model references best for animations, product demos, or cinematic scenes?
White model references work well for all three formats. They are especially helpful for product reveal shots, architectural tours, and planned cinematic scenes where camera precision matters more than random text drift.
How long can my reference video and generated output be?
The add-on accepts reference videos between 2 and 30 seconds (up to 100 MB). Generated outputs can be set from 4 to 30 seconds at 480p, 720p, or 1080p resolution.
Can I attach reference audio files to my generation?
No. The add-on interface doesn't support reference audio uploads. You can describe scene sounds directly in your text prompt, though there's no guarantee the sound will match your description exactly, so you'll want to review the generated audio.



