Text defines intent
Use language for action, story beat, camera instruction, exclusions, and the final delivery format.
Plan text, image, audio, and source-video inputs around the creative decision each reference should control.
Use text when you want the model to invent the subject, setting, motion, and camera from a blank starting point.
More references do not automatically improve the result. Clear responsibility prevents contradictory direction.
Use language for action, story beat, camera instruction, exclusions, and the final delivery format.
Use high-signal frames for identity, product detail, wardrobe, composition, or visual style.
Use audio references for voice, music, timing, or emotional cadence when the selected workflow supports them.
Use source clips for movement, timing, editing, extension, and actions that are difficult to specify with text alone.
Resolve conflicts before they reach the generation queue.
Write one sentence describing what the finished clip should communicate.
Name what each file controls and what it must not change.
Call out identity, product geometry, logos, camera axis, or timing that must remain stable.
Validate the direction with a short clip before requesting a larger batch.
A clean input set is often more valuable than a longer prompt.
Use text for open exploration, an image for visual consistency, audio for voice or rhythm, and source video when motion or timing already exists.
No. Add an input only when it controls a specific part of the result. Extra references can create conflicting instructions.
Use clear identity references, avoid conflicting faces or wardrobes, state what must remain unchanged, and revise one visual variable at a time.
This site does not offer a SynthID, C2PA, or watermark-removal workflow. Keep applicable disclosure and provenance information with generated media.
Move from input planning to generation and reusable prompt structure.