Shot descriptions, camera moves and the clip lengths that actually hold together. Below are 7 copy-ready prompts. Fill in the [BRACKETS], copy, and paste into ChatGPT, Claude, Gemini or any capable assistant.
Video models are where the gap between demo and daily use is widest. They are genuinely impressive in short bursts and fall apart over length, which shapes how you prompt.
The 7 prompts
Write a prompt for an AI video clip
Describe a shot so a video model can produce it.
Write a video generation prompt. WHAT I WANT TO SEE: [DESCRIBE THE SHOT] CLIP LENGTH: [SECONDS] MODEL: [WHICH VIDEO MODEL] WHAT IT IS FOR: [b-roll / product / social / concept / film] ASPECT RATIO: [DIMENSIONS] Video prompts need everything an image prompt needs plus motion, and motion is where they fail. Build it in layers: 1. THE SUBJECT AND SETTING - as specific as an image prompt. 2. THE CAMERA - static, pan, tilt, dolly, tracking, handheld, or crane. State the direction and speed. Ambiguous camera instructions produce drifting, unstable shots. 3. THE SUBJECT MOTION - what moves in the frame, how fast, in which direction. One clear motion beats several. 4. THE LIGHTING - as for an image, plus whether it changes during the clip. Usually it should not. 5. THE SHOT DURATION AND PACING - what happens across the clip. In a few seconds, one thing happens. Then: - THE ASSEMBLED PROMPT in my model's conventions - THE ONE-MOTION RULE - current video models handle a single clear motion far better than several. If my description has two, say which to cut and why. - WHAT WILL GO WRONG - the known weak points: hands and faces in motion, text, object permanence (things appearing and disappearing), physics, consistent detail across frames, and anything requiring a specific count. - WHAT NOT TO ASK FOR - dialogue and lip sync, precise choreography, and anything needing narrative across a cut. These are not clip-level capabilities. - THE SHOT LIST APPROACH - if what I described is really several shots, break it into separate clips to be assembled in an editor. This is almost always better than one complex prompt. - THE ITERATION PLAN - what to change first if the motion is wrong versus if the subject is wrong. Be honest about capability limits rather than writing an elaborate prompt for something the model cannot do.
What you get: A layered video prompt with camera and subject motion separated, the one-motion rule applied, and explicit capability limits.
Tip: The one-motion rule is the practical constraint. A prompt asking the camera to pan while the subject turns and the light changes produces incoherent footage.
Plan a sequence of generated clips
Assemble multiple clips into something coherent.
Help me plan a video made from generated clips. WHAT THE VIDEO IS: [PURPOSE AND MESSAGE] TOTAL LENGTH: [DURATION] STYLE: [LOOK AND FEEL] MODEL: [WHICH] EDITING SOFTWARE: [WHAT YOU WILL ASSEMBLE IN] AUDIO PLAN: [music / voiceover / both / none] Produce: 1. THE SHOT LIST - each clip: number, duration, what it shows, why it is there, and what it cuts to. Keep clips short; generated footage is more convincing in brief cuts than in sustained shots, where artefacts accumulate. 2. THE CONSISTENCY SPECIFICATION - what must hold across every clip: lighting quality, colour palette, level of realism, camera height and lens character, and subject treatment. Write it once and include it in every prompt verbatim. 3. THE PROMPTS - one per shot, combining the consistency specification with that shot's content. 4. WHERE CONSISTENCY WILL FAIL - be honest. Generated clips will not match perfectly. Plan around it: cut between different subjects rather than the same subject from different angles, use colour grading in the edit to unify, and avoid shots that would reveal a continuity break. 5. THE CUTTING PLAN - generated footage benefits from faster cutting. Say the rough shot lengths, and where a cut hides a problem. 6. WHAT SHOULD NOT BE GENERATED - shots that would be better as stock footage, a screen recording, a static graphic, or text on screen. Being selective about which shots are generated is what makes the result look deliberate rather than uniformly synthetic. 7. THE AUDIO - what carries the piece. Generated video with good sound reads far better than the reverse. Say what music or voiceover should do, and where sound design would cover a visual weakness. 8. THE REJECT RATE - realistically how many generations per usable clip, and therefore how long this takes. 9. THE ASSEMBLY ORDER - generate the most important shot first, establish the look, then match the rest. 10. THE GRADE - unifying the clips in the edit with a single colour treatment does more for coherence than any prompt consistency.
What you get: A shot list with per-shot prompts, a shared consistency specification, a cutting plan that hides weaknesses and an honest reject rate.
Tip: Point 10 does more than anything at the prompt stage. A single grade applied across mismatched clips in the edit unifies them in a way no prompt wording achieves.
Turn a still image into a moving shot
Animate a photograph or generated image.
Help me animate this image. THE IMAGE: [DESCRIBE WHAT IS IN IT] WHAT SHOULD MOVE: [YOUR INTENTION] WHAT MUST STAY STILL: [WHAT TO PRESERVE] CLIP LENGTH: [SECONDS] TOOL: [IMAGE-TO-VIDEO MODEL, OR PARALLAX/2.5D TOOL] PURPOSE: [WHERE IT WILL BE USED] Produce: 1. THE MOTION CHOICE - decide which kind of movement suits this image: - CAMERA MOVE ONLY: a slow push, pull or pan across a still scene. The most reliable, least likely to distort anything, and often the most effective. - AMBIENT MOTION: small natural movement - water, cloth, hair, steam, foliage. Convincing when subtle. - SUBJECT MOTION: the subject acts. Highest risk of distortion, especially with faces and hands. Recommend one for my image and say why. 2. THE PROMPT - describing the motion specifically: what moves, in which direction, how fast, and what stays fixed. Include what should not move; models will animate things you did not intend. 3. THE SUBTLETY RULE - less motion looks better. A barely perceptible drift reads as cinematic; vigorous motion reveals every artefact. Say what magnitude suits this image. 4. WHAT WILL DISTORT - based on my image description, the elements most likely to warp: faces, hands, text, straight architectural lines, repeating patterns, and anything with fine detail. Say which are present and whether to crop them out or restrict motion near them. 5. THE DURATION LIMIT - artefacts accumulate over time. Say what length this image can sustain before degrading, and that a short loop is usually better than a longer clip. 6. THE LOOP OPTION - if it can loop seamlessly, how to prompt for it and where a loop is more useful than a one-off clip. 7. THE ALTERNATIVE - for many purposes, a simple camera move added in an editor over a high-resolution still is more reliable and looks better than generated motion. Say whether that applies here. Generated motion is worth it when something in the scene genuinely needs to move. 8. THE SELECTION APPROACH - generate several, and judge them at full speed and full size rather than frame by frame.
What you get: A motion type chosen for the image, a prompt specifying what stays fixed, distortion risks identified and an honest alternative.
Tip: Point 7 is worth considering first. A slow digital push on a sharp still often looks better than generated motion and takes two minutes in any editor.
Write a brief for video content you will shoot
Plan a real video shoot properly.
Write a shoot brief. WHAT THE VIDEO IS FOR: [PURPOSE AND WHERE IT GOES] MESSAGE: [WHAT THE VIEWER SHOULD TAKE AWAY] LENGTH: [DURATION] WHO IS IN IT: [PRESENTER, SUBJECTS, OR NOBODY] WHERE: [LOCATION] EQUIPMENT: [WHAT YOU HAVE - phone, camera, lights, audio] TIME AVAILABLE ON THE DAY: [HOURS] BUDGET: [IF ANY] Produce: 1. THE SHOT LIST - every shot needed: number, description, framing, camera position, duration, and what it is for in the edit. Include the coverage shots that make editing possible: cutaways, b-roll, and a wide establishing shot. Running out of cutaways in the edit is the most common regret from an under-planned shoot. 2. THE SHOOTING ORDER - grouped by setup, not by edit order. Every lighting or location change costs time; shoot everything in one setup before moving. 3. THE AUDIO PLAN - the most neglected element and the one viewers notice most. Bad audio loses an audience faster than bad picture. What to record with, where to place it, how to monitor it, and what to do about room noise. If the only audio option is the camera microphone, say what that means and how to mitigate it. 4. THE LIGHTING PLAN - for my equipment and location. If it is available light, where to position the subject relative to the window and what time of day. If a window is behind the subject, say so now rather than discovering it in the edit. 5. THE SHOT DISCIPLINE - hold each shot longer than feels necessary, record before and after the action, and get a clean version of every take. Editors need handles. 6. THE CHECKLIST FOR THE DAY - equipment, batteries, cards, backup, and the settings to confirm before rolling: resolution, frame rate, audio levels, focus mode, and white balance. 7. THE TIME BUDGET - allocate my stated hours across setup, shooting and wrap. Be realistic; setup always takes longer than planned. If the shot list does not fit the time, say what to cut. 8. WHAT TO GET IF TIME RUNS OUT - the shots that are non-negotiable, ranked. 9. THE BACKUP PLAN - what to do if the location, the weather, the presenter or the equipment fails. 10. THE ON-SET CHECK - review footage on set before wrapping. Discovering a problem in the edit means reshooting.
What you get: A shot list including coverage, a setup-grouped shooting order, an audio-first plan, a time budget and a ranked must-get list.
Tip: Point 3 is the one to act on. Viewers tolerate mediocre picture and abandon bad sound, and audio is the cheapest thing to fix before the shoot.
Edit a rough cut into something watchable
Turn raw footage into a finished video.
Help me edit this video. WHAT I HAVE: [DESCRIBE THE FOOTAGE - what was shot, how much, what quality] WHAT IT IS FOR: [PURPOSE AND PLATFORM] TARGET LENGTH: [DURATION] MY CURRENT CUT: [DESCRIBE IT, AND WHAT FEELS WRONG] SOFTWARE: [WHAT YOU USE] SKILL LEVEL: [beginner / comfortable / experienced] Produce: 1. THE STRUCTURE - what the video should be, in order, with rough timings. Every section needs a reason to exist. 2. THE OPENING - the first five seconds decide everything. What should be there, and specifically what should not: a logo animation, a slow build, an introduction, or the presenter saying hello. 3. WHAT TO CUT - the specific categories: - The beginning of every clip before the action starts - Pauses, ums, restarts, and throat-clearing - Anything said twice - Explanation of what is about to be explained - The section where the point is already made and it continues anyway Say roughly what proportion of raw footage typically survives, so my target length is realistic. 4. THE PACING DIAGNOSIS - from what I described feels wrong, the likely cause: shots held too long, no variation in shot type, a flat middle with no change, or too much information with no pauses to absorb it. 5. THE CUT POINTS - where to cut. Cut on action, cut on the end of a thought, cut before someone finishes rather than after. Never let a shot sit after its purpose is served. 6. THE AUDIO PASS - the single biggest quality gain available in the edit: levelling, removing noise, cutting breaths, and making sure music sits under speech rather than competing. Do this before any visual polish. 7. THE MUSIC DECISION - whether it needs any, what it should do, where it should start and stop, and the licensing question. 8. TEXT AND GRAPHICS - what genuinely helps the viewer: names, key numbers, and section markers. What does not: decorative transitions, animated logos, and text that repeats what is being said. 9. THE TRANSITIONS - hard cuts for almost everything. Say the rare cases where anything else is warranted. 10. THE FINAL CHECK - watch it once at full length without touching anything, once at double speed to spot pacing problems, and once with the sound off to check it reads visually. Then export for the platform.
What you get: A structure with timings, a specific cut list, a pacing diagnosis, an audio-first polish order and a three-pass final check.
Tip: Point 6 is the highest-return edit work. An hour spent on audio improves a video more than a day spent on colour and transitions.
Write captions and subtitles that work
Make a video watchable with the sound off.
Help me caption this video. TRANSCRIPT OR SCRIPT: """ [PASTE] """ VIDEO LENGTH: [DURATION] PLATFORM: [WHERE IT IS POSTED] STYLE: [burned-in captions / subtitle file / both] AUDIENCE: [WHO, AND ANY ACCESSIBILITY REQUIREMENTS] Produce: 1. THE FORMAT DECISION - burned-in captions cannot be turned off and are what most social viewers need; a subtitle file is selectable, translatable and better for accessibility. For my platform, which and why. Often both. 2. THE CAPTIONED TEXT - broken into caption-length segments: - Two lines maximum, roughly 40 characters per line - Break at natural phrase boundaries, never mid-phrase or after an article - Each caption on screen long enough to read: roughly 20 characters per second is a comfortable reading rate - Do not split a short sentence across two captions 3. THE CLEAN-UP DECISION - how much to tidy speech. Remove filler and false starts for readability, but do not rewrite what was said into something better. Say where the line is, and note that verbatim matters for accessibility and for anything quoted. 4. THE NON-SPEECH INFORMATION - for accessibility, captions must convey meaningful sound: [door slams], [laughter], music that matters, and who is speaking when it is not visible. This is what separates captions from subtitles and it is routinely omitted. 5. THE PLACEMENT - where on screen, and what must not be covered. Platform interfaces cover the bottom portion; say how much to keep clear for my platform. 6. READABILITY - font size for small screens, contrast against varying backgrounds, and a background or outline so text stays legible over bright footage. 7. THE SUBTITLE FILE - the properly formatted SRT or VTT with timings, if I asked for one. 8. THE AUTO-CAPTION WARNING - platform auto-captions get names, technical terms, numbers and accents wrong, and errors in captions are more noticeable than errors in speech. Say what to check specifically in my transcript. 9. THE OPENING CAPTION - the first caption appears before most viewers have decided to watch. It should carry the hook.
What you get: Properly segmented captions at a readable rate, non-speech sound included, placement guidance and a formatted subtitle file.
Tip: Point 4 is what makes captions accessible rather than merely present. A caption track with no sound information leaves deaf viewers missing half the scene.
Decide whether to use AI video at all
Choose between generated footage, stock, and shooting it.
Help me decide how to produce this video content. WHAT I NEED: [DESCRIBE THE FOOTAGE OR VIDEO] WHAT IT IS FOR: [PURPOSE AND AUDIENCE] BUDGET: [AMOUNT OR 'none'] TIME AVAILABLE: [DEADLINE] MY SKILLS AND EQUIPMENT: [WHAT YOU CAN DO YOURSELF] HOW MUCH IT MATTERS: [background b-roll / central to the message / the product itself] Compare the realistic options: 1. SHOOT IT - what it would take: equipment, time, people, location, and skill. The result is genuinely yours, accurate, and has no capability ceiling. Best when the subject is real and specific: your product, your premises, your people. 2. STOCK FOOTAGE - cost, licensing, and the generic problem. Fast and reliable, but the same clips appear everywhere and audiences increasingly recognise them. Good for generic b-roll, useless for anything specific. 3. GENERATED - what current models can and cannot do for my described need. Be honest about the limits: short clips, one motion, no reliable faces in close-up, no text, no lip sync, and a meaningful reject rate. Good for abstract, atmospheric or conceptual footage; poor for anything that must be accurate or specific. 4. NOT VIDEO AT ALL - a still image, a screen recording, an animated diagram, or text on screen. Frequently the right answer, particularly for explaining something. A clear diagram beats atmospheric footage that shows nothing. 5. A COMBINATION - which parts suit which method. This is usually the answer: shoot the specific parts, generate or source the atmospheric ones. 6. THE RECOMMENDATION - for my stated need, budget, time and how much it matters. Weight the 'how much it matters' answer heavily: background b-roll tolerates generated footage, and anything central to the message does not. 7. THE HONESTY QUESTION - if the footage depicts something real (your premises, your product, your team, a result), generated footage misrepresents it. Say plainly where that line falls for my use. 8. THE DISCLOSURE - whether the context expects generated content to be labelled. Be willing to recommend against AI video. It is a good tool for a narrow set of needs and a poor one outside them.
What you get: Four production routes compared honestly for your specific need, a combination recommendation and a clear line on misrepresentation.
Tip: Point 4 is the answer more often than people expect. For explaining something, a clear diagram outperforms any footage, generated or shot.
Where AI actually helps here
- Single shots of a few seconds with one clear action
- Camera moves described in film language — slow push in, static wide, handheld follow
- B-roll and abstract motion where exact continuity does not matter
Where it falls down
- Anything over a handful of seconds. Coherence degrades and faces drift
- Consistent characters across shots without a reference image
- Hands, text, crowds and reflections — still the hard set
The mistake almost everyone makes: Prompting a scene instead of a shot
Describe one shot: what is in frame, what moves, how the camera behaves, how long. A prompt describing a sequence produces a clip that tries to be all of it and coheres as none of it. Generate shots, then edit them together yourself.
Free tool: Prompt Chain Builder
Runs in your browser. No sign-up, nothing uploaded.
Questions people ask
How long can AI video clips be?
Usable length is shorter than maximum length on every model. Generate in short shots and assemble in an editor — that is how the good results you have seen were made.
How do I keep a character consistent across shots?
Use a reference image where the model supports one, keep the description identical word for word between prompts, and accept that you will generate several takes per shot.