YouTube, explainers, ads and voiceover — written for the ear, not the page. Below are 7 copy-ready prompts. Fill in the [BRACKETS], copy, and paste into ChatGPT, Claude, Gemini or any capable assistant.
Scripts fail on the page-versus-ear problem: sentences that read fine are unspeakable. Models write for the page unless told otherwise, and the tell is a subordinate clause nobody can say out loud.
The 7 prompts
Write a video script that holds attention
Script a video where every section earns the next.
Write a video script. TOPIC: [WHAT IT IS ABOUT] LENGTH: [MINUTES] PLATFORM: [YouTube / short-form vertical / course / internal training / ad] VIEWER: [WHO AND WHY THEY CLICKED] WHAT THEY SHOULD DO OR KNOW AFTERWARDS: [OUTCOME] WHAT I KNOW THAT OTHERS DO NOT: [YOUR ANGLE OR EXPERIENCE] PRESENTATION: [talking head / voiceover over screen / animation / no face] Structure: **0-5 seconds**: the hook. State the specific promise or show the result. No introduction, no 'hey guys', no channel branding, no 'in this video I'm going to'. **5-20 seconds**: why this viewer should stay. What they get and roughly when. **Body**: sections, each with a timestamp, the point it makes, and how it opens. Every section must open with something that earns attention - a question, a claim, a result - not a transition phrase. **Close**: the actual conclusion, then one call to action. Not three. For each line of script, mark whether it is spoken, on-screen text, or a visual note. Also give me: - THE RETENTION RISKS - the three moments most likely to lose the viewer, and what to do at each. Usually: right after the hook, at any point where you explain background, and before the payoff. - WHAT TO CUT if this runs long, in order - THE TITLE AND THUMBNAIL CONCEPT that matches this hook Rules: - Write to be spoken. Short sentences. No subordinate clauses that only work on the page. - Do not front-load context. Get to the substance and backfill. - No 'before we get started', no 'don't forget to like and subscribe' mid-video. - If this content does not need [LENGTH] minutes, say so and give the honest length.
What you get: A timestamped, speakable script with visual notes, retention risk points and an honest length check.
Tip: 'Do not front-load context' is the most-violated rule in video. The background explanation belongs after the viewer is invested, not before.
Write a short-form vertical video script
Script for the formats where the first second decides everything.
Write a short-form vertical video script. PLATFORM: [TikTok / Reels / Shorts] LENGTH: [SECONDS - 15, 30, 45 or 60] TOPIC: [WHAT IT IS ABOUT] VIEWER: [WHO] THE ONE THING they should take away: [TAKEAWAY] MY FORMAT: [talking to camera / voiceover with b-roll / text on screen / screen recording] Structure it second by second: **Second 0-1**: the visual and the first words. These must work with sound off - most viewers see before they hear. Give the on-screen text for this moment. **Second 1-3**: the promise or the tension. Why keep watching. **Middle**: the content, in beats. Something must change every 3-5 seconds: a new point, a cut, a visual change, a question. **Last 3 seconds**: the payoff. Deliver what the hook promised. Optionally a loop back to the start. Output as a table: Time | Spoken | On-screen text | Visual Rules: - One idea only. Short-form fails by trying to cover two. - The hook must be the most interesting thing, not a build-up to it. - No introduction of yourself. No 'part 1 of a series'. - Assume sound off for the first two seconds - the on-screen text carries the hook. - Write spoken lines to be said in one breath. - No 'follow for more' before the payoff. Then give: - THREE ALTERNATIVE HOOKS for the same content - THE CAPTION, written to add something rather than repeat the video - THE HONEST CHECK: is this idea too big for [LENGTH] seconds? If so, say what to cut or split.
What you get: A second-by-second table of speech, on-screen text and visuals, with three alternative hooks.
Tip: Writing the first two seconds to work with sound off is the highest-leverage rule in short-form. Most viewers decide before audio registers.
Write a podcast interview plan
Prepare questions that produce a good conversation rather than a press release.
Plan a podcast interview. GUEST: [WHO THEY ARE AND WHAT THEY ARE KNOWN FOR] WHAT I KNOW ABOUT THEM: [BACKGROUND, THEIR WORK, ANYTHING YOU HAVE READ] MY AUDIENCE: [WHO LISTENS AND WHAT THEY WANT] EPISODE LENGTH: [MINUTES] WHAT I WANT THE EPISODE TO BE ABOUT: [ANGLE] Produce: 1. THE ANGLE - the specific thing this conversation is about. Not 'their career'. The narrower the better. 2. WHAT THEY HAVE ALREADY SAID EVERYWHERE - the questions they have answered a hundred times, which will produce their rehearsed answer. List them so I avoid them. 3. THE QUESTION ARC - 8-12 questions in an order that builds. Start with something concrete and specific rather than broad, because opening with 'tell us about your journey' produces a rehearsed monologue. For each question: the question itself, why it will produce something interesting, and a follow-up probe if the answer is thin. 4. THE THREE BEST QUESTIONS - the ones most likely to produce something they have not said before. Usually these are about: a specific decision and its alternative, something that did not work, a disagreement with their own field, or a concrete detail of how they actually work. 5. THE UNCOMFORTABLE ONE - the question worth asking that they might not enjoy, phrased fairly. Whether to ask it is my call; give me the best version of it. 6. WHERE TO LET IT WANDER - the topics where a tangent would be better than the next question. 7. THE OPENING AND CLOSING - the first question, and how to end without 'where can people find you'. Rules: - No question answerable with yes or no. - No two-part questions; they always answer the second and forget the first. - No 'what advice would you give'. It produces platitudes every time. - Prefer 'tell me about the time when' over 'what do you think about'.
What you get: A built question arc with probes, the three highest-yield questions, and a list of questions to avoid because they are over-asked.
Tip: Point 2 is the preparation that separates good interviews from bad ones. Asking what everyone asks gets you what everyone got.
Write a demo or product walkthrough script
Show software without narrating every click.
Write a product demo script. PRODUCT: [WHAT IT IS] AUDIENCE: [WHO IS WATCHING AND WHAT THEY CARE ABOUT] CONTEXT: [sales demo / onboarding / conference talk / recorded walkthrough] LENGTH: [MINUTES] THE PROBLEM IT SOLVES: [THE PAIN] WHAT I MUST SHOW: [REQUIRED FEATURES] WHAT USUALLY GOES WRONG IN MY DEMOS: [IF ANYTHING] Structure: 1. THE BEFORE - 30 seconds on the painful current state, in the audience's terms. Concrete, not abstract. This is what makes the rest land. 2. THE PAYOFF FIRST - show the end result before showing how to get there. Audiences need to know where you are going. 3. THE WALKTHROUGH - the path from problem to result. For each step: what you click, what you say, and what the audience should notice. Say the benefit, not the feature name. 4. THE MOMENT - the single point where the audience should be impressed. Identify it and say how to slow down and let it land. 5. THE CLOSE - what they do next. Rules: - Never narrate the interface. Nobody needs 'and here I'll click the settings icon'. Say what you are achieving. - Use realistic data. Demo data called 'Test Customer 1' undermines everything. - Do not show every feature. Show the path that solves the stated problem. - Have an answer ready for 'can it do X' that is not a detour into another feature. Also give me: - THE CUT LIST if it runs long - THE THREE QUESTIONS the audience will interrupt with, and short answers that do not derail the demo - THE FAILURE PLAN - what to say if something does not load or breaks live - WHAT NOT TO SHOW - features that will raise questions you do not want to answer yet
What you get: A problem-first demo script with a designated impressive moment, interruption answers and a live-failure plan.
Tip: Showing the payoff before the walkthrough is the structural fix. Audiences tolerate a lot of clicking if they already know what it builds to.
Write a conference talk from an idea
Turn something you know into a talk with a shape.
Help me build a conference talk. THE IDEA: [WHAT YOU WANT TO TALK ABOUT] WHY YOU: [YOUR EXPERIENCE WITH THIS] AUDIENCE: [WHO ATTENDS, THEIR LEVEL] LENGTH: [MINUTES INCLUDING OR EXCLUDING Q&A] WHAT THEY SHOULD DO DIFFERENTLY AFTERWARDS: [OUTCOME] CONFERENCE TYPE: [TECHNICAL / INDUSTRY / ACADEMIC / INTERNAL] Produce: 1. THE ONE SENTENCE - what this talk argues. If it is 'here is an overview of X', that is not a talk, it is a documentation page read aloud. Push me toward a claim. 2. THE TALK'S SHAPE - one of: problem→journey→solution, a thesis defended, a story with a lesson, or a myth dismantled. Recommend one for this material and say why. 3. THE STRUCTURE - sections with minute budgets. Budget 80% of the stated time; talks always run long. 4. THE OPENING 90 SECONDS - written out. No agenda slide, no 'a bit about me' beyond one sentence, no apology. Start with the problem or a specific moment. 5. THE SPINE - the 5-7 points the audience should retain, in order. If they remember nothing else, these. 6. THE SPECIFIC BITS - talks are remembered for concrete details: a number, a failure, a real example, a screenshot of something real. From what I told you, which specifics should anchor each section? Mark [NEED] where I have not given you one. 7. THE HONEST PART - the talk needs a moment where you say something true and slightly uncomfortable: what did not work, what you got wrong, what the trade-off actually is. Audiences trust talks that admit cost. Identify where it goes. 8. SLIDE GUIDANCE - roughly how many, and what belongs on them versus what you say. No slide should be readable and listened to at once. 9. THE Q&A PREP - the four questions you will be asked, including the hostile one. 10. THE ABSTRACT - the submission version, under 200 words.
What you get: A talk built around a claim with a chosen shape, a written opening, a retained spine, and the uncomfortable honest moment placed.
Tip: Point 7 is what makes a talk memorable. The speaker who says 'this cost us six months and here is what I would do differently' gets remembered; the one with only successes does not.
Write a voiceover script for existing footage
Narrate video you have already shot.
Write a voiceover script for this footage. FOOTAGE DESCRIPTION: [WHAT IS ON SCREEN, WITH TIMINGS IF YOU HAVE THEM] TOTAL LENGTH: [DURATION] PURPOSE: [explainer / brand film / tutorial / documentary / ad] AUDIENCE: [WHO] TONE: [DESCRIPTION] MUSIC OR AMBIENT SOUND: [IF ANY] Produce a table: Timecode | Visual | Voiceover | Words | Notes Pacing rules: - Conversational narration runs roughly 140-160 words per minute. Slower for emotional or technical content, faster for energetic. Calculate the word budget per segment and stay inside it. - Leave silence. Footage that speaks for itself should not be narrated over. Mark these segments [NO VO] deliberately. - Never describe what is visible. If the viewer can see it, say something else: why it matters, what happened before, what it leads to. - Land key lines on visual cuts, not across them. Writing rules: - Write for the ear: short sentences, active voice, no clauses that require rereading. - No tongue-twisters or awkward consonant clusters. - Numbers written as they should be spoken ('twenty-twenty-six', not '2026'). - Mark [PAUSE], [EMPHASIS] and pronunciation for anything unusual. Then give me: - WHERE THE SCRIPT IS FIGHTING THE FOOTAGE - moments where what I want to say does not match what is on screen, and whether to change the words or the edit - TOTAL WORD COUNT AND RUNTIME estimate - THE THREE LINES that carry the piece
What you get: A timecoded voiceover table within a calculated word budget, with deliberate silence and script/footage conflicts flagged.
Tip: 'Never describe what is visible' is the rule that separates professional narration from a slideshow with captions read aloud.
Write dialogue that sounds like people talking
Fix dialogue that reads as written rather than spoken.
Write or improve this dialogue. SCENE/SITUATION: [WHAT IS HAPPENING] CHARACTERS: [WHO THEY ARE, THEIR RELATIONSHIP, WHAT EACH WANTS IN THIS SCENE] WHAT MUST HAPPEN IN THE SCENE: [PLOT OR INFORMATION REQUIREMENT] MEDIUM: [screenplay / novel / audio drama / game / training video] EXISTING DRAFT: [PASTE, or 'write it from scratch'] Write the dialogue, applying these principles: 1. CHARACTERS WANT DIFFERENT THINGS - dialogue without conflicting wants is exposition with quote marks. If both characters want the same thing here, tell me the scene has a structural problem. 2. PEOPLE DO NOT ANSWER THE QUESTION ASKED - they deflect, change subject, answer a different question, or answer the emotion rather than the words. 3. DISTINCT VOICES - each character should be identifiable with the names removed. Vary: sentence length, vocabulary, how directly they speak, what they avoid, verbal habits (sparingly). 4. SUBTEXT - the most important thing in the scene should mostly not be said aloud. 5. INTERRUPTION AND OVERLAP - real conversation is not turn-taking. 6. CUT THE GREETING - start the scene as late as possible, after the small talk. Rules: - No character explains something both already know for the audience's benefit - No speech longer than four lines unless the character is deliberately holding the floor and that is the point - Sparing dialect spelling. Rhythm and word choice convey accent better than phonetic spelling. - Cut dialogue tags where the speaker is obvious Then tell me: - THE VOICE TEST - cover the names: can you tell who is speaking? Where you cannot, say so. - WHAT IS BEING SAID ALOUD THAT SHOULD BE SUBTEXT - WHERE THE SCENE COULD START LATER
What you get: Dialogue with conflicting wants and subtext, plus a voice-distinctness test and a note on where the scene should start.
Tip: The name-covering voice test is the fastest dialogue diagnostic. If every character sounds the same, you have one character talking to themselves.
Where AI actually helps here
- Hooks — ask for ten and two will be usable, which is a good rate
- Restructuring a written piece into spoken order
- Pacing notes, B-roll suggestions and beat structure
Where it falls down
- Sounding like a person. Default script output has a YouTube-narrator cadence everyone recognises
- Knowing your format’s conventions unless you describe them
- Timing. Its word counts for a duration are unreliable — read it aloud with a timer
The mistake almost everyone makes: Not making it write for the mouth
Add: write this to be spoken. Short sentences. No subordinate clauses. If a line cannot be said in one breath, split it. Then read the draft aloud — the places you stumble are the places the model was writing prose.
Free tool: Prompt Builder
Runs in your browser. No sign-up, nothing uploaded.
Questions people ask
How long should a script be for a 10-minute video?
Roughly 1,300–1,600 words at a normal speaking pace, but it varies enormously with delivery and B-roll. Do not trust a model’s estimate — read a minute aloud and time yourself, then multiply.
Can AI write a hook that works?
It writes hook-shaped sentences reliably and genuinely surprising ones rarely. Generate ten, discard eight. The point is the volume, not the quality of any single one.