Skip to content
PromptifyLab

Image, Video & Audio

Audio, Music & Voice prompts

7 prompts Free · no sign-up Works in ChatGPT, Claude & Gemini Search & filter these

Voiceover direction, music briefs and sound design. Below are 7 copy-ready prompts. Fill in the [BRACKETS], copy, and paste into ChatGPT, Claude, Gemini or any capable assistant.

Audio tools take direction the way a session musician or a voice actor does, which means the useful vocabulary is a director’s, not a prompt engineer’s.

The 7 prompts

Intermediate 5 blanks to fill

Write a script for text-to-speech

Write words that sound right when a synthetic voice reads them.

Prompt
Write or adapt this for text-to-speech.

THE CONTENT:
"""
[PASTE, or describe what you need written]
"""

TTS SERVICE: [WHICH ONE]
PURPOSE: [video voiceover / audio article / phone system / accessibility / audiobook]
LENGTH: [TARGET DURATION OR WORD COUNT]
VOICE: [DESCRIBE THE VOICE OR ITS SETTINGS]

Produce:

1. THE ADAPTED SCRIPT - rewritten for the ear:
   - Short sentences. Long subordinate clauses that work on the page become incomprehensible spoken.
   - One idea per sentence
   - No parenthetical asides; synthetic voices handle them badly and listeners cannot see the brackets
   - No bullet lists read aloud as lists; convert to prose with clear sequencing words

2. THE PRONUNCIATION MARKUP - every item that will be read wrongly:
   - Numbers: write them as they should be spoken. '2026' is 'twenty twenty-six' or 'two thousand and twenty-six' depending on context, and the engine will pick one, often wrongly.
   - Abbreviations and acronyms: say which are spelled out and which are read as words
   - Currency, percentages, dates, times, and measurements
   - Technical terms, names, and any word with an unusual pronunciation
   - Homographs where stress or pronunciation depends on meaning
   Provide the phonetic spelling or the SSML tag for each, in my service's syntax if it supports it.

3. THE PACING MARKUP - where pauses go. Synthetic speech runs at a constant pace and sounds relentless without deliberate breaks. Mark: the pause after a key statement, the longer pause between sections, and the shorter one before a list item.

4. THE EMPHASIS - the words that carry the meaning of each sentence, marked in my service's syntax. Without emphasis, every sentence has the same shape.

5. THE THINGS THAT WILL STILL SOUND WRONG - what current TTS handles poorly regardless of markup: questions with rising intonation, sarcasm, emotional emphasis, and sentences whose meaning depends on stress placement. Rewrite these to avoid the problem rather than marking them up.

6. THE READ-ALOUD TEST - read the script aloud yourself. Anything you stumble over, the engine will get wrong.

7. THE TIMING - estimated duration at typical speaking rates, so I can check against my target.

8. THE DISCLOSURE - whether the context requires saying the voice is synthetic.

9. WHAT NOT TO USE TTS FOR - anything needing genuine emotional range, a personal message, or where a listener would feel deceived. Say if my use is one.

What you get: A script rewritten for the ear with pronunciation and pacing markup, emphasis marked and unfixable problems rewritten around.

Tip: Point 2's number handling is where TTS most often embarrasses itself. '1990s' read as 'one thousand nine hundred and ninety s' is a two-minute fix that nobody makes.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Intermediate 6 blanks to fill

Plan and record a podcast episode

Produce an episode that is worth listening to.

Prompt
Help me plan a podcast episode.

SHOW: [WHAT IT IS ABOUT, WHO LISTENS]
EPISODE TOPIC: [SUBJECT]
FORMAT: [solo / interview / co-hosted / narrative]
TARGET LENGTH: [MINUTES]
MY MATERIAL: [NOTES, RESEARCH, OR GUEST DETAILS]
RECORDING SETUP: [EQUIPMENT AND SPACE]

Produce:

1. THE EPISODE'S POINT - what a listener should take away. An episode about a topic meanders; an episode making a point does not.

2. THE STRUCTURE - segments with rough timings. For each: what it covers and why it earns its minutes.

3. THE OPENING - the first 30 seconds. Lead with the most interesting thing in the episode, not with the show's name and a greeting. Listeners decide in the first minute, and most shows spend that minute on housekeeping.

4. THE ARC - what changes across the episode. Even a conversational show needs somewhere to end up. Name the arc.

5. IF INTERVIEW - the question plan: the opening question that is specific rather than broad, the arc of questions, the probes, and the questions to avoid because the guest has answered them everywhere.

6. IF SOLO - the structure that prevents rambling: the points in order, the examples for each, and the transitions. Solo episodes fail by having no structure and by having too much.

7. THE RECORDING SETUP ADVICE - for my equipment and space. In order of impact: get the microphone close to your mouth, record in a soft-furnished room rather than a hard one, record each participant on a separate track, and record a backup. Room acoustics matter more than microphone price.

8. THE COMMON RECORDING MISTAKES - levels set too hot and clipping, no backup recording, recording over a video call rather than locally at each end, background noise nobody noticed until the edit, and starting to record after the best part of the conversation.

9. THE EDIT PLAN - what to cut: the warm-up, the false starts, the tangent that went nowhere, the part where the point was already made. Roughly what proportion survives.

10. THE SHOW NOTES - what they need: the episode's point, timestamps for sections, anything referenced with links, and guest details. Written so someone deciding whether to listen can tell.

11. THE SUSTAINABILITY CHECK - given how long this planning takes, whether my publishing schedule is realistic.

What you get: An episode built around a point with a structured arc, a strong opening, recording setup priorities and a realistic edit plan.

Tip: Point 7's ordering is worth following. A cheap microphone close to your mouth in a carpeted room beats an expensive one across a bare office.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Intermediate 5 blanks to fill

Fix audio that sounds bad

Diagnose and repair a recording.

Prompt
Help me fix this audio.

THE PROBLEM: [DESCRIBE WHAT IS WRONG - hiss, echo, muffled, too quiet, distorted, inconsistent levels, background noise, plosives]
HOW IT WAS RECORDED: [EQUIPMENT, SPACE, SETUP]
WHAT IT IS FOR: [podcast / video / voiceover / archive]
SOFTWARE: [WHAT YOU HAVE]
CAN I RE-RECORD: [yes / no]

Produce:

1. THE RECOVERABILITY ASSESSMENT first, honestly:
   - FIXABLE: hiss and steady background noise, inconsistent levels, too quiet, mild plosives, minor mouth noise, slight EQ problems
   - PARTIALLY FIXABLE: echo and room reverb (reducible, never removable), muffled recording, sibilance
   - NOT FIXABLE: clipping and digital distortion (the information is gone), heavy reverb, severe background noise that overlaps the voice frequencies
   Tell me which category mine falls into. If it is not fixable and I said I can re-record, say so immediately rather than describing repairs that will not work.

2. THE CAUSE - what in my described setup produced this, so it does not recur. Almost every audio problem is a recording problem, and the fix next time is usually simpler than the repair now.

3. THE REPAIR ORDER - sequence matters:
   - Remove noise first, before any processing that would amplify it
   - Then corrective EQ
   - Then de-essing and plosive repair
   - Then compression for consistency
   - Then loudness normalisation for the destination, last

4. THE SPECIFIC STEPS in my software, with settings and what each does.

5. THE NOISE REDUCTION WARNING - over-applied noise reduction produces a watery, artificial sound that is worse than the original noise. Use the lightest setting that helps, and say what the artefacts sound like so I can hear when I have gone too far.

6. THE LOUDNESS TARGET - for my destination, the appropriate level. Platforms normalise audio, so mastering too loud achieves nothing and costs dynamic range.

7. WHAT WILL STILL BE AUDIBLE after the repair, realistically.

8. THE RE-RECORD DECISION - if I can re-record, an honest comparison: the time spent repairing versus re-recording properly. Repairing bad audio often takes longer than recording it again and produces a worse result.

9. NEXT TIME - the three changes to the recording setup that would prevent this. In most cases: closer microphone, softer room, and check levels before recording rather than after.

What you get: An honest recoverability verdict, the cause, an ordered repair sequence with settings, and a re-record versus repair comparison.

Tip: Point 1's clipping entry is the one to internalise. Distorted audio cannot be repaired because the waveform information was never captured.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Intermediate 5 blanks to fill

Use AI voice cloning responsibly

Understand what is acceptable and what is not.

Prompt
Help me think through using AI voice generation or cloning.

WHAT I WANT TO DO: [DESCRIBE THE USE]
WHOSE VOICE: [my own / a colleague who agreed / a hired voice actor / a public figure / a synthetic voice]
PURPOSE: [audiobook / video / accessibility / customer service / marketing / personal project]
WHO WILL HEAR IT: [AUDIENCE]
JURISDICTION: [WHERE YOU OPERATE]

Produce:

1. THE CONSENT QUESTION first, and it is not negotiable. Cloning a voice requires that person's informed, specific, documented permission. Informed means they understand what it will be used for, for how long, and whether it can be used for things they have not individually approved. If my described use involves someone who has not given that, say so plainly and stop there. If it involves a public figure without permission, the answer is no.

2. WHAT CONSENT SHOULD COVER - if permission is being obtained: the specific uses permitted, the duration, whether the model can be reused for new content, whether it can be transferred, how to revoke it, and what happens to the voice model afterwards. A verbal yes is not sufficient for anything commercial.

3. THE DISCLOSURE QUESTION - when listeners should be told the voice is synthetic. The practical test: would a listener feel deceived if they found out? If yes, disclose. Narration and accessibility uses generally warrant a note; anything presented as a person speaking directly to the listener certainly does.

4. THE USES THAT ARE NOT ACCEPTABLE regardless of technical capability: impersonating someone to make it appear they said something, any use intended to deceive about who is speaking, voices of deceased people without the family's agreement, and anything that would be used to authenticate identity. Say plainly if my described use is in this territory.

5. THE LEGAL LANDSCAPE - personality and voice rights vary considerably by jurisdiction and are changing. State that this needs checking locally for anything commercial, and that I should not rely on this as advice.

6. THE PRACTICAL QUALITY NOTES - if the use is legitimate: what makes a cloned voice sound convincing, how much source audio is needed, what it will still get wrong (emotional range, emphasis, unusual words), and where a real recording is still better.

7. THE SECURITY IMPLICATION - a voice model of yourself is a credential-adjacent asset. Voice authentication systems exist, and family members can be deceived by a familiar voice. Say what to consider about storing and controlling access to a voice model.

8. THE ALTERNATIVE - a good synthetic voice that belongs to nobody is available for most uses and avoids all of this. Say whether my use actually requires a specific person's voice.

If the described use is not acceptable, say so directly and offer what is.

What you get: A consent gate, what consent must cover, disclosure guidance, unacceptable uses named plainly and a synthetic-voice alternative.

Tip: Point 8 resolves most of these questions. Very few uses genuinely need a specific person's voice rather than a good voice, and choosing a synthetic one avoids every complication.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Beginner 6 blanks to fill

Choose and use music in your content

Add music without a licensing problem.

Prompt
Help me with music for this project.

THE PROJECT: [WHAT IT IS]
WHERE IT WILL BE PUBLISHED: [PLATFORMS]
COMMERCIAL OR NOT: [monetised / brand content / personal / educational]
WHAT THE MUSIC SHOULD DO: [set a mood / carry energy / fill silence / underscore narration]
LENGTH AND STRUCTURE: [WHERE MUSIC IS NEEDED]
BUDGET: [AMOUNT OR 'none']

Produce:

1. THE LICENSING REALITY first, because getting this wrong is expensive:
   - COMMERCIAL MUSIC: needs a licence, usually two (the recording and the composition). Not available for most small projects.
   - ROYALTY-FREE LIBRARIES: a one-off or subscription fee for a broad licence. The practical option for most work. Read what the licence actually permits: some exclude broadcast, some exclude client work, some require attribution.
   - CREATIVE COMMONS: free with conditions. Attribution requirements are binding, and some licences prohibit commercial use or derivative works. Say how to check and comply.
   - PUBLIC DOMAIN: the composition may be free while a specific recording of it is not. This trips people up constantly.
   - AI-GENERATED: check the service's terms for commercial rights, and be aware the legal position on ownership is unsettled.
   For my project and budget, which route.

2. THE PLATFORM CLAIM PROBLEM - automated content identification will flag music, including music you are licensed to use and sometimes music that is public domain. What this means practically: keep your licence documentation, know how to dispute a claim, and understand that a claim can affect monetisation while it is disputed.

3. WHAT THE MUSIC SHOULD DO - given my stated purpose. Music under narration must be unobtrusive: no prominent melody competing with speech, no lyrics, and limited dynamic range. Music carrying a montage can do more.

4. THE SELECTION BRIEF - what to search for: tempo, instrumentation, energy, and mood, in the terms music libraries use. Also what to avoid for this project.

5. THE MIXING - the practical part most people get wrong:
   - Music under speech should sit far lower than feels right when you are mixing it
   - Duck the music when speech starts, and bring it back in the gaps
   - Fade in and out rather than cutting
   - Check the mix on phone speakers and on headphones; they are very different

6. THE STRUCTURE - where music starts and stops. Continuous music through a whole piece is fatiguing. Silence is a tool.

7. THE ALTERNATIVE - many pieces are better with no music at all. Say whether mine is one. Music added to fill silence usually signals that the content is not holding attention, and music will not fix that.

8. THE DOCUMENTATION - what to record for every track used: the source, the licence, the date, and the attribution required. Keep it with the project.

What you get: A licensing route matched to your project, the platform claims problem explained, a selection brief and practical mixing guidance.

Tip: Point 5's first line is the most common mixing error. Background music always feels too quiet to the person mixing it and is too loud for everyone listening.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Beginner 6 blanks to fill

Write and record a voiceover that sounds natural

Deliver a script without sounding like you are reading.

Prompt
Help me write and record a voiceover.

THE SCRIPT OR CONTENT: [PASTE, or describe what it covers]
PURPOSE: [explainer / advert / tutorial / narration / e-learning]
LENGTH: [DURATION]
TONE: [DESCRIPTION]
WHO IS RECORDING: [me / someone else / a professional]
EQUIPMENT AND SPACE: [WHAT YOU HAVE]

Produce:

1. THE SCRIPT, WRITTEN TO BE SPOKEN:
   - Sentences short enough to say in one breath
   - Contractions, because people use them and reading without them sounds stiff
   - No subordinate clauses stacking up
   - Words that are easy to say. Flag any tongue-twisters and replace them.
   - Written in the rhythm of speech, not of prose

2. THE MARKED-UP VERSION - for the person reading it:
   - Line breaks at breathing points, not at margin width
   - Emphasis marked on the word carrying each sentence's meaning
   - Pauses marked, with their length
   - Pronunciation noted for anything unusual
   Formatted so it can be read from a screen without scrolling mid-sentence.

3. THE DELIVERY NOTES - how to sound natural:
   - Read to one person, not to an audience. Picture them.
   - Smile where the content is warm; it is audible.
   - Slow down. Everyone reads faster than they think, and nervousness accelerates delivery.
   - Stand up rather than sitting; it changes the voice.
   - Record a few sentences of conversation first and match that energy.

4. THE RECORDING SETUP - for my equipment and space:
   - Microphone position: close, slightly off-axis to avoid plosives
   - Room treatment with what I have: soft furnishings, a wardrobe, blankets, a corner rather than the centre of the room
   - Levels set so the loudest parts have headroom
   - Record a few seconds of silence for noise reduction reference

5. THE TAKE STRATEGY - record in sections rather than one long take, keep going through mistakes and repeat the line rather than stopping, and leave a pause before each retake so the edit point is clean.

6. THE COMMON PROBLEMS - reading pace accelerating through the take, energy dropping in the second half, mouth noise, breathing that is too close to the microphone, and the final word of each sentence dropping in volume.

7. THE EDIT - removing breaths but not all of them, since a voice with no breathing sounds inhuman. Levelling, light compression, and a gentle high-pass filter.

8. THE TIMING CHECK - word count against target duration at a natural pace, so the script is not being rushed to fit.

What you get: A speakable script plus a marked-up performance version, delivery notes, recording setup and a take strategy.

Tip: Point 7's note about breaths is worth knowing. Removing every breath produces an unsettling, inhuman delivery that listeners notice without knowing why.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026
Intermediate 5 blanks to fill

Turn written content into audio

Adapt an article or document for listening.

Prompt
Help me turn this written content into audio.

THE CONTENT:
"""
[PASTE]
"""

WHAT IT IS: [article / report / documentation / newsletter / book chapter]
WHO WILL LISTEN AND WHEN: [commuting / working / accessibility / studying]
TARGET LENGTH: [DURATION]
VOICE: [me reading / TTS / a narrator]

Produce:

1. WHAT DOES NOT TRANSLATE - the elements that fail in audio and what to do with each:
   - Headings: invisible to a listener. Convert to spoken signposts.
   - Bullet lists: hard to follow spoken. Convert to prose with sequencing, or keep short and announce the count first.
   - Tables and data: almost impossible. Summarise the finding instead of reading the numbers.
   - Links and references: cannot be followed. Mention the source in words or defer to show notes.
   - Images and charts: describe the point they make, not the chart.
   - Code, formulas and anything requiring exact characters: generally should not be read aloud.
   - Parentheses and footnotes: restructure into the main sentence or cut.

2. THE ADAPTED SCRIPT - rewritten for listening:
   - Signposting, because the listener cannot see where they are: what is coming, and what was just covered
   - More repetition than written text needs. A listener cannot re-read a sentence.
   - The key point stated before the explanation, not after
   - Concrete language; abstract sequences are hard to hold in the ear

3. THE STRUCTURE - audio needs clearer sections than text. Give the spoken transitions between them.

4. THE OPENING - a listener needs to know within twenty seconds what this is and whether it is for them.

5. THE LENGTH ADJUSTMENT - spoken content takes longer than reading and holds less. Say what the written word count implies in duration, and what to cut if it exceeds my target. Audio adaptations usually need to be shorter in substance, not just in words.

6. THE ATTENTION CURVE - listeners are usually doing something else. Where to place the important material, and how often to re-orient them.

7. THE SUPPORTING TEXT - what should exist alongside the audio: a transcript, links, and any table or figure that could not be spoken.

8. THE HONEST CHECK - is this content suited to audio at all? Reference material, anything with tables, and anything the reader needs to scan rather than follow linearly is better left as text. Say so if that is the case here.

What you get: An element-by-element translation plan, an adapted script with signposting and repetition, and an honest check on suitability.

Tip: Point 2's repetition note is the main difference between good and bad audio adaptation. Listeners cannot glance back, so a well-placed restatement is not padding.

Open in Written for Claude, ChatGPT · Reviewed September 18, 2026

Where AI actually helps here

  • Voiceover direction — pace, energy, where to breathe, which word carries the line
  • Music briefs by function: what it should do under a scene, not what genre it is
  • Sound design lists for a scene

Where it falls down

  • Cloning a voice without permission. Legally exposed and getting more so
  • Emotional range. Generated delivery is even, and evenness reads as flat
  • Long-form consistency — tone drifts across a long read

The mistake almost everyone makes: Not directing the read

Give it the direction you would give a person: ‘conversational, slightly faster than natural, small pause before the number, emphasis on “only”‘. Text with no direction gets a neutral read, and neutral is the thing that sounds synthetic.

Free tool: Prompt Builder

Runs in your browser. No sign-up, nothing uploaded.

Open the Prompt Builder →

Questions people ask


Is it legal to clone a voice with AI?

Not without the person’s permission, and several jurisdictions have added specific protections for voice and likeness. Synthetic voices that are not modelled on a real person avoid the question entirely.


How do I make AI voiceover sound natural?

Direct it, punctuate for breath, break long sentences, and spell tricky words phonetically. Most of what sounds robotic is the script, not the voice.