Skip to content
PromptifyLab

Prompt Engineering

Troubleshooting prompts

7 prompts Free · no sign-up Works in ChatGPT, Claude & Gemini Search & filter these

Diagnosing the specific ways prompts fail, and what to change. Below are 7 copy-ready prompts. Fill in the [BRACKETS], copy, and paste into ChatGPT, Claude, Gemini or any capable assistant.

Almost every prompt failure is one of five things, and each has a specific fix. Knowing which one you are looking at saves rewriting the whole prompt.

The 7 prompts

Intermediate 6 blanks to fill

Diagnose why a prompt is not working

Work out what is wrong before changing anything.

Prompt
My prompt is not working. Help me diagnose it before I start rewriting.

THE PROMPT:
"""
[PASTE]
"""
WHAT I WANT: [THE INTENDED OUTPUT]
WHAT I GET: [PASTE AN ACTUAL BAD OUTPUT]
HOW OFTEN IT FAILS: [always / usually / sometimes]
WHAT I HAVE ALREADY CHANGED: [IF ANYTHING]
MODEL: [WHICH]

Produce:

1. THE FAILURE CLASSIFICATION - what kind of wrong this is, because each has different causes:
   - WRONG TASK: it did something other than what I asked
   - RIGHT TASK, WRONG OUTPUT: correct approach, poor result
   - WRONG FORMAT: right content, wrong shape
   - TOO GENERIC: correct but says nothing specific
   - MISSING PARTS: some requirements followed, others dropped
   - FACTUALLY WRONG: confident and incorrect
   - INCONSISTENT: fine sometimes, not others
   Say which mine is.

2. THE FREQUENCY CLUE - consistent failure points to a prompt problem: something is wrong or missing. Intermittent failure points to underspecification: the prompt leaves a decision open and the model makes it differently each time. This distinction determines where to look.

3. THE LIKELY CAUSE - for my classification, the usual causes in order, each checked against my actual prompt:
   - Missing context the model needs and I have not supplied
   - Ambiguous instruction that can be read two ways
   - Conflicting instructions
   - The key instruction buried in the middle of a long prompt
   - Constraint stated only negatively
   - Format unspecified
   - Too many requirements, so some are dropped
   - The task itself underspecified

4. THE OUTPUT, ANALYSED - I pasted a bad output. Work backwards: what would the prompt have to mean for this output to be a reasonable response? That reading usually reveals the ambiguity.

5. THE ONE CHANGE - the single most likely fix. Change one thing and test, rather than rewriting everything, because a wholesale rewrite that works teaches nothing about what was wrong.

6. WHAT I HAVE ALREADY TRIED - if my previous changes did not help, that is evidence about the cause. Say what it rules out.

7. THE ISOLATION TEST - strip the prompt to its minimum, confirm that works, then add elements back one at a time. This finds the offending instruction faster than reasoning about it.

8. THE MODEL QUESTION - whether this is a prompt problem or a capability problem. Some tasks are not reliably doable by prompting, and no wording fixes that. Say if mine is one.

9. THE REWRITTEN PROMPT - after the diagnosis, not instead of it.

10. THE TEST - what to check in the output to confirm the fix, and what to keep as a test case so the problem is caught if it recurs.

What you get: The failure classified, the frequency clue interpreted, likely causes checked against your prompt, the bad output read backwards and one change to try.

Tip: Point 4 is the fastest diagnostic. Asking what the prompt would have to mean for this output to be reasonable usually points straight at the ambiguity.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Intermediate 4 blanks to fill

Fix output that is generic and says nothing

Get specifics instead of plausible filler.

Prompt
My output is technically correct and says nothing useful. Help me fix it.

THE PROMPT:
"""
[PASTE]
"""
THE OUTPUT:
"""
[PASTE A TYPICAL BAD ONE]
"""
WHAT I ACTUALLY NEED: [WHAT WOULD MAKE IT USEFUL]
WHAT I KNOW THAT THE PROMPT DOES NOT SAY: [YOUR CONTEXT]

Produce:

1. WHY OUTPUT IS GENERIC - the causes, in order of frequency:
   - THE PROMPT IS GENERIC: a broad question gets a broad answer, correctly
   - NO CONTEXT: the model has no situation to be specific about, so it covers all situations
   - NO CONSTRAINTS: without limits it produces the safe middle
   - ASKING FOR ADVICE RATHER THAN A DECISION: 'what should I consider' produces a list; 'what should I do' produces an answer
   - THE MODEL LACKS THE INFORMATION: it cannot be specific about things it does not know
   Say which applies to mine.

2. THE MISSING CONTEXT - I said what I know that the prompt does not. That gap is usually the whole problem. The model is answering a general question because a general question is what it was asked.

3. THE SPECIFICITY INSTRUCTIONS - what to add:
   - Require a concrete example, number or named case in every section
   - Ban the safe hedge: 'it depends', 'various factors', 'consider your needs'
   - Ask for a recommendation rather than options, with the reasoning stated
   - Ask what it would not do, which forces a position

4. THE DECISION FRAMING - the single most effective change for advice-type prompts. Instead of asking what to consider, describe the situation and ask what to do. A model asked to decide is specific; a model asked to inform produces a checklist.

5. THE BANNED-PHRASES LIST - the specific filler in my output, quoted, with what to require instead.

6. THE NARROWING - a broad prompt cannot produce a specific answer. If mine covers several situations, split it or pick one. Say where my prompt is too broad.

7. THE 'WHAT WOULD AN EXPERT SAY' TEST - generic output is what a cautious generalist says. What would someone with real experience of my specific situation say instead? That is the target, and describing it in the prompt helps.

8. THE INFORMATION LIMIT - if the specificity I want requires knowledge the model does not have about my situation, no prompt change produces it. I have to supply it. Say what is needed.

9. THE REWRITTEN PROMPT.

10. THE CHECK - the test for whether output is specific enough: could this sentence appear in an answer to a different question? If yes, it is filler.

What you get: Why the output is generic, your missing context identified, specificity instructions, decision framing and a test for filler.

Tip: Point 4 changes advice-type output more than anything else. Asking what to consider produces a list; describing the situation and asking what to do produces an answer.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Intermediate 4 blanks to fill

Stop the model agreeing with everything you say

Get honest assessment instead of accommodation.

Prompt
The model agrees with whatever I say. Help me get an honest assessment.

WHAT I AM ASKING ABOUT: [YOUR PLAN, WORK, IDEA OR CLAIM]
MY PROMPT:
"""
[PASTE]
"""
WHAT I SUSPECT IS HAPPENING: [YOUR SENSE OF IT]
WHAT I ACTUALLY WANT TO KNOW: [THE REAL QUESTION]

Produce:

1. WHY THIS HAPPENS - models are trained to be helpful and agreeable, and they take the premise of a question as given. If I ask 'is my plan good?', the framing already suggests the answer I want. Agreement is the path of least resistance, and it takes deliberate instruction to get past it.

2. THE LEADING FRAMING IN MY PROMPT - quote where I have signalled the answer I want. Common tells: stating my position before asking, asking whether something is good rather than what is wrong with it, and describing my reasoning approvingly.

3. THE NEUTRAL REFRAMING - the same question asked without signalling. Frequently this alone changes the answer substantially.

4. THE ADVERSARIAL INSTRUCTION - asking for what is wrong rather than for an assessment. 'Identify the three weakest points and why a critic would attack them' produces more than 'what do you think'. Write it for my case.

5. THE WITHHOLD-MY-POSITION APPROACH - presenting the situation without saying what I think or what I have decided, so there is nothing to agree with. Often the most effective single change.

6. THE STEELMAN REQUEST - asking for the strongest case against, written as its best advocate would put it. This is more productive than asking for criticism, which produces hedged concerns.

7. THE 'DO NOT REASSURE ME' INSTRUCTION - explicit, and it works. Also: instructing it not to end on an encouraging note, because models reliably undo a critique in the final paragraph unless told not to.

8. THE FALSIFICATION QUESTION - asking what evidence would show I am wrong, and whether it exists. This converts an opinion request into an analytical one.

9. THE SEPARATE-EVALUATION APPROACH - presenting the work without saying it is mine. Models are less accommodating about work presented neutrally.

10. THE REWRITTEN PROMPT.

11. THE HONEST LIMIT - these techniques reduce accommodation and do not eliminate it. A model has no independent stake in my being right or wrong, which is useful, but it also has no real judgement about my situation. For a decision that matters, a person who knows the domain is not replaceable by a better-worded prompt.

What you get: The leading framing in your prompt quoted, a neutral reframing, adversarial instructions, the withhold-your-position approach and an honest limit.

Tip: Point 5 works better than any instruction. Describing the situation without saying what you think leaves nothing to agree with.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Beginner 4 blanks to fill

Deal with a model that will not do what you ask

Work out why a request is being declined or deflected.

Prompt
The model will not do what I am asking. Help me work out why.

WHAT I ASKED FOR:
"""
[PASTE YOUR PROMPT]
"""
WHAT IT SAID:
"""
[PASTE THE RESPONSE]
"""
WHAT I AM ACTUALLY TRYING TO DO: [THE REAL PURPOSE]
MODEL: [WHICH]

Produce:

1. WHAT KIND OF REFUSAL THIS IS - the categories are different and only some are addressable:
   - MISREAD: the request resembles something it should decline but is not. Rephrasing with context usually resolves it.
   - AMBIGUOUS: the request could be benign or not, and it defaulted to caution
   - CAPABILITY, PRESENTED AS REFUSAL: it cannot do this and is saying so awkwardly
   - PARTIAL: it did some of it and stopped
   - GENUINE: the request is something it should not do, and no rephrasing changes that
   Say which mine is, honestly.

2. IF IT IS A MISREAD - what in my wording triggered it. Usually a term with a dual meaning, a scenario resembling a harmful one, or a request stripped of the context that makes it ordinary. Quote the likely trigger.

3. THE CONTEXT FIX - the most effective response to a misread: state what the work is for and who you are doing it for. A request that is ambiguous without context is usually unambiguous with it.

4. THE REPHRASING - the same request stated clearly and with purpose, not disguised. Say plainly that dressing a request up to get past a refusal is not what this is: if the underlying request is one the model should decline, rewording it is trying to circumvent a judgement rather than clarify a misunderstanding, and I should not do it.

5. IF IT IS A CAPABILITY LIMIT - what the model actually cannot do here, and the alternative approach. Common cases: real-time information, acting on external systems, reading something it was not given, guaranteeing accuracy, and counting or calculating reliably.

6. IF IT IS PARTIAL - why it stopped. Usually length, or an instruction conflict, or reaching something it treats differently. Say which, and how to get the rest.

7. IF IT IS GENUINE - say so directly. Some requests will not be fulfilled by any model, and the useful response is to tell me that and, where there is one, what legitimate alternative achieves my actual goal.

8. THE ALTERNATIVE ROUTE - given my stated real purpose, whether there is a different way to get there that does not run into the issue at all.

9. THE MODEL DIFFERENCE - models vary in where they draw lines, and a different one may handle a genuinely benign request differently. Worth noting for misreads, and not a workaround for genuine refusals.

10. THE HONEST ASSESSMENT - based on what I said my real purpose is, whether this is a misunderstanding worth resolving or a request I should reconsider.

What you get: The refusal type classified, the likely trigger quoted, a context fix for misreads, capability limits distinguished and an honest assessment.

Tip: Point 3 resolves most of these. A request that looks concerning without context is usually obviously fine with two sentences of purpose.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Intermediate 4 blanks to fill

Fix a long conversation that has gone off track

Recover when a chat has drifted or degraded.

Prompt
This conversation has gone off track. Help me get it back.

WHAT I AM TRYING TO ACHIEVE: [THE GOAL]
WHAT HAS HAPPENED SO FAR: [SUMMARISE THE CONVERSATION]
WHAT IS GOING WRONG NOW: [repeating itself / lost the thread / ignoring earlier instructions / quality has dropped / stuck on a wrong approach]
HOW LONG THE CONVERSATION IS: [ROUGHLY]

Produce:

1. WHAT HAPPENS IN LONG CONVERSATIONS - the common degradations:
   - EARLY INSTRUCTIONS FADE: constraints set at the start carry less weight later
   - ACCUMULATED CONTEXT COMPETES: everything said is still present and pulls in different directions
   - COMMITMENT TO A WRONG APPROACH: once an approach is established, it persists even when corrected, because the whole conversation supports it
   - REPETITION: the same points recycled with small variations
   - DRIFT: gradual movement away from the original task
   Say which is happening in mine.

2. THE RESTART DECISION - the honest answer is frequently to start a new conversation. A long thread carrying a wrong approach is harder to redirect than to replace, because the accumulated context keeps pulling back. Say whether mine is recoverable or whether restarting is faster.

3. IF RESTARTING - the handover prompt: a fresh statement of the goal, the constraints, what has been established that is worth keeping, and what approach did not work and should not be repeated. This last part is the value of the failed conversation.

4. IF CONTINUING - the reset message: restating the goal and the constraints as a fresh instruction rather than as a correction. Explicitly setting aside what came before works better than asking it to change course, which is treated as an adjustment to the existing approach.

5. THE COMMITMENT PROBLEM - once a wrong approach is established, corrections produce variations on it rather than a different approach. The fix is to state explicitly that the previous approach should be abandoned entirely and to describe the new one positively, rather than saying what was wrong with the old one.

6. THE REPETITION FIX - if it is recycling, listing what has already been covered and asking for what is not in that list.

7. THE CONSOLIDATION - a message that summarises what has been established and agreed, which then acts as the working context rather than the whole thread. Useful mid-project.

8. THE PREVENTION - for next time: setting the task and constraints clearly at the start, restating key constraints when the conversation gets long, and starting a new conversation when moving to a genuinely different task rather than continuing in the same thread.

9. THE RECOVERY MESSAGE - written for my situation, ready to send.

10. WHAT TO CARRY FORWARD - from what I described, what is worth keeping from this conversation and what should be left behind.

What you get: The degradation type identified, an honest restart-versus-continue decision, a handover prompt and the fix for a model committed to a wrong approach.

Tip: Point 5 is the practical insight. Correcting a wrong approach produces variations of it; abandoning it explicitly and describing the new one positively does not.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Intermediate 5 blanks to fill

Work out whether the model can do this at all

Stop trying to prompt your way past a capability limit.

Prompt
I cannot get this to work. Help me work out whether it is possible.

WHAT I AM TRYING TO GET IT TO DO: [THE TASK]
WHAT I HAVE TRIED: [YOUR ATTEMPTS]
WHAT KEEPS HAPPENING: [THE FAILURE]
HOW MANY DIFFERENT APPROACHES I HAVE TAKEN: [ROUGHLY]
MODEL: [WHICH]

Produce:

1. THE CAPABILITY ASSESSMENT - whether this task is something a language model can reliably do. Things they are structurally poor at, regardless of prompting:
   - Exact counting: characters, words, items in a long list
   - Reliable arithmetic, particularly multi-step
   - Guaranteed adherence to a hard numeric constraint
   - Knowing anything after training, or about my private data unless supplied
   - Perfect consistency across runs
   - Verifying its own factual accuracy
   - Anything requiring real-time information or action in the world
   Say whether mine falls into one of these.

2. THE RELIABILITY QUESTION - separate from capability. Many tasks a model can do sometimes it cannot do reliably. If I need it right every time and it is right most of the time, that is a different problem from an impossible task, and it is solved by validation rather than by prompting.

3. THE ATTEMPT ANALYSIS - I said what I have tried. If several genuinely different approaches have failed in the same way, that is evidence of a capability limit rather than a prompting problem. Say what my pattern suggests.

4. THE DIMINISHING RETURNS POINT - after a few serious attempts, further prompt variations rarely help. Say whether I have reached it.

5. IF IT IS A CAPABILITY LIMIT - the alternatives:
   - Do the unreliable part in code and use the model for the rest. Counting, arithmetic and exact constraints belong in code.
   - Use a platform feature: structured output, tool calling, code execution, or retrieval
   - Change the task so the model does the part it is good at
   - Accept an imperfect result with validation
   - Do it another way entirely
   For my task, which applies.

6. THE HYBRID DESIGN - usually the right answer. Which parts of my task suit a model and which suit code or a rule. The division is normally obvious once stated.

7. IF IT IS A RELIABILITY PROBLEM - the validation approach: check the output, retry on failure, and route failures to review. Say what to check for my task.

8. THE HONEST VERDICT - possible, possible but unreliable, or not something to pursue this way. Be direct rather than encouraging another attempt.

9. IF IT IS POSSIBLE - what I have not tried, specifically. Only offer this if there is genuinely a materially different approach, not another rewording.

10. THE TIME SPENT - if I have been at this a while, say plainly whether continuing is worth it. Recognising a limit is faster than working around it indefinitely.

What you get: A capability assessment, capability distinguished from reliability, your attempt pattern interpreted, hybrid alternatives and a direct verdict.

Tip: Point 6 is the answer to most of these. Counting and arithmetic belong in code, and the model handles the part that needs judgement.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026
Intermediate 5 blanks to fill

Test a prompt properly before relying on it

Find out whether a prompt works beyond the case you built it on.

Prompt
Help me test this prompt before I rely on it.

THE PROMPT:
"""
[PASTE]
"""
WHAT IT DOES: [THE TASK]
HOW IT WILL BE USED: [by me occasionally / by colleagues / in something automated / customer-facing]
HOW I HAVE TESTED IT SO FAR: [WHAT YOU HAVE TRIED]
WHAT WOULD GO WRONG IF IT FAILS: [CONSEQUENCE]

Produce:

1. THE DEVELOPMENT BIAS - a prompt developed against one or two examples is tuned to those examples. It works on what it was built on and fails on the range of real inputs. Say how much of my testing is this.

2. THE TEST CASE CATEGORIES - what to test with:
   - TYPICAL: the ordinary case, several of them, not one
   - BOUNDARY: shortest and longest plausible input, minimum and maximum values
   - DEGENERATE: empty, whitespace, a single word, unusual characters, another language
   - AMBIGUOUS: input where a reasonable person would need to ask a question
   - OUT OF SCOPE: input the prompt was not designed for, to check what it does
   - ADVERSARIAL: input containing instructions, if the prompt will receive content from elsewhere
   Write specific cases in each category for my prompt.

3. THE REPEAT RUNS - the same input several times. Variation between runs on identical input reveals underspecification that single runs hide, and it is the test people most often skip.

4. THE PASS CRITERIA - what counts as a correct output, defined before testing rather than judged afterwards. Without this, testing becomes reading outputs and feeling that they seem fine.

5. THE AUTOMATABLE CHECKS - which criteria can be checked by code: format validity, required elements present, length bounds, absence of forbidden content. Automating these means the test set can be re-run cheaply, which is what makes it useful over time.

6. THE PROPORTIONATE EFFORT - given how I said it will be used and what happens if it fails. For occasional personal use, a handful of varied inputs is enough; for anything automated or customer-facing it is not. Say what mine warrants.

7. THE FAILURE EXPECTATION - roughly what failure rate is acceptable for my use, and what happens to the failures. Nothing is perfect, and deciding the tolerance in advance is what determines whether validation is needed.

8. THE REGRESSION SET - keeping the test cases, including every failure ever found, to re-run whenever the prompt changes or the model updates. This is the durable output of testing and the reason it is worth doing properly once.

9. THE BLIND SPOT - what my testing will not catch, so I know what remains uncertain.

10. THE FIRST TEN CASES - written out, ready to run.

What you get: Your development bias named, test cases across six categories, repeat runs, pre-defined pass criteria and a reusable regression set.

Tip: Point 3 finds problems nothing else does. Running the same input five times exposes every decision your prompt left open.

Open in Written for Claude, ChatGPT, Gemini · Reviewed September 18, 2026

Where AI actually helps here

  • Isolating the failure: remove instructions one at a time until it works
  • Checking for contradictions, which are the most common cause of ignored instructions
  • Testing the same prompt several times, because output varies between identical runs

Where it falls down

  • Adding emphasis. Capitals and ‘IMPORTANT’ do less than people believe
  • Assuming it will improve on a retry. If the prompt is wrong, ten retries are ten wrong answers
  • Debugging a prompt and its input at the same time

The mistake almost everyone makes: Adding instructions instead of removing them

When a prompt is not followed, the instinct is to add a firmer instruction. Usually the cause is that an existing instruction contradicts it, or there are so many that each gets partial attention. Cut the prompt in half and see whether the problem persists — that one test localises most failures in a minute.

Free tool: Prompt Improver

Runs in your browser. No sign-up, nothing uploaded.

Open the Prompt Improver →

Questions people ask


Why does the model ignore part of my prompt?

Most often a contradiction elsewhere in the prompt, or instruction overload. Buried instructions in the middle of a long prompt also get less attention than ones at the start or end.


Why do I get different answers to the same prompt?

Generation is probabilistic by default. If you need determinism, lower the temperature where the API exposes it — and even then, expect variation rather than identity.