Prompting advice spread for three years faster than anyone could test it. The testing has now caught up, and the results are unkind to most of the famous techniques. What follows is ten beliefs, each with what the evidence actually says, and each with something usable in place of the belief.

One framing note before the list. These are findings about measured accuracy on benchmark tasks. A technique that does not move accuracy may still be useful for controlling tone or format, and that is a different claim from the one usually made for it.

10 BELIEFS / TESTED RITUAL SPECIFICATION
A clear brief works on every model generation. A trick works on one.

1. Offering a tip or making a threat improves output

Tested directly and at scale. A Wharton study ran roughly 19,800 model responses across tipping offers, threats and neutral prompts and found no statistically significant effect on accuracy at the aggregate level. Individual questions moved in both directions, which is exactly what noise looks like.

The origin of this belief is instructive. It came from public anecdote, was repeated by prominent people, and was never accompanied by a controlled test until someone bothered to run one.

2. Assigning an expert persona makes the model smarter

A study across 162 personas, 2,410 questions and nine models found no improvement, and in places a small negative effect. "You are a world-class expert in X" does not add knowledge the model lacked.

What personas do reliably is control register and vocabulary. If you want the answer written like a cinematographer, say so. If you want it to be more correct, that is not the lever.

3. Being polite gets better results

Effects show up in individual runs and vanish on aggregation. Politeness is free and costs nothing to keep, so keep it if you prefer. Just do not build a workflow around it or teach it as technique.

4. Emotional pressure raises performance

The often-cited uplift is under 2%, from a single-model study with a circular design. Worse, there is a documented cost: emotional pressure is one of the reliable ways to induce sycophancy, where the model agrees with you rather than assessing the thing. For anyone using a model to check their own work, that is the opposite of what you want.

5. Always ask the model to think step by step

For reasoning models this is now documented as counterproductive by the vendor. OpenAI's reasoning best practices guide says plainly: "Avoid chain-of-thought prompts." The model already runs an internal reasoning pass, and instructing it to narrate one on top is redundant and can degrade the result.

Chain-of-thought remains useful on non-reasoning models. The rule is conditional on which model you are talking to, which is true of far more prompting advice than people acknowledge.

6. Structured tags and role prompting are best practice

Anthropic has retracted its own earlier advice on XML tags and role prompting. This is a lab withdrawing guidance it published itself, which is about as strong a signal as this field produces.

The general point: prompting technique is coupled to model generation. Advice written for a 2023 model was correct for a 2023 model. Repeating it in 2026 as a timeless principle is a category error, and most prompting listicles are exactly that.

7. A bigger context window means you can just paste everything in

The most useful finding in this list. Research across 18 models shows that performance degrades as input length grows even when task difficulty is held constant. Length itself is the variable.

The counterintuitive detail: shuffled haystacks sometimes outperform coherent ones, which means the degradation is not simple distraction by related material. Whatever is happening, it is not the intuitive story.

Operationally: a million-token window is a capacity, not a recommendation. Curate what goes in. Retrieval into a short context beats dumping into a long one.

8. Asking for JSON is harmless

It is not, when you want ideas. Across 44 models, requesting structured output collapsed answer diversity, with the modal answer's share rising from 41% to 64%. The model converges on one answer far more often.

Asking for JSON narrows the answers

Share held by the modal (most common) answer, across 44 models

0%25%50%75%100% Without structured outputWithout structured output: 41%Without structured output: 41%41% Structured output requestedStructured output requested: 64%Structured output requested: 64%64%
Asking for JSON narrows the answers
ConditionModal answer's share
Without structured output41%
Structured output requested64%
Source: Structured Output Collapses Answer Diversity Across 44 Language Models.

Which is exactly what you want for extraction and exactly what you do not want for brainstorming. Ask for prose when you want range, JSON when you want a field filled.

9. Longer, more detailed prompts are better prompts

Half true, and the half that is false is expensive. Specificity helps. Length does not, and by finding 7 above, length actively hurts past a point. The useful discipline is to add constraints and remove decoration.

A test worth running on your own prompt library: cut every prompt by 40% while keeping every constraint, then compare. Most prompt libraries have a lot of ceremony in them that survives only because nobody has measured it.

10. Prompting is the skill

The through-line of every finding above is that the tricks do not carry much weight. What carries weight is the specification: what you want, in what form, with what constraints, judged against what standard. That is a writing and thinking problem, not a technique problem.

This is also why prompting advice keeps expiring while specification quality does not. A clear brief works on every model generation. A trick works on one.

What to do with this

  • Strip the folklore out of your prompt templates. Tips, threats, expert personas and emotional framing can come out with no measurable loss. That is dead weight you are paying tokens for.
  • Split your templates by model type. Reasoning models and non-reasoning models now want different prompts, and step-by-step instructions belong only in the second set.
  • Set a context budget per task and retrieve into it rather than pasting whole documents. Treat the window size as a ceiling you stay well under.
  • Choose output format by intent. Prose for divergence, JSON for extraction. Never ask for JSON in an ideation step.
  • Measure one change at a time. Twenty runs per variant on a task you actually do. Almost nobody does this, which is precisely why folklore survived three years.

The short version

The famous prompting tricks were mostly never tested, and when they were tested most of them did nothing. What survives contact with evidence is unglamorous: be specific, be short, know which kind of model you are talking to, control the context deliberately, and pick the output format for the job. Everything else is ritual.

Sources

Gabriel Brien

Gabriel Brien

Founder of Crimson Spark Agency. AI filmmaker and creative technologist, writing from running this work daily.

More about Gabriel →