Free tool

Prompting in 2026

Most of what you were taught about prompting is now measurably wrong. Not out of fashion: tested, at scale, and wrong. Answer six questions about how you actually work and I will email you which of your habits the 2026 research has overturned, what to do instead, and the papers so you can check me.

  • Answer six questions about how you prompt today
  • Your habits are scored against the five 2026 papers
  • The guide and your personalised feedback are emailed to you

Five studies published in 2026 overturned advice that was standard practice a year earlier. Cornell Tech found that ending a prompt with a confident tag like right makes newer models agree with you rather than assess the evidence. IBM Research ran over 430,000 evaluations and found think step by step lost to plain asking. Microsoft found roughly 1 in 20 NeurIPS 2025 papers carried at least two likely hallucinated citations that passed peer review. Meta measured that models satisfy all of eight simultaneous rules just 5.7% of the time. EPFL, Apple and Mistral found worked examples now drag modern models down, with one model rising from 74% to 83.8% once they were deleted. The pattern matters more than any single finding: prompting knowledge decays, so AI fluency is maintenance rather than a certificate.

The check

Six questions about how you actually prompt

Answer honestly rather than aspirationally. The feedback is only useful if it describes what you really do. 0 of 6 answered.

1. When you want a model to check your reasoning, how do you usually phrase it?

What the research found. Cornell Tech tested 45 models. Ending with a confident tag like "right?" made newer models disagree with you; swapping it for a tentative "maybe?" made all 45 more agreeable. The model mirrors your confidence, not the evidence. Tag Questions, Parikh, Cornell Tech, Jul 2026 · arxiv.org/abs/2607.23976

2. Do you add "think step by step" or similar to your prompts?

What the research found. Across more than 430,000 evaluations, IBM Research found the most famous prompt phrase in the world lost to plain asking. What edged ahead was the question plus a short role. The Simplicity Paradox, IBM Research, 2026 · arxiv.org/abs/2607.14109

3. When a model gives you a citation, a statistic, or a source, what do you do with it?

What the research found. Microsoft researchers audited references at the top AI conferences and found roughly 1 in 20 NeurIPS 2025 papers carried at least two likely hallucinated citations. Those papers passed expert peer review. Phantom References, Russinovich, Siva Kumar and Salem, Microsoft, Jul 2026 · arxiv.org/abs/2607.00738

4. How many separate requirements do you put in a single prompt?

What the research found. Meta Superintelligence Labs measured the ceiling: at 8 simultaneous rules, models satisfy each individual rule about 41% of the time, and all eight together just 5.7%. Phase Transitions in Compositional Constraint Satisfaction, Vasileva, Meta, Aug 2026 · arxiv.org/abs/2608.12426

5. Do you include worked examples in your prompts?

What the research found. EPFL, Apple and Mistral researchers found worked examples now drag modern models down. Deleting them lifted one model from 74% to 83.8%. Soft Guidance Starts to Outperform CoT Prompting, Pushkin et al., Aug 2026 · arxiv.org/abs/2608.03550

6. When did you or your team last update how you prompt, based on new evidence?

What the research found. Every finding above invalidates advice that was best practice twelve months ago. Prompting knowledge has a half-life, which makes fluency a maintenance activity rather than a certificate. The pattern across all five 2026 papers · rushi.knowwhatson.com/ai-fluency-workshops/

Where should I send your result?

Answer all six questions so the feedback describes how you actually work.

I send you the result and, occasionally, something genuinely useful about AI adoption in Australia. No list swapping, no spam, unsubscribe any time.

Questions

What people ask before they start

Is prompt engineering still worth learning in 2026?

Yes, but not as a fixed skill. Every technique in this check was best practice within the last two years and has since been overturned by measurement. What is worth learning is the habit of retesting, because the advice has a half-life.

Does think step by step still work?

Not reliably. Across more than 430,000 evaluations, IBM Research found plain asking beat it, and a question plus a short role beat both. Reasoning models already do the stepwise work internally, so the instruction mostly adds constraint.

How many instructions can a model follow at once?

About three before quality falls away. Meta Superintelligence Labs measured that at eight simultaneous rules, each individual rule is satisfied roughly 41% of the time and all eight together just 5.7%. Ask for three, then revise one requirement at a time.

Can I trust citations an AI gives me?

No, and the reason is uncomfortable. Microsoft researchers found roughly 1 in 20 NeurIPS 2025 papers carried at least two likely hallucinated citations, and those papers passed expert peer review. Fabricated references read as fluent and correctly formatted, so intuition will not catch them. Open the source.