Prompt Engineering: The Techniques That Still Matter
Models got better at guessing what you meant, and most prompt tricks stopped mattering. Seven techniques did not: a clear task, context, examples, structure, an output format, iterating and evaluating. Here they are, from the vendors’ own guides, with a template you can reuse.
7 min read
Prompt engineering is the practice of writing instructions so that a language model does what you meant, the first time and every time after. The tricks that once filled guides, from capital letters to magic phrases, matter much less with current models. What still matters is what would matter for a capable new colleague: a clear task, the context they cannot guess, an example of good work, a prompt organized so instructions and material do not blur, the format you want back, and the habit of testing and revising. Anthropic, OpenAI and Google each publish a prompting guide, and on those points they agree.
Two neighbors cover the rest. How prompt engineering differs from managing everything else a model sees is in context engineering vs prompt engineering. Advice specific to Claude, with seven before-and-after rewrites, is in Claude prompting best practices.
What the three vendor guides agree on
- Start with a definition of success. Anthropic’s prompt engineering overview (opens in a new tab) assumes you already have success criteria, a way to test against them and a first draft, and says to set those up before tuning anything.
- Say exactly what you want. Google’s prompt design strategies (opens in a new tab) call clear and specific instructions the most effective way to customize a model’s behavior.
- Show examples. Google recommends always including a few examples; OpenAI describes the same thing as few-shot learning.
- Mark the parts. OpenAI’s prompt engineering guide (opens in a new tab) suggests Markdown and XML tags to show the model where instructions end and material begins.
- Give it the facts. All three say to put the information the model needs in the prompt rather than assume it knows.
- Test and revise. Each treats the first prompt as a draft to be checked against real cases.
Seven prompt engineering techniques that still matter
1. A clear task
Name the output, the audience and the limits. “Write about our refund policy” leaves the model to guess the reader, the length and the purpose. “Write a 120-word answer for our help center explaining how a customer in the US returns an item within 30 days, for someone who has never ordered from us” leaves nothing to guess. If order matters, number the steps.
2. Context the model cannot guess
The model knows a great deal in general and nothing about your situation: who the customer is, what was decided last week, why a rule exists. Give the reason behind an instruction as well as the instruction, because a model that knows why generalizes better. Put long material first and the question at the end; Google’s guide gives the same advice for large amounts of context, and Anthropic’s says the same for Claude.
3. Examples
One good example fixes a format better than a paragraph describing it. Use a few, make them close to your real case, vary them so the model does not copy an accident of one, and keep their structure identical, which Google warns is how you avoid answers in formats you did not want. Show the pattern to follow rather than a list of mistakes to avoid.
4. Structure
A prompt that mixes instructions, background and a pasted document becomes ambiguous about which is which. Separate them with headings or tags such as <document> and <instructions>. It costs a few characters and stops the model treating a sentence in your material as an order.
5. An output format
Say what the answer should look like: a table with named columns, three bullets under 15 words, JSON with given fields, or plain paragraphs. Say what to do rather than what not to do: “write in plain paragraphs” works better than “no bullet points”. Where a fact is missing, tell the model what to write instead of inventing one, such as “not stated”.
6. Iterate on real cases
Google’s guide says plainly that prompt design can take a few iterations. Run the draft on five or ten real inputs, not the one you wrote it for. When an answer is wrong, change one thing at a time, so you know which change fixed it.
7. Evaluate before you trust it
For anything you will reuse, write down what a good answer is and check against it. Anthropic’s page on success criteria and evaluations (opens in a new tab) asks for criteria that are specific, measurable, achievable and relevant, tests that mirror your real task including edge cases, and automated grading where possible, and it prefers more test cases graded automatically to a few graded by hand. For a work prompt that can be as simple as ten saved inputs and a checklist.
A reusable prompt template
Every part of the template maps to one of the techniques. Delete what a small job does not need; keep the order.
## Role You are [who the model should act as, in one sentence]. ## Task [What to produce, for whom, and why it matters.] ## Context [Facts it cannot guess: the situation, decisions already made, constraints.] ## Material <material> [Paste documents, data or text here. Long material goes before the questions.] </material> ## Example of a good answer <example> [One to three short examples, identical in structure.] </example> ## Output format [Length, structure, tone. What to write when something is missing.] ## Check before answering [One to three tests the answer must pass.]
A worked example
An online store in Ohio wants a weekly summary of support tickets. The first attempt was “Summarize these tickets.” Filled in, the template becomes this:
## Role You are a support lead at a small online store that ships across the US. ## Task Summarize this week's support tickets for the owner, who decides on Monday what to fix first. She reads it on her phone. ## Context Checkout moved to a new payment provider on Tuesday. Sales tax is charged by ZIP code. We can fix about three problems a week. ## Material <material> [paste this week's tickets] </material> ## Example of a good answer <example> bug: Sales tax wrong for some ZIP codes at checkout (9 tickets). Customers in two Ohio counties were charged the state rate only. </example> ## Output format At most five lines, most tickets first. Start each line with bug:, feature: or enhancement:, then the problem in under 15 words, then the ticket count. If a cause is not in the tickets, write "cause not stated". ## Check before answering Every line is backed by at least one ticket. Nothing about the payment change is blamed unless a ticket says so.
What changed: a reader and a decision, the one fact that explains this week (the payment change), a limit that forces a ranking, an example that fixes the format, a rule against invented causes, and a check the model runs on itself.
What matters less than it used to
- Shouting. Capital letters and “CRITICAL” were a way to make older models pay attention. Current models follow plain instructions and can overreact to emphatic ones; the Claude post above covers Anthropic’s advice on this.
- “Think step by step.” Reasoning models decide how much to think on their own. OpenAI compares a reasoning model to a senior coworker you give a goal and a GPT model to a junior one who needs explicit instructions, so match the prompt to the model.
- Fiddling with settings. Google now recommends leaving temperature and similar parameters at their defaults for its current Gemini models.
- Clever phrasing in general. A plain request with the right context beats a clever one without it.
When the prompt is not the problem
Anthropic notes that not every failing test is best solved by prompt engineering; sometimes a different model is the easier fix for speed or cost. If the model lacks knowledge it cannot be told in a prompt, retrieval or fine-tuning may fit better, as prompt engineering vs RAG vs fine-tuning explains. If it reads material written by someone else, a web page or an email, that material can carry instructions of its own; see indirect prompt injection. And if you want to know why a model sometimes answers fluently and wrongly, what is an LLM explains how it produces text in the first place.
Keep the prompts that work
A good prompt lost in a chat history gets rewritten from memory next month, worse. Keep the ones that earn it where the team and its assistants can find them. On a fenbs board, AI context holds the short standing notes every connected assistant reads first with fenbs_get_context, a task’s note holds the problem and its plan holds how it will be done, and the test status and test notes record how the result was checked. That covers the context, the task and the check. fenbs does not version prompts or run evaluations for you, so keep saved test cases beside the prompt, and see how to write a task for an AI agent for turning a prompt into a task an agent can finish.
Related
Copyable prompts for running a board: prompts for project management AI. Choosing which assistant to prompt: ChatGPT alternatives. Connecting one to a board: the MCP docs.