How to Write Instructions for an AI Agent: A System Prompt Template

Standing instructions tell an agent what is true for every task; the task card tells it what to do now. What belongs in each, a six-part template to copy, the mistakes that make agents ignore instructions, and how to test a prompt before you trust it.

7 min read

Good instructions for an AI agent come in two layers. The standing instructions, a system prompt or a file such as AGENTS.md, hold what is true for every task: who the agent is working for, what it needs to know about the project, the rules and why each exists, which tools to use, when to stop and ask, and what its report should look like. The task itself holds what is true only now: the outcome, the place, the acceptance criteria. Keep the two apart, give every rule a reason, write what to do rather than only what to avoid, and test the instructions on real tasks before you trust them. The template below has the six parts.

Standing instructions versus the task card

The commonest mistake is mixing the layers. A system prompt that mentions this week’s bug is out of date by Friday, and a task card that repeats the coding standards is ten lines longer than it needs to be and disagrees with the next card. A simple test sorts any line: would it still be true for the next task? If yes, it is standing. If no, it belongs on the card.

  • Standing: the agent’s job and who it works for; what the project is; where things live; how to run and check the work; rules that never change; tools and when to use them; when to stop; the shape of the report.
  • The task: the outcome, where the work is, what has been tried, constraints that apply only here, acceptance criteria and how they will be checked. Writing that part well is covered in giving an AI agent a task it can finish.

What the model makers advise

Anthropic’s prompting best practices (opens in a new tab) describe the model as a brilliant but new employee who lacks context on your norms, and offer a golden rule: show the prompt to a colleague with little context on the task, and if they would be confused, so will the model. The same page recommends explaining why an instruction matters, because the model generalises from the reason, and telling it what to do instead of what not to do.

Anthropic’s engineering team adds, in its piece on context engineering for agents (opens in a new tab), that a system prompt should sit at the right altitude: not brittle if-else logic, not vague guidance that assumes shared context, but the minimal set of information that fully outlines the expected behaviour. It advises against stuffing in a laundry list of edge cases, and in favour of a few diverse, canonical examples.

OpenAI’s prompt engineering guide (opens in a new tab) suggests a developer message usually runs identity, instructions, examples, then context, marked out with Markdown headings and XML tags.

OpenAI’s GPT-5 prompting guide (opens in a new tab) warns that contradictory or vague instructions do more damage to a model that follows instructions precisely, because it spends effort reconciling them, and recommends stating the stop conditions, which actions are safe and which are not, and when the agent should hand back to the user.

A system prompt template for agents

Six sections, in this order. Replace everything in angle brackets; delete any line you cannot make specific.

Standing instructions
# Role
You are the coding agent for <project>, working for <team or person>.
Your job is <the kind of work you hand it>. You do not <what is out of bounds>.

# Context
- What this is: <product, who uses it, what stage it is at>.
- Where things are: <main folders, services, the one file to read first>.
- How to check work: <test command>, <lint command>, <how to run it locally>.
- Before any task, read <the task board / the context notes> for current rules.

# Rules, each with its reason
- Never run anything against the live database, because it holds customer
  data and there is no undo.
- Change only the files the task needs, because reviewers read the diff.
- Never edit a test to make it pass, because the test is how we know it works.
- Write user-facing text in British English, because our customers are in the UK.

# Tools
- Use <tool> for <job>. Prefer it to <alternative>, because <reason>.
- The task board is the list of work: read your task before you start and
  comment on it when you stop.

# When to stop and ask
Ask before you delete anything you did not create, push, deploy, send a
message, or change anything other people can see. Ask when the task can be
read two ways that lead to different results. If you are blocked, say what you
tried and what you need, then stop. Do not guess.

# Output
When you finish, reply with:
1. What changed, one line per file.
2. How you checked it: the command and the result.
3. What you did not do, and anything you noticed but left alone.

The six parts, briefly

  • Role: a job, not a compliment. “The coding agent for the billing service” tells it what to attend to; “a world-class engineer” tells it nothing it can act on.
  • Context: only what the agent cannot find out quickly by itself. Point at files rather than pasting them, so the instructions do not go stale when the files change.
  • Rules with reasons: the reason lets the agent handle the case you did not foresee. “Never touch live, because there is no undo” also covers the live queue you forgot to mention.
  • Tools: which to use for what, and which to prefer when two could do the job. An agent with a search tool and a shell will otherwise pick at random.
  • When to stop: name the irreversible and the visible. Everything else it may do and report. This is what turns “autonomous” into something you can leave running.
  • Output: a fixed report shape makes a day of agent work quick to review, and makes a missing check obvious.

Anti-patterns

  • Rules without reasons. They are followed literally and fail at the edges.
  • Only prohibitions. A list of “do not” lines leaves the agent to guess what you want instead.
  • Two rules that disagree, often in two files. The agent either picks one or spends its effort trying to satisfy both.
  • This week’s work in the standing instructions. It goes stale and gets obeyed long after it stops being true.
  • Every edge case you have ever met. Long lists dilute the rules that matter; keep a few examples instead.
  • Secrets, tokens or connection strings. Instructions are read, logged and shared; credentials belong in the tool’s own configuration.
  • The same rules copied into four tool-specific files by hand. They drift apart. Keep one source and point the others at it, as AI context files compared describes.

Testing your instructions

  1. Run the colleague test first. Give the prompt to someone who has not seen the project and ask what they would do on day one.
  2. Pick five to ten real tasks from your history, including one that should make the agent stop and ask. These are your fixtures.
  3. Run each in a fresh session, with nothing but the instructions and the task, and score the result against a short list: did it read what it should, stay in scope, run the check, stop where it should, and report in the right shape?
  4. Change one thing at a time and rerun the fixtures. Anthropic suggests starting from a minimal prompt and adding instructions for the failures you actually see; OpenAI recommends evaluation suites and pinning model versions, so a model upgrade does not change behaviour unnoticed.
  5. Check the file loaded at all. Each tool has a command that shows it, such as /context in Claude Code or /instructions in Copilot CLI. Many “ignored” instructions were never read.

Where the instructions live, per tool

Standing rules that outlive the tool

Some standing instructions are about the project rather than the codebase: what never to do for a client, how tasks should be written, a decision you made last month. They are the same for every assistant, so they are better kept where every assistant reads them. On fenbs that is AI context: notes kept on the board, which any assistant you connect fetches with fenbs_get_context before it starts, and which an assistant can add to, signed with its name. Decisions have their own page and their own numbers, and an assistant searches them with fenbs_list_decisions before it changes something that looks deliberate. The instructions file then needs one line for the board, and the task card on the board carries the rest.

Related

Writing the other half: giving an AI agent a task it can finish and AI agent task examples. What fills the context window around your instructions: context engineering for AI agents. Connecting an assistant to a board: MCP docs.

Questions people ask.

What should an AI agent system prompt include?

Six things: the agent’s role, the project context it cannot discover quickly, the rules with a reason for each, the tools and when to use them, when to stop and ask, and the shape of the report it gives when it finishes. Anything that is true for only one task belongs on the task instead.

How long should agent instructions be?

As short as they can be while still covering the behaviour you expect. Anthropic describes the goal as the minimal set of information that fully outlines it, which is not the same as short. Cut any line the agent would do correctly without it, and move task details to the task.

Should I tell an AI agent what not to do?

Say what to do first, and keep prohibitions for the few things that must never happen, each with its reason. A prohibition alone leaves the agent guessing at the alternative; a reason lets it apply the rule to cases you did not list.

How do I know if my agent instructions work?

Run a handful of real past tasks in fresh sessions and score the results against a short checklist: stayed in scope, ran the check, stopped where it should, reported in the right shape. Change one instruction at a time and rerun, and repeat after any model upgrade.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.