Prompt Engineering vs RAG vs Fine-Tuning: Which Fixes Your Problem

Three ways to make a model do better, and each one fixes a different failure. Start from the symptom you can see, and the choice between prompting, retrieval and training mostly makes itself.

8 min read

Prompt engineering changes what you ask. Retrieval-augmented generation (RAG) changes what the model can see when it answers. Fine-tuning changes the model itself. So the right choice depends on what is actually wrong: if the model has the facts but answers in the wrong shape or voice, fix the prompt; if it lacks your private or current facts, retrieve them; if a narrow task is done thousands of times a day and a well-built prompt is still too long, too slow or not consistent enough, consider fine-tuning. Try them roughly in that order, because each step costs more effort than the one before and is harder to undo.

What each technique changes

  • Prompt engineering changes the instructions, examples and structure in the request. Nothing outside the request moves, so a change is a text edit you can test in minutes and revert just as fast.
  • RAG changes the material in the request. Before the model answers, your application searches a body of documents and adds the relevant passages. Anthropic’s glossary (opens in a new tab) describes it as grounding the answer in an external knowledge base retrieved at runtime, and notes it is most useful for up-to-date information, domain-specific knowledge and answers that must cite their sources.
  • Fine-tuning changes the weights. You train a pretrained model further on examples of the inputs and outputs you want, and you get a new model that behaves that way without being told each time. The same glossary entry says it can adapt a model to a domain, a task or a writing style, and warns that it needs care over the training data and its effect on the model.

The three are not rivals. A production system often uses all of them: a tuned model, a retrieval step, and a carefully written prompt around both. The useful question is which one to reach for next, given the failure in front of you.

Choose by symptom

The answer is right but in the wrong format

This is a prompt problem first. Say exactly what shape you want, show one or two examples of it, and say what to leave out. If the output feeds another program and must parse every time, use the API’s schema features rather than training: Claude’s structured outputs (opens in a new tab) constrain the response to a JSON schema you supply, and other vendors offer their own equivalents. Fine-tuning for format makes sense only when the format is unusual, hard to describe and needed at very high volume.

It does not know your private information

Your contracts, your product docs and your customer history were never in the training data, and no prompt wording will make the model know them. Put them in front of the model at the moment it needs them, with retrieval or with a tool it can call. Training the facts in is the weaker option: the model cannot tell you which document an answer came from, you cannot take one fact back out without retraining, and anyone who can use the model can draw on everything it was trained on.

Its facts are out of date

Stale facts are the same problem with a clock on it. A model knows the world up to its training data, and a fine-tuned model knows it up to your training run. If the answer depends on today’s price list, this week’s policy or the current state of a ticket, fetch it at request time. Retrieval over documents works for text that changes weekly; a tool that reads the live system is better for anything that changes by the minute.

The tone or voice is wrong

Start with the prompt: describe the voice, give a short example of it, and say what to avoid. That fixes most cases. Fine-tuning earns its place when you need the same voice across a very large volume of output and the examples needed to hold it steady have made the prompt long. Microsoft’s fine-tuning considerations (opens in a new tab) list exactly these cases: reducing prompt overhead that has built up from many examples, modifying style and tone, and producing specific formats or schemas.

A narrow task, repeated thousands of times

Classifying support tickets, extracting fields from invoices, tagging products: the task is fixed, the inputs vary, and volume is high. Here fine-tuning is at its strongest, especially to teach a smaller model to do one job as well as a larger one does with a long prompt. OpenAI’s model optimisation guide (opens in a new tab) lists the benefits over prompting alone: more examples than fit in one request, shorter prompts that save tokens and can lower latency, training on data you would rather not send with every request, and a smaller, cheaper model that excels at one task. Build an evaluation set before you train, so you can prove the tuned model is better and not merely different.

It is too slow or too expensive

Check the cheap levers first. A smaller or faster model may already pass your tests; Anthropic’s prompt engineering overview points out that latency and cost are sometimes easier to fix by selecting a different model than by rewriting the prompt. Trim the prompt and the retrieved passages to what the model actually uses. Only if a small model still falls short on a high-volume task is fine-tuning the answer, because it lets the small model carry behaviour that would otherwise need a long prompt on every call.

What each one costs in effort

  • Prompt engineering is the cheapest to try and to reverse. It needs a way to test outputs against clear criteria, and a person who writes precisely. Its running cost grows with prompt length, since every token is paid for on every call.
  • RAG costs more to build and to keep. You need a pipeline that splits and indexes documents, a search step that finds the right passages, and someone who notices when retrieval misses. Its quality is capped by the quality of what it retrieves.
  • Fine-tuning costs the most up front. You need hundreds to thousands of good examples, an evaluation set, training runs and a way to deploy and version the result. When the base model is retired or improved, you may have to do it again.

Google’s guide to tuning Gemini models (opens in a new tab) gives the same order: start with prompting to find the best prompt, then move to fine-tuning if required, and look at where the model makes mistakes before adding more data.

Where context engineering, tools and MCP fit

These three labels come from an era of single questions and single answers. With agents, most of what the model reads is not your prompt at all but tool results, file contents and earlier turns. Deciding what should be in the window at each step is context engineering, and it contains prompt engineering rather than replacing it; the boundary is set out in context engineering vs prompt engineering.

Tools are the other half of the “missing facts” answer. Classic RAG decides what to fetch before the model runs. A tool, often published through an MCP server, lets the model decide what to look up, look again when the first search misses, and act on what it finds. For live systems such as a task board, a database or a ticket queue, a tool usually beats pre-indexed retrieval. The trade-offs are covered in MCP vs RAG, and the writing side of prompts in Claude prompting best practices.

Who offers fine-tuning today

Availability changes, so check the vendor’s page before you plan around it. As documented today:

  • Anthropic: the Claude API does not currently offer fine-tuning; the glossary says to ask your Anthropic contact if you want to explore it. Claude is steered through prompts, context, tools and structured outputs.
  • Amazon Bedrock: supervised fine-tuning, reinforcement fine-tuning and distillation. Its list of fine-tunable models (opens in a new tab) includes Anthropic’s Claude 3 Haiku in one US region, alongside Amazon Nova and Meta Llama models.
  • Google: supervised fine-tuning and preference tuning for several Gemini models, with the list of supported models on the tuning page above.
  • Microsoft Foundry: fine-tuning for a range of models, with guidance on when it is worth it.
  • OpenAI: its model optimisation guide says it is winding down the fine-tuning platform. It is closed to new users; existing users can create training jobs for the coming months, and fine-tuned models stay available until their base models are deprecated.

The practical reading: if your stack is built on Claude, plan around prompts, retrieval and tools, and treat fine-tuning as an exception you would take up with Anthropic or run on an older model in Bedrock.

A decision order you can copy

Which fix to try next
1. Write down the failure with 5-10 real examples.
2. Did the model have the facts it needed?
   No  -> missing or stale facts: add retrieval or a tool.
   Yes -> go to 3.
3. Is the output wrong in shape, tone or steps?
   -> rewrite the prompt: clearer rules, 1-3 examples.
   -> machine-read output: use a schema feature.
4. Still failing, on a narrow task at high volume?
   -> try a smaller model first, then fine-tune one
      against an evaluation set you built in step 1.
5. Too slow or costly? Smaller model, shorter prompt,
   fewer retrieved passages, then fine-tuning.

Keeping the experiments straight

The expensive mistake here is not choosing the wrong technique; it is forgetting what you already tried. Each attempt is a small piece of work with a hypothesis, a change and a result, and on a shared board it reads naturally as a task. On fenbs, file the failure as a bug, with the examples in the note. Put the technique you are trying, and why, in the plan. When the evaluation runs, set the test status to Tested, Partly tested or Failed and write the scores and what was not checked in the test notes. The next person, or the next assistant, can then see that the prompt rewrite got half the way and retrieval did the rest, instead of starting again.

Standing rules that every assistant should follow, such as “cite the document you used”, belong in the board’s AI context, which connected assistants read when they arrive. That is prompt and context work done once, for everyone, rather than retyped in each chat.

Related

The boundary between prompts and everything else the model sees: context engineering vs prompt engineering. Retrieval versus tools: MCP vs RAG. Building the test set you need before any of this: how to evaluate AI agents.

Questions people ask.

Should I use RAG or fine-tuning to add company knowledge?

Usually RAG, or a tool the model can call. Retrieval shows the model the current document at the moment it answers, can cite the source, and changes when the document changes. Fine-tuning bakes a snapshot into the model, cannot tell you where an answer came from, and needs retraining to update.

Is prompt engineering enough on its own?

For format, tone, steps and most reasoning problems, often yes, provided the model already has the facts it needs. It cannot supply private or recent information, which is where retrieval or tools come in.

Can I fine-tune Claude?

Anthropic’s documentation says the Claude API does not currently offer fine-tuning, and suggests asking your Anthropic contact if you want to explore it. Amazon Bedrock lists Claude 3 Haiku among its fine-tunable models in one US region.

Can I combine all three?

Yes, and many systems do: a tuned model for a narrow task, retrieval for the facts, and a prompt that ties them together. Add each layer only when a test shows the previous one is not enough.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.