RAG or Fine-Tuning? Choosing the Right Approach for Knowledge
Fine-tuning gets attention and a budget line. Retrieval gets results. For anything that changes — prices, policies, inventory, specs — the answer is almost always retrieval.
Nearly every conversation about company AI starts as a debate between fine-tuning and retrieval, as if they were competitors. They are not. They answer different questions, and most projects that get stuck have mixed them up.
The question behind the question
Ask what fails today. Usually it is one of three things:
- The model does not know something. That is a knowledge problem, and knowledge retrieval solves.
- The model knows it but behaves wrongly — wrong format, wrong tone, wrong structure, ignoring steps. That is a behaviour problem, and fine-tuning or better prompting solves it.
- The model knows and behaves correctly, but the answer is expensive or slow. That is an architecture problem, and neither of these is the fix.
Fine-tuning cannot make a model aware of a policy you revised last Tuesday. It has no mechanism to look anything up. It will answer from its weights, which are a snapshot, and it will answer with the same confidence either way. This is the most important sentence in this article.
What fine-tuning actually changes
Fine-tuning adjusts distribution. It changes what the model tends to produce — format, tone, structure, refusal behaviour, domain vocabulary. With good data it reliably makes output shape predictable, which matters when downstream code parses it.
It does not add knowledge. It does not add facts about your company. It cannot reference a document it has never seen, and it will not tell you where an answer came from.
What retrieval actually changes
Retrieval changes the input. Before the model answers, your system fetches relevant passages from a corpus you control and puts them in the context window with instructions to answer only from them.
This gives you three things fine-tuning cannot: freshness, because the corpus updates when documents do; provenance, because you know exactly which passage was used; and reversibility, because fixing an answer means fixing a document rather than retraining a model.
The cost is retrieval quality. If the wrong paragraph arrives, no model saves you.
Choosing by requirement
| Requirement | Retrieval | Fine-tuning |
|---|---|---|
| Answers reflect a policy revised this week | Yes | No |
| Every answer cites a source | Yes | No |
| Consistent JSON or fixed output shape | Partly | Yes |
| Specific tone or house style | Partly | Yes |
| Refuse reliably on out-of-scope questions | With guardrails | Yes |
| Remove sensitive data before inference | Yes | No |
| Cost per query drops materially | Yes | Sometimes |
| Improve a capability without a corpus | No | Yes |
The order to try things
I have a fixed order, and skipping early steps is why projects end up fine-tuning a retrieval problem.
- Prompt properly. A clear system instruction, delimiters for retrieved context, and an explicit instruction to refuse when the context does not answer the question. This alone fixes a surprising share of "the AI is bad" complaints.
- Fix retrieval. Structure-aware chunking, hybrid keyword plus vector search, a reranking pass, metadata filters for permissions.
- Add citations and verify faithfulness. If you can measure whether the answer follows from the retrieved text, you can improve everything else.
- Route between models. Use a cheap model for classification and routing, an expensive one for hard answers. Cost drops, quality holds.
- Fine-tune. Only now, and only for behaviour: output format, classification, tone, refusal.
If you reach step five and are still unhappy, the problem is almost certainly at step two.
Where fine-tuning genuinely wins
There are real cases, and they are narrower than the marketing suggests.
Format compliance at scale: you need every output to parse, you have thousands of examples, and a small fine-tune replaces prompt gymnastics. Domain classification and routing, where a fine-tuned classifier beats a general model plus prompt. Tone that must be exactly right, such as legal or medical register. Behaviour removal — stopping a specific unwanted pattern more reliably than an instruction can.
In all four, fine-tuning changes how the model behaves, and retrieval still supplies the facts. The pattern is not either/or.
The hybrid pattern
The systems that hold up in production look like this: a retrieval layer supplies fresh, permission-filtered, cited context; a model gateway routes the request to the smallest model that can handle it; a fine-tuned layer enforces output shape and refusal behaviour; an evaluation harness proves all of it on every release.
Each piece does the job it is actually good at. Fine-tuning is in there, but it is not carrying the knowledge, and nobody is pretending otherwise.
How to tell which one you need
Three questions, in order.
Does the answer need to be current? Retrieval.
Does the answer need a source a person can check? Retrieval.
Does the output need to look or behave in a specific way? Fine-tuning, prompting, or plain code.
If your honest answer to the first two is no, you may not need either — you may need a rule engine, a normal database query, and considerably less infrastructure. That is a legitimate and underused answer.