What is the difference between training, fine-tuning and retrieval?
Three different ways of getting a model to work with particular information, differing enormously in cost, and frequently confused when people ask about "training it on our data".
Pre-training. Building the base model from scratch on an enormous general corpus. Costs a very large amount in computation, requires specialist infrastructure, and is undertaken by a small number of organisations. This is essentially never what a company means when it talks about training a model on its data.
Fine-tuning. Taking an existing trained model and continuing training on a smaller specific dataset, adjusting the weights.
What it is genuinely good for: teaching a style, format or behaviour — responding in a house tone, producing consistently structured output, following a specialist convention, or handling a task type reliably.
What it is bad for: teaching facts. Knowledge added by fine-tuning is diffuse, cannot be cited, cannot be updated without retraining, and can be recalled incorrectly. Using fine-tuning to add a knowledge base is the most common expensive mistake, and it produces a model that sounds authoritative about things it half-remembers.
It also risks catastrophic forgetting, degrading general capability, and requires curated training data that is more work than people expect.
Retrieval-augmented generation (RAG). The documents stay in a searchable store. At query time, relevant passages are retrieved and inserted into the prompt, and the model answers from the supplied text.
Why it is usually the right answer for factual work:
Updating is instant — change the document, and the next answer reflects it.
Sources can be cited, so answers are checkable.
Access control is possible, since retrieval respects permissions. A fine-tuned model cannot forget what one user should not see.
Far cheaper, with no training required.
Its weaknesses: retrieval quality determines answer quality, so poor chunking or search returns the wrong passages; it consumes context; and the model can still misread what it was given.
The practical rule: RAG for knowledge, fine-tuning for behaviour, and frequently both. Prompt engineering should be exhausted first, since it costs nothing.