What is a large language model and how is it different from a search engine?
A large language model (LLM) is a statistical model trained to predict likely continuations of text. A search engine retrieves documents that exist. They answer questions in fundamentally different ways, and the difference explains their respective failure modes.
How a search engine works. It crawls and indexes existing pages, then ranks them against a query and returns pointers to sources. Every result corresponds to a real document you can inspect. Its failures are failures of retrieval — the right page may be missing, buried or outranked.
How an LLM works. It is trained on very large quantities of text to predict the next token — roughly a word or word-fragment — given what came before. Through that training it acquires statistical structure about language, facts, reasoning patterns and style. When you ask a question it generates a response token by token.
Crucially, it is not looking anything up. The response is constructed, not retrieved, which has three consequences:
It can produce fluent text that is wrong. Commonly called hallucination — plausible, confident and fabricated. Invented citations, false statistics and non-existent case law are characteristic failures, and fluency is not evidence of accuracy.
It has a training cutoff and does not inherently know recent events.
It cannot always cite sources, because the answer was not taken from a particular document.
Where each is better. Search is better when you need the source itself, current information, or something specific and verifiable. An LLM is better at synthesis, explanation, rephrasing, drafting, summarising supplied text, and questions where the answer is a pattern rather than a document.
The distinction is blurring. Retrieval-augmented generation (RAG) combines them — searching for relevant documents, then having the model answer from those, with citations. This reduces fabrication substantially but does not eliminate it, since the model can still misread or over-extrapolate from what it retrieved.
Verify anything consequential.