How does a search engine actually work?
Through three separate processes that people conflate: crawling, indexing and ranking — and a page can succeed at the first two and still never appear, which is the source of most confusion about why something is not found.
Crawling. Automated programs follow links from page to page, discovering content. They respect instructions in robots.txt, allocate limited attention to each site — the crawl budget — and revisit based on how often a page changes. A page nothing links to may never be discovered, which is why sitemaps and internal linking matter.
Indexing. Fetched pages are parsed, rendered — increasingly including running JavaScript, which historically was a major limitation — and analysed. The engine extracts text, structure, links and metadata, and builds an inverted index: a map from each word to every document containing it, which is what makes searching billions of pages feasible in milliseconds. Not everything crawled is indexed; near-duplicate, thin or low-value pages are discarded.
Ranking. When a query arrives, the engine retrieves candidate documents, then scores them using a large number of signals: relevance of the text; link-based authority, the original insight that made modern search work; user behaviour signals; freshness where the query warrants it; location and language; page experience; and increasingly semantic understanding of intent rather than keyword matching, so a page can rank for words it does not contain.
Query processing happens first — interpreting what was meant, correcting spelling, expanding synonyms, and classifying intent as informational, navigational or transactional, which determines what kind of result is appropriate.
Why results differ between people: location, language, device, search history where personalisation applies, and ongoing experiments, since engines continuously run tests on subsets of users.
What is changing. Generated answers synthesised from multiple sources are increasingly placed above links, which alters what traffic reaches sites — and the underlying crawl-index-rank machinery still operates beneath them.
Why a page may not appear: not crawled, not indexed, indexed but outranked, or deliberately excluded.