Question

What is an embedding, and what is a vector database?

Vault Verified
Curated Intelligence
Definitive Source
Answer

An embedding turns a piece of content — text, an image, audio — into a list of numbers positioned so that similar things land near each other. A vector database stores those lists and finds the nearest ones quickly.

Why this is useful. Traditional search matches words. Embedding-based search matches meaning: a query about "how to stop my dog barking" can retrieve a document titled "reducing nuisance vocalisation in canines" despite sharing almost no words, because the two are close in the numerical space.

How embeddings are produced. A neural network trained on enormous quantities of data learns to place content in a space of typically several hundred to a few thousand dimensions, such that related items are near one another. The individual numbers mean nothing interpretable; only the relative positions matter.

What a vector database adds. Comparing a query against millions of vectors exhaustively is too slow, so these systems use approximate nearest neighbour indexes that trade a small amount of accuracy for enormous speed. They also handle filtering by metadata, updates, and scaling.

Where this is used: semantic search; retrieval-augmented generation, where relevant documents are fetched and supplied to a language model so it answers from your content rather than from memory; recommendation; deduplication and near-duplicate detection; image and audio search; and clustering.

What determines quality, in rough order of importance:

The embedding model, and whether it suits your domain and languages.

Chunking. How documents are split before embedding is the most underestimated decision — chunks too large dilute meaning, too small lose context, and splitting mid-idea produces retrieval that is technically correct and useless.

Hybrid search, combining semantic similarity with traditional keyword matching, which consistently outperforms either alone — exact terms, names and identifiers are precisely where embeddings are weak.

Reranking the top candidates with a more expensive model.

The limitations: embeddings encode the biases of their training data; similarity is not relevance; and they cannot tell you whether something is true, only whether it is related.

Related Questions