Question

How does OCR work?

Vault Verified
Curated Intelligence
Definitive Source
Answer

By locating text in an image and then recognising it — two separate problems, and the second is far easier than the first, which is why scanning a clean document works well and photographing a sign at an angle does not.

The traditional pipeline:

Preprocessing — converting to greyscale, removing noise, correcting contrast, and deskewing so lines are horizontal, which matters enormously.

Binarisation, separating text from background. Adaptive methods handle uneven lighting, which is the usual failure point on photographs.

Layout analysis, identifying blocks, columns, tables and reading order. This is where most real-world OCR fails — recognising the characters in a two-column page is easy; knowing which order to read them in is not.

Segmentation into lines, words and characters — difficult with joined, overlapping or damaged text.

Recognition, matching shapes against learned models.

Post-processing with a dictionary and language model to correct implausible results.

What changed with neural networks. Modern systems recognise whole lines at once rather than segmenting characters, which handles joined handwriting and unusual fonts far better, and end-to-end models increasingly do detection and recognition together — which is what makes live translation through a phone camera possible.

Where it still struggles: handwriting, particularly cursive and historical hands; low-resolution or compressed images, where JPEG artefacts destroy fine detail; tables and forms, where structure carries meaning; multi-column and complex layouts; degraded historical documents; unusual fonts and stylised type; and scripts underrepresented in training data.

The distinction worth knowing. A scanned PDF is an image — it looks like a document and contains no text. Running OCR adds a searchable text layer beneath the image, which is what makes it findable. Confusing the two is why documents cannot be searched and why accessibility tools cannot read them.

Practical points: scan at adequate resolution, typically 300 dots per inch for text; keep pages flat and square; prefer scanning to photographing; and always verify critical figures, since a misread digit in a number carries no linguistic context to correct it.

Related Questions