Why is my scanned PDF so large?
Because a scan is not a document — it is a collection of photographs, and it is stored as image data rather than as text.
The difference. A PDF created from a word processor stores text characters plus font and layout instructions, which is extremely compact — a long document may be a few hundred kilobytes. A scanned PDF stores a picture of each page, and the file size depends on pixel count and colour depth, not on how much is written.
What drives the size:
Resolution (DPI). Size scales roughly with the square of resolution. Scanning at 600 DPI produces a file around four times larger than 300 DPI. For text documents, 300 DPI is sufficient for both readability and OCR accuracy, and higher adds size without benefit. 600 DPI is worth it only for very small print or archival photographs.
Colour mode, which is the biggest single lever. A page scanned in 24-bit colour is roughly three times the data of greyscale and dramatically more than bitonal (pure black and white). Scanning plain printed text in full colour is the most common cause of an oversized file.
Compression choice. JPEG compression suits photographs and produces artefacts around text. Group 4 / CCITT compression is designed for bitonal text and is extremely efficient. JBIG2 and MRC (mixed raster content) separate text from background and compress each appropriately, achieving very large reductions — this is what document scanners and "reduce file size" features use.
No downsampling, leaving the scan at capture resolution.
Embedded thumbnails and metadata, a smaller contributor.
What to do:
Scan text at 300 DPI in black and white or greyscale. This alone often reduces a file by 90%.
Run OCR, which adds a searchable text layer — usually small — and makes the document genuinely useful rather than an image.
Use a PDF optimiser to downsample and recompress existing files.