Skip to content

Why PDF compression sometimes does nothing

What actually occupies space in a PDF, which files can shrink, and the trade-off you accept when pages are rasterised.

Last reviewed 8 September 2026

Run a 40 MB scanned contract through a compressor and it may come out at 4 MB. Run a 2 MB report from a word processor through the same tool and nothing happens. Both results are correct, and the reason is that the two files are large for entirely different reasons.

What is actually inside a PDF

A PDF is a container of objects: text with font references, vector drawing instructions, embedded images, fonts, metadata, and — in interactive documents — form fields and scripts. Their contributions to file size are wildly unequal.

  • Text is tiny. A page of prose is a few kilobytes of glyph positions.
  • Vector graphics are small, being drawing commands rather than pixels.
  • Embedded fonts are moderate, typically tens to hundreds of kilobytes each, and are paid once per document.
  • Images dominate everything. A single 300 DPI scan of an A4 page can be several megabytes on its own.

So the useful question is not “how do I compress this PDF” but “what is taking up the space”. If the answer is images, a lot can be done. If the answer is text, very little can.

Why a text PDF will not shrink

PDF already compresses its own content streams internally, usually with the same algorithm behind ZIP. A document generated cleanly by a word processor arrives with almost no redundancy left in it.

A structural pass can still help in specific cases — removing orphaned objects, packing the cross-reference table more efficiently, discarding metadata left behind by an editor — but on a well-formed 2 MB report that might be a few percent. Occasionally the rewritten file is actually slightly larger. That is not a failure of the tool; there was simply nothing to remove.

Why a scan shrinks dramatically

A scanned page is a photograph. Scanners default to high resolution and conservative compression because they cannot know what you need, so a twenty-page scan can carry twenty full-resolution images sized for printing when it will only ever be read on screen.

Reducing that resolution, and re-encoding at a sensible JPEG quality, is where large savings come from. A 300 DPI page reduced to 150 DPI has a quarter of the pixels before any quality setting is applied, because resolution affects both dimensions.

The trade-off nobody mentions

Aggressive PDF compressors work by rasterising: each page is rendered to an image and the image becomes the page. That produces impressive reductions on any document and destroys something on documents that had real text.

After rasterising:

  • Text can no longer be selected or copied.
  • The document is no longer searchable.
  • Links and form fields stop working.
  • Screen readers cannot read it, which for many organisations is an accessibility failure rather than an inconvenience.
  • The file cannot be edited meaningfully afterwards.

For a scan, none of that is a loss — there was no text layer to begin with. For a text document, it is a significant one. This is why the PDF compressor here separates the two: a structural mode that changes nothing visible, and rasterising modes that state what they remove before you run them.

Reading the numbers honestly

A compressor that always reports a saving is not measuring; it is reassuring. If a rewritten file comes out larger, the honest response is to say so and offer the original, which is what happens here. Treat a tool that never reports a failure with suspicion.

What to do in practice

  1. Identify the content. Try selecting text in the document. If you can, it has a text layer and structural compression is the only safe option. If you cannot, it is a scan and rasterising costs you nothing.
  2. Compress the images before assembling, where you can. If you are building a PDF from photographs, compress them first with the image compressor and then combine them with Image to PDF. That gives far better results than compressing the finished document.
  3. Scan at a sensible resolution. 300 DPI is for printing and for text recognition. 150 DPI is ample for reading on screen and is a quarter of the data.
  4. Split instead of compressing. If you only need to send four pages, send four pages. The split tool often solves an attachment limit better than any compressor.
  5. Check the result. Open the output and read the smallest text on the busiest page before you discard the original.

When compression is the wrong answer entirely

If a file is too large for email, consider whether it should be an attachment at all. Compressing a 60 MB document to 18 MB still exceeds most limits, and a shared link avoids the problem without degrading anything. And if the document must remain accessible — a public report, anything covered by accessibility requirements — rasterising it to hit a size target trades one obligation for another.

Tools referred to in this guide