OCR PDF
Extract English text from a scanned PDF using optical character recognition that runs on your device.
What OCR PDF does
A scanned document is a picture of text. It cannot be searched or copied until character recognition has read it, which is why a scan behaves so differently from a PDF produced by a word processor.
Each page is rendered and passed to a recognition engine that runs in the page. The engine and its English language data are downloaded from a public CDN the first time you use the tool; after that, the recognition itself happens on your device and your document is never uploaded.
- English text recognition across every page
- Live progress while pages are processed
- Copy the extracted text or download it as a .txt file
How to use OCR PDF
- 1
Add your scanned PDF
A clear scan at a reasonable resolution gives much better results.
- 2
Allow the first-run download
The recognition engine and English data are fetched on first use, which takes a moment and needs a connection.
- 3
Wait for recognition
Pages are processed one at a time. Long documents take a while, because the work happens on your own processor.
- 4
Check the text before using it
Copy or download the result, and proofread it. Recognition is never perfect.
Limits and known behaviour
- English only. No other language data is loaded, so other scripts will produce unusable output.
- Accuracy depends heavily on the scan. Clean, straight, high-contrast pages at 300 DPI read well; photographs of pages, skewed scans, faint print and unusual fonts read badly.
- Handwriting is not recognised.
- Layout is not preserved. The output is plain text, so columns, tables and headers are flattened into a linear stream.
- The output is a text file, not a searchable PDF; the recognised text is not written back into the document.
- The first run downloads roughly 15 MB from a third-party CDN, so this specific tool is not fully offline on first use.
- Password-protected PDFs cannot be opened. The PDF library used here has no decryption support, so an encrypted file is rejected rather than partially processed. Remove the password in a PDF reader first.
Privacy and data handling
Runs in your browser after a one-time download
- Your PDF is not uploaded. It is rendered and recognised in this page.
- The recognition engine, its WebAssembly core and the English training data are downloaded from a public CDN (jsDelivr) on first use. That request tells the CDN your IP address and which files you requested; it does not carry your document.
- Once loaded, recognition runs entirely on your device.
- Neither the document nor the extracted text is stored anywhere after the tab is closed.
Site-wide data handling, including analytics and advertising, is described in the privacy policy.
Frequently asked questions
Why does the first run take so long?
The recognition engine and the English language data have to be downloaded and initialised before any page can be processed. Later documents in the same session reuse them.
How accurate is it?
On a clean 300 DPI scan of ordinary printed text, accuracy is high but not perfect. On a photograph of a page, a faint fax, or an unusual typeface, expect meaningful errors. Always proofread anything you rely on.
Can it read my language?
Only English. Other languages require different training data, which this tool does not load.
Does it produce a searchable PDF?
No. It gives you the recognised text as a separate plain-text file.