1Upload the scanned PDF documents; a clean 300 DPI grayscale scan gives the recogniser the most to work with.
2Tell it which language to expect — naming the language beats letting it guess by a wide margin.
3Run the recognition pass; the words are written back as an invisible layer sitting over the original page image.
4Download the searchable PDF file — it looks identical, but you can now search and select the text in it.
OCR PDF FAQ
What should I scan at for the best result?
+
300 DPI in grayscale is the sweet spot. Below 200 DPI character shapes start breaking down; above 400 DPI you gain almost nothing and pay for it in file size and processing time.
How does OCR PDF turn a scan into text?
+
Concretely, each page is rendered, run through a Tesseract text-recognition pass in over 100 languages, and the recognised words are written back as an invisible text layer positioned over the original image — so the page looks unchanged but is searchable and selectable. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
Anything specific to PDF here?
+
Yes — a PDF is an object graph, not a page image, so text stays selectable and vectors stay sharp no matter what happens to the raster content inside it. It affects what the recogniser can see.
How accurate is it?
+
On a clean 300 DPI scan of printed text, high enough that errors are rare and mostly in unusual proper nouns. Accuracy falls off sharply with low resolution, skew, shadows from a phone camera, unusual typefaces and handwriting — handwriting in particular is not what this is for.
What can I upload to OCR PDF?
+
Any standard PDF, including ones produced by a scanner, by a word processor or by a print driver. Files with an owner password that only restricts editing are handled; files that need a user password to open are not, and we do not attempt to break encryption.
Is there a file size limit on OCR PDF?
+
Yes: free accounts process documents up to 25 MB each, which covers most reports and contracts but not a long scanned document at 600 DPI; Ghostscript and qpdf do the work. Scanned PDFs are the usual thing that exceeds it, and compressing the document first is normally enough to bring it back under.
Will OCR PDF lower the quality of my PDF documents?
+
Text and vector artwork are objects rather than pixels, so they stay perfectly sharp at any zoom no matter what happens. Only the embedded raster images can degrade, and only if the operation you chose resamples them.
Can I run OCR PDF on several PDF documents at once?
+
Yes — upload the set and they process in parallel under one set of settings, which is the point of doing a document workflow here rather than clicking through a desktop reader.
Why does a EPUB site host OCR PDF?
+
EPUB.to is built around the open eBook format — XHTML in a zip, reflowable by design, and readable on hardware nobody has thought about the layout for. A book is a zip full of markup, images and fonts, so almost every job people bring to an eBook site is really a job on one of those things. OCR PDF runs on the same upload and the same account as the conversions because that is where it is needed.
What should I do with the result once OCR PDF is finished?
+
The converter on this site moves books between EPUB, MOBI, AZW3, PDF and DOCX, so a title ends up on whatever hardware is actually going to read it. Doing that after OCR PDF means the conversion is made from the version you settled on, not from the one you were still fixing.
Is OCR PDF here the same tool the sibling sites run?
+
The engines are shared — the same toolchain, the same workers, the same limits. What a EPUB site adds is a view on reflow: which of these operations a reflowable book genuinely supports, and which ones only make sense on the fixed-layout media inside it. It also starts from one fact about the format this site is named after: the book is a zip of XHTML with a manifest, so the text costs almost nothing and every decision that matters is about the embedded images and font subsets.
Do I need an account, and does anything get kept?
+
No account, and nothing is kept: uploads are deleted from the workers shortly after the job finishes, nothing is read and nothing is indexed. Free accounts exist for history and batch size, not for access.