1Upload the scanned PDF documents; a clean 300 DPI grayscale scan gives the recogniser the most to work with.
2Tell it which language to expect — naming the language beats letting it guess by a wide margin.
3Run the recognition pass; the words are written back as an invisible layer sitting over the original page image.
4Download the searchable PDF file — it looks identical, but you can now search and select the text in it.
OCR PDF FAQ
Anything specific to PDF here?
+
Yes — a PDF is an object graph, not a page image, so text stays selectable and vectors stay sharp no matter what happens to the raster content inside it. It affects what the recogniser can see.
What should I scan at for the best result?
+
300 DPI in grayscale is the sweet spot. Below 200 DPI character shapes start breaking down; above 400 DPI you gain almost nothing and pay for it in file size and processing time.
How does OCR PDF turn a scan into text?
+
Concretely, each page is rendered, run through a Tesseract text-recognition pass in over 100 languages, and the recognised words are written back as an invisible text layer positioned over the original image — so the page looks unchanged but is searchable and selectable. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
Which languages can it recognise?
+
Over 100, including Latin, Cyrillic, Greek, Arabic, Hebrew, Chinese, Japanese, Korean and the Indic scripts. Telling it which language to expect materially improves accuracy over letting it guess.
Does OCR PDF cost anything?
+
No. OCR PDF is free without an account, and nothing is stamped onto the pages. Uploaded documents are deleted from the workers shortly after the job completes.
What can I upload to OCR PDF?
+
Any standard PDF, including ones produced by a scanner, by a word processor or by a print driver. Files with an owner password that only restricts editing are handled; files that need a user password to open are not, and we do not attempt to break encryption.
Is there a file size limit on OCR PDF?
+
Yes: free accounts process documents up to 25 MB each, which covers most reports and contracts but not a long scanned document at 600 DPI; Ghostscript and qpdf do the work. Scanned PDFs are the usual thing that exceeds it, and compressing the document first is normally enough to bring it back under.
Do bookmarks, links and form fields survive?
+
Outlines, internal links and annotations are preserved wherever the operation allows it. Digital signatures are the exception: any change to the file necessarily invalidates a signature, because that is precisely what a signature is for.
Why does a MP3 site host OCR PDF?
+
MP3.to is built around the format that made portable audio ordinary: universally playable, small enough to ignore, and lossy in a way that only matters if you keep re-encoding it. Audio jobs cluster: the same person who needs a format changed usually needs the file cut, levelled or made smaller in the same sitting. Putting OCR PDF behind the same upload, the same caps and the same account is what turns three tabs into one.
What should I do with the result once OCR PDF is finished?
+
The converter on this site moves audio between MP3, WAV, FLAC, M4A, OGG and Opus, so the output format is a decision you can leave until the end rather than commit to at the start. Doing that after OCR PDF rather than before means the encode happens once, on the file you actually settled on, instead of twice.
Is OCR PDF here the same tool the sibling sites run?
+
The engines are shared — the same ffmpeg build, the same workers, the same limits. What a MP3 site adds is a point of view about bitrates, about what survives an encode and about when a lossless intermediate is worth the upload time. It also starts from one fact about the format this site is named after: MP3 frames are fixed-size blocks, so a cut lands on a frame boundary and gapless playback markers do not survive a re-encode.
Do I need an account, and does anything get kept?
+
No account, and nothing is kept: uploads are deleted from the workers shortly after the job finishes, nothing is listened to and nothing is indexed. Free accounts exist for history and batch size, not for access.