Suitable only for clear, large print
Small characters and thin strokes can disappear. Use it when the source is exceptionally clean and compact size matters more than recognition reliability.
Convert PDF tool
Make scanned PDF pages searchable and selectable with a language-aware text layer.
Drag & drop or select a file
30 MB per PDFPlease keep this page open.
Optical character recognition
A scanned PDF usually stores each page as an image, so a PDF reader cannot understand the words inside it. OCR analyzes those pixels and adds a searchable-text layer behind the visible page. A born-digital or already searchable PDF already contains text objects and normally does not need OCR.
Available recognition models
Choose the primary language printed in the document. Mixed-language pages, names, codes and specialist vocabulary still require review.
Scan preparation
Small characters and thin strokes can disappear. Use it when the source is exceptionally clean and compact size matters more than recognition reliability.
A practical starting point for ordinary printed text. It usually retains useful character detail without the processing and storage cost of very high-resolution scans.
May help with very small type or fine originals, but it increases file size and processing time. It cannot recover detail already lost to blur, glare or compression.
Recognition factors
Output behavior
OCRmyPDF and Tesseract process the document on the server. The workflow keeps the page image as the visible source and adds recognized text for searching and selection. Small page-skew correction can alter page-image geometry when enabled.
Search results and copied text depend on the recognized text layer—not solely on what the page looks like. Complex reading order, columns, tables, symbols and unusual fonts can therefore copy differently from their visual arrangement.
Practical limits
OCR cannot reconstruct characters missing from an out-of-focus scan, remove heavy glare, infer obscured words or promise correct handwriting. A 30 MB per-file upload limit applies, and complex pages can also encounter processing-time or server resource limits.
Troubleshooting
Run the untouched scan again and choose its primary printed language. A single selection cannot fully model arbitrary mixed-language content.
Check scan sharpness, contrast, compression artifacts and the selected language. Rescanning is more effective than repeatedly processing a poor source.
Try selecting text before processing. Existing text pages are skipped, so OCR is most useful for image-only pages.
Return to the paper source and capture it again with steady focus, even light and the camera parallel to the page.
Use Rotate PDF for quarter-turn correction first. Enable Straighten only for small scan-skew angles.
The recognized words may be present while reading order is imperfect. Verify visually and use a specialist extraction workflow for structured tables.
File handling
OCR requires server processing because the recognition engines run on the configured VPS worker. Generated files become eligible for cleanup after 15 minutes and, while scheduled cleanup is healthy, are normally removed by a later run. Avoid uploading highly sensitive material unless this processing model meets your requirements.
OCR questions
OCR analyzes page images and adds an invisible searchable-text layer. The page image remains the visible document while compatible PDF readers can search, select and copy recognized text.
This tool currently supports English, German, French, Spanish, Italian, Portuguese and Dutch. Choose the main language used in the document.
300 DPI is a practical starting point for most printed documents. A 150 DPI scan may lose character detail, while 600 DPI increases processing and file size without guaranteeing better recognition.
The visible page is generally preserved because OCR adds a text layer rather than rebuilding the document. Reading order, copied text, columns and tables may not reproduce the visual layout perfectly.
This workflow uses a printed-text OCR engine. Neat handwriting may produce partial results, but handwriting recognition is not a supported or reliable use case.
The OCR engine must be able to open the document. If you know the password and have permission, create an authorized unlocked copy before running OCR.
No. Recognition quality varies by scan, typeface, language and layout, and this service has not published a controlled accuracy benchmark. Always verify important text against the page image.
Related workflows
Continue your workflow
Use your document in a closely connected PDF task.
Browse all PDF tools →Useful reading
Learn what affects this tool’s output and how to verify it.