Convert PDF tool

OCR PDF Online

Make scanned PDF pages searchable and selectable with a language-aware text layer.

Step 1 · Upload

Upload and prepare

Add your file

Drag & drop or select a file

30 MB per PDF

Optical character recognition

Turn page images into searchable documents

A scanned PDF usually stores each page as an image, so a PDF reader cannot understand the words inside it. OCR analyzes those pixels and adds a searchable-text layer behind the visible page. A born-digital or already searchable PDF already contains text objects and normally does not need OCR.

Available recognition models

Supported OCR languages

Choose the primary language printed in the document. Mixed-language pages, names, codes and specialist vocabulary still require review.

  • English
  • German
  • French
  • Spanish
  • Italian
  • Portuguese
  • Dutch

Scan preparation

Resolution matters, but more is not always better

150 DPI

Suitable only for clear, large print

Small characters and thin strokes can disappear. Use it when the source is exceptionally clean and compact size matters more than recognition reliability.

600 DPI

For demanding small print

May help with very small type or fine originals, but it increases file size and processing time. It cannot recover detail already lost to blur, glare or compression.

Recognition factors

What affects printed-text recognition

Typeface and print
Clean, conventional fonts with distinct character shapes are easier than decorative, condensed or damaged print. Verify names, totals and reference numbers manually.
Columns
Text may be recognized correctly but copied in the wrong reading order. Review multi-column pages after processing.
Tables
Cells, borders and visual alignment are not a guarantee of structured table extraction. Copied rows can lose their column relationships.
Rotation and skew
The Straighten option corrects small skew angles; it is not a substitute for rotating sideways pages. Use Rotate PDF before OCR when a page is turned 90° or 180°.
Contrast and lighting
Faint ink, colored backgrounds, shadows and uneven lighting reduce separation between characters and the page. Rescan with even light when possible.
Blur and handwriting
Motion blur merges character edges and cannot be reliably reversed. Handwriting is outside this printed-text OCR workflow and should be treated as unsupported.

Output behavior

Visible layout and searchable-text layer

OCRmyPDF and Tesseract process the document on the server. The workflow keeps the page image as the visible source and adds recognized text for searching and selection. Small page-skew correction can alter page-image geometry when enabled.

Search results and copied text depend on the recognized text layer—not solely on what the page looks like. Complex reading order, columns, tables, symbols and unusual fonts can therefore copy differently from their visual arrangement.

Practical limits

Documents OCR cannot reliably repair

OCR cannot reconstruct characters missing from an out-of-focus scan, remove heavy glare, infer obscured words or promise correct handwriting. A 30 MB per-file upload limit applies, and complex pages can also encounter processing-time or server resource limits.

Troubleshooting

When the searchable text is not right

OCR detects the wrong language

Run the untouched scan again and choose its primary printed language. A single selection cannot fully model arbitrary mixed-language content.

Text contains garbled characters

Check scan sharpness, contrast, compression artifacts and the selected language. Rescanning is more effective than repeatedly processing a poor source.

The PDF is already searchable

Try selecting text before processing. Existing text pages are skipped, so OCR is most useful for image-only pages.

The scan is too blurry

Return to the paper source and capture it again with steady focus, even light and the camera parallel to the page.

Pages are rotated

Use Rotate PDF for quarter-turn correction first. Enable Straighten only for small scan-skew angles.

Columns or tables copy badly

The recognized words may be present while reading order is imperfect. Verify visually and use a specialist extraction workflow for structured tables.

File handling

Privacy and temporary processing

OCR requires server processing because the recognition engines run on the configured VPS worker. Generated files become eligible for cleanup after 15 minutes and, while scheduled cleanup is healthy, are normally removed by a later run. Avoid uploading highly sensitive material unless this processing model meets your requirements.

Read the File Deletion Policy Review the Privacy Policy

OCR questions

Frequently asked questions

What does OCR add to a scanned PDF?

OCR analyzes page images and adds an invisible searchable-text layer. The page image remains the visible document while compatible PDF readers can search, select and copy recognized text.

Which OCR languages are supported?

This tool currently supports English, German, French, Spanish, Italian, Portuguese and Dutch. Choose the main language used in the document.

What scan resolution should I use?

300 DPI is a practical starting point for most printed documents. A 150 DPI scan may lose character detail, while 600 DPI increases processing and file size without guaranteeing better recognition.

Will OCR preserve the original layout?

The visible page is generally preserved because OCR adds a text layer rather than rebuilding the document. Reading order, copied text, columns and tables may not reproduce the visual layout perfectly.

Can OCR recognize handwriting?

This workflow uses a printed-text OCR engine. Neat handwriting may produce partial results, but handwriting recognition is not a supported or reliable use case.

Can I process a password-protected PDF?

The OCR engine must be able to open the document. If you know the password and have permission, create an authorized unlocked copy before running OCR.

Does this page display an OCR accuracy percentage?

No. Recognition quality varies by scan, typeface, language and layout, and this service has not published a controlled accuracy benchmark. Always verify important text against the page image.