A scan may look like a normal PDF while containing only page images. Optical character recognition (OCR) detects characters and places a text layer behind the image, enabling search, selection and better accessibility.
OCR is recognition, not proofreading. Names, numbers and low-contrast text deserve manual review.
Capture a clean source
Use even lighting, keep the camera parallel to the page, fill the frame, and avoid shadows over text. A sharp, straight image produces better recognition than software correction applied to a blurred photo.
- Use the correct document language.
- Deskew and rotate pages before OCR.
- Prefer adequate resolution over extreme JPEG compression.
Run OCR and test it
Upload the scan to OCR PDF, select the appropriate language where available, and process it. Download the output, search for a distinctive sentence, then copy several paragraphs into a plain-text editor to reveal recognition errors.
Understand the limits
Handwriting, decorative fonts, tables, faint carbon copies and complex multi-column layouts are harder to recognize. For legal, medical or financial records, compare critical fields with the page image and do not treat OCR text as authoritative without review.