Over the last [months=6] months, quantify when scanned [pages] were judged illegible and why: scanner settings, paper …
Open the scanned PDF in Adobe Acrobat.
Check if the document is an image-only PDF by attempting to select or search for text.
NoteOCR cannot recognize text if the PDF does not contain images or if the text is already selectable.
Go to Tools and select Recognize Text.
Click Recognize Text and choose In This File.
Review the OCR output for errors or missed text by scrolling through the document.
If OCR accuracy is poor, check the scan quality: ensure the document is at least 300 dpi, not skewed, and has high contrast.
Best practiceLow-resolution, blurry, or skewed scans are the most common causes of OCR failure.
Rescan the original document at 300 dpi or higher, using RGB mode for discolored or older pages.
Best practiceFor best results, avoid scanning with excessive brightness and ensure pages are flat and unmarked.
Repeat the OCR process on the improved scan.
Manually correct any remaining OCR errors by clicking on suspect words to edit and accept corrections directly.
NoteYou can click on suspect words to edit and accept corrections directly.
Save the corrected PDF.
If OCR still fails, try rescanning your document at a higher resolution and check for skewed or blurry pages before running OCR again.
This runs in the Acrobat app - there is no separate API for this task.
Recognizing text in scanned PDFs is important because it makes the document searchable. Without OCR, a scanned document is just an image, and you cannot select text, copy it, or find specific words. Making it searchable helps with archiving, indexing, and quickly finding information, saving you time and effort.
If Acrobat says 'page contains renderable text', it means the page already has real text, not just an image. OCR is only for image-only pages. You don't need to run OCR on this document because the text is already recognized and searchable.
OCR accuracy changes a lot with scan resolution. Scanning at 300 dpi gives the best accuracy because it captures more detail, making it easier for Acrobat to recognize characters. Scanning at lower resolutions, like 72 dpi, can lead to more errors and less accurate text recognition, making the document less reliable for searching or copying.
Choosing the correct language before recognizing text is very important because OCR software uses language-specific rules and dictionaries to identify characters and words. If you select the wrong language, Acrobat might misinterpret characters, leading to many errors in the recognized text. This makes the document less accurate and harder to search.