Looking at our backlog of scanned forms and OCR logs, show the recurring error types (e.g., decimal misreads, column s…
Open the scanned PDF in Adobe Acrobat.
Check if the document is an image-only PDF by attempting to select or search for text.
NoteOCR cannot recognize text if the PDF does not contain images or if the text is already selectable.
Go to Tools and select Recognize Text.
Click Recognize Text and choose In This File.
Review the OCR output for errors or missed text by scrolling through the document.
If OCR accuracy is poor, check the scan quality: ensure the document is at least 300 dpi, not skewed, and has high contrast.
Best practiceLow-resolution, blurry, or skewed scans are the most common causes of OCR failure.
Rescan the original document at 300 dpi or higher, using RGB mode for discolored or older pages.
Best practiceFor best results, avoid scanning with excessive brightness and ensure pages are flat and unmarked.
Repeat the OCR process on the improved scan.
Manually correct any remaining OCR errors by clicking on suspect words to edit and accept corrections directly.
NoteYou can click on suspect words to edit and accept corrections directly.
Save the corrected PDF.
If OCR still fails, try rescanning your document at a higher resolution and check for skewed or blurry pages before running OCR again.
This runs in the Acrobat app - there is no separate API for this task.
Acrobat's OCR feature is designed for pages that are only images. If a page already contains actual text, not just a picture of text, Acrobat will not perform OCR. You might see an error message like 'page contains renderable text'. This means the page already has text that can be selected and copied, so OCR is not needed for it.
For the most accurate OCR results on paper forms, you should scan your documents at 300 dpi (dots per inch). Scanning at a higher resolution helps Acrobat recognize text more clearly, reducing the number of errors. While 72 dpi is the lowest acceptable quality, 300 dpi gives much better accuracy.
If your scanned document is in a language other than English, it's very important to select the correct OCR language before you start the recognition process. In the 'Scan & OCR' tool, you will find an option to set the language. Choosing the right language helps Acrobat's OCR engine understand the specific characters and grammar rules, leading to much more accurate text recognition.
| Feature | Scanned Image of Text | Searchable Text Layer |
|---|---|---|
| Nature | A picture of words, like a photo. | Actual, editable text on top of the image. |
| Select/Copy | Cannot select or copy words. | Can select, copy, and edit words. |
| Search | Cannot search for words. | Can search for any word in the document. |
| File Size | Often larger, as it's an image. | Slightly larger than image, but adds functionality. |