I’m responsible for converting a rare diary series into durable digital records for a grant report in [deadline=6 week…
Open your scanned PDF document in Adobe Acrobat.
Go to the Tools menu and select Recognize Text.
In the Recognize Text dialog, choose Recognize Text in This File.
Set the language and other OCR settings as needed.
CautionFor best OCR accuracy, source scans should be at least 300 dpi.
Click Recognize Text to start the OCR process.
To export as plain text, go to File > Export To > Text.
To export as an editable Word document, go to File > Export To > Microsoft Word.
Choose your save location and complete the export.
Best practiceReview the exported text or Word file for formatting or recognition errors, especially with complex layouts or handwriting.
After OCR, you can search, select, and copy text directly from your scanned PDF before exporting.
This runs in the Acrobat app - there is no separate API for this task.
Recognizing text in a scanned PDF converts images of text into actual, editable, and searchable text. This is important because it allows you to find specific words within the document, copy text, and make the document accessible for archiving and indexing purposes.
This error happens because Acrobat's OCR feature is designed to work only on pages that are pure images, without any existing text data. If a page already has renderable (real) text, even if it looks like an image, Acrobat will not try to recognize text on it. To fix this, ensure your document pages are entirely image-based before attempting OCR.
| Feature | Regular PDF (Image-based) | OCR'd PDF (Searchable) |
|---|---|---|
| Text Selectability | Cannot select or copy text | Can select and copy text |
| Searchability | Cannot search for words | Can search for words |
| File Type | Image of a document | Image with a hidden text layer |
| Editing | Limited to image editing | Text can be edited (after conversion) |