I received a scanned [invoice] from a supplier and need line items extracted into our budget sheet today. The scan is …
Open the scanned PDF file in Adobe Acrobat.
Go to the Tools menu.
Select Scan & OCR from the tools list.
Click Recognize Text and choose In This File.
In the settings panel, ensure PDF Output Style is set to Searchable Image.
Click Recognize to start the OCR process.
Save the PDF after OCR is complete.
Best practiceFor best OCR accuracy, scan your documents at 300 dpi and ensure the images are clear before running OCR.
After OCR completes, you can search, copy, and highlight text in your previously scanned PDF.
This runs in the Acrobat app - there is no separate API for this task.
A searchable text layer is an invisible layer of actual text placed over an image-based PDF page. It matters for scanned invoices because it allows you to select, copy, and search for text within the invoice, even though the original document was just a picture. Without it, your invoice would just be an uneditable image.
If your scanned document is blurry, the accuracy of text recognition (OCR) will be lower. Acrobat's OCR tool needs clear text to correctly identify characters. Blurry text can lead to errors, where letters or numbers are misidentified, or some text is not recognized at all. For the best accuracy, always try to scan documents at a high resolution, like 300 dpi, to ensure the text is as sharp as possible before running OCR.
Acrobat shows the 'page contains renderable text' error because the page already has real, selectable text. The OCR tool is designed to convert images of text into selectable text. If the page already has real text, OCR sees no image text to convert. This means you don't need to run OCR; the text is already selectable.
| PDF Type | Text Selectability (Before OCR) | File Content |
|---|---|---|
| From Scanner | No (it's an image) | An image of the document |
| From Word Processor | Yes (it's real text) | Actual text and graphics |