Convert submitted [charts=GIS chart PDFs] to searchable Excel tables, run OCR where needed, verify density and parcel …
Open the scanned PDF file in Adobe Acrobat.
Go to the Tools menu.
Select Scan & OCR from the tools list.
Click Recognize Text and choose In This File.
In the settings panel, ensure PDF Output Style is set to Searchable Image.
Click Recognize to start the OCR process.
Save the PDF after OCR is complete.
Best practiceFor best OCR accuracy, scan your documents at 300 dpi and ensure the images are clear before running OCR.
After OCR completes, you can search, copy, and highlight text in your previously scanned PDF.
This runs in the Acrobat app - there is no separate API for this task.
Renderable text means the text is already real text that you can select and copy in your PDF. Acrobat's OCR feature is designed to work only on images where the text is just a picture. If a page already has renderable text, Acrobat thinks there's no image text to convert, so it stops the OCR process. This prevents it from creating a duplicate text layer or trying to process something that is not an image.
If your scanned density figures are blurry, Acrobat's OCR might struggle to recognize the text accurately. Blurry images make it hard for the software to identify letters and numbers. While it might still create a searchable layer, the quality of the search results could be poor. For better results, try to rescan the figures at a higher quality or use image editing tools to improve clarity before running OCR.
Selecting the correct language is very important because OCR software uses language-specific rules and dictionaries to recognize characters. For example, the letter 'ñ' in Spanish is different from 'n' in English. If you choose the wrong language, Acrobat might misinterpret words, leading to many errors in the searchable text layer. This means your searches might not find what you expect.
| Feature | Regular PDF (image-only) | Searchable PDF (OCR) |
|---|---|---|
| Text Selection | Cannot select text directly from images. | Can select and copy text from images. |
| Search Function | Cannot search for words within the document. | Can search for any word in the document. |
| File Origin | Often from scanned documents or photos. | Created by adding a hidden text layer to an image-only PDF. |