Work with [IT] to configure storage and searchable retention for born‑digital fieldwork files, ensure completed docume…
Open the scanned PDF file in Adobe Acrobat.
Go to the Tools menu.
Select Scan & OCR from the tools list.
Click Recognize Text and choose In This File.
In the settings panel, ensure PDF Output Style is set to Searchable Image.
Click Recognize to start the OCR process.
Save the PDF after OCR is complete.
Best practiceFor best OCR accuracy, scan your documents at 300 dpi and ensure the images are clear before running OCR.
After OCR completes, you can search, copy, and highlight text in your previously scanned PDF.
This runs in the Acrobat app - there is no separate API for this task.
OCR, or Optical Character Recognition, is technology that converts images of text into actual, editable, and searchable text. It is important for scanned documents because most scanners create image-only files. Without OCR, you cannot search for words, copy text, or edit the document. OCR makes your scanned files useful for archiving and data extraction.
No, if your scanned PDF already contains 'renderable text,' Acrobat's OCR feature will not run on those specific pages. OCR is designed for image-only pages. If you get an error message, it means the page probably already has text. You only need to apply OCR to pages that are purely images.
Choosing the correct language is very important for OCR accuracy because different languages have unique character sets, accents, and grammar rules. If you select the wrong language, Acrobat's OCR engine might misinterpret words or characters, leading to many errors in the recognized text. Always set the language to match the document's content.
For the best OCR accuracy, you should scan your documents at 300 DPI (Dots Per Inch). While 72 DPI is the lowest acceptable setting, a higher resolution like 300 DPI provides much clearer images, which helps the OCR software recognize text more precisely.