We booked the [reading room] and [volunteers] for tomorrow to scan fragile 1870s [diaries] because a course needs imag…
Open the scanned course material PDF in Adobe Acrobat Pro.
Select Tools from the global bar in the upper left.
Scroll down and select Scan & OCR.
Click Recognize Text and choose In This File.
Set the language to match the document's primary language for best accuracy.
RecommendationSelecting the correct language improves OCR results, especially for non-English texts.
Click Recognize Text to start the OCR process.
After OCR completes, select Tools again, then choose Accessibility.
RecommendationRunning the accessibility checker helps ensure that the PDF is usable by screen readers.
Follow the guided steps to address any accessibility issues flagged by Acrobat.
Save the PDF with a new filename to preserve the accessible, searchable version.
Best practiceFor best OCR accuracy, scan documents at 300 dpi and avoid skewed or blurry pages.
After OCR, use the accessibility checker to ensure your scanned readings are fully usable by all students.
This runs in the Acrobat app - there is no separate API for this task.
OCR, or Optical Character Recognition, is a technology that converts images of text into actual text data. It's important for scanned documents because it makes them searchable and editable. Without OCR, a scanned document is just a picture, but with OCR, you can easily find specific words or phrases, making it much more useful for archiving and research.
If your scanned diary pages are blurry, OCR might struggle to recognize the text accurately. Blurry images reduce the clarity of characters, which can lead to errors in the text recognition process. To improve accuracy, try to rescan the pages at a higher resolution, like 300 dpi, and ensure the original document is flat and well-lit during scanning. The clearer the image, the better OCR will perform.
Acrobat shows the 'page contains renderable text' message because the page already has real, selectable text. The OCR feature is designed to work on pages that are only images, turning picture-based text into searchable text. If a page already has text, Acrobat doesn't need to perform OCR. You should check if the text is already searchable before trying to run OCR again.
| Feature | Scanned PDF | OCR'd PDF |
|---|---|---|
| Content | An image of the original document. | An image with a hidden, searchable text layer. |
| Searchability | Not searchable (like a photo). | Fully searchable (you can find words). |
| Editing | Cannot select or copy text directly. | Can select and copy text from the hidden layer. |
| File Size | Can be larger if not optimized. | Slightly larger than image-only but highly functional. |