Skip to main content
PDFCraft
Back to blog
Guides

Make a scanned PDF searchable with OCR

Turn photographed or scanned pages into text you can search, copy and convert — and get the best accuracy from Tesseract.

By PDFCraft Team2 min read

A scanned PDF looks like a document but behaves like a photograph: you cannot search it, select a sentence, or convert it to Word. Optical character recognition (OCR) fixes that by reading the pixels and writing an invisible text layer over each page. The scan still looks exactly the same; underneath it, every word becomes real text.

What OCR PDF does

PDFCraft's OCR tool is built on OCRmyPDF and Tesseract, the open-source engine used by most document management systems. It accepts a PDF or a single JPG/PNG image and produces a PDF with a hidden text layer and, if you enable it, a plain-text file with the recognised text that you can download alongside it. Pages that already contain text are skipped by default so existing digital pages are never damaged.

Because Tesseract is a native program, OCR runs on our servers. Your file is processed in an isolated job, the result is stored only until you download it, and both are deleted automatically after the retention period.

Step by step

  1. Open OCR PDF and upload the scan. Images are rasterised at 300 dpi, which is the sweet spot for recognition.
  2. Select the languages present in the document — English, French and Arabic are always installed; German, Spanish, Italian, Portuguese, Dutch, Turkish and Russian can be enabled by the operator. You can pick up to four for mixed documents; the first one should be the dominant language.
  3. Tick Deskew if the pages were scanned slightly crooked. Straightening before recognition noticeably improves accuracy.
  4. Leave Force OCR off unless the file already has a broken or garbage text layer (common with faxes and old scanner software), in which case the existing text is discarded and every page is re-recognised.
  5. Start the job. You can follow its progress on the page; a 30-page scan typically takes a minute or two depending on server load.

Getting better accuracy

  • Resolution matters most. 300 dpi is ideal; 200 dpi is usable; anything from a phone camera at arm's length is a gamble. Rescan rather than upscale.
  • Contrast beats colour. Grey or black-and-white scans recognise as well as colour ones and are much smaller.
  • Choose the right languages. Adding a language the document does not contain lowers accuracy, because the engine considers more candidate words.
  • Deskew crooked pages. Even a two-degree tilt hurts line segmentation.
  • Expect imperfect results on handwriting, decorative fonts and tables. OCR is built for printed text; review anything critical.

After OCR

Once the text layer exists, the rest of the toolkit works on the scan: PDF to Text extracts the recognised text, PDF to Word rebuilds editable paragraphs, and the browser's find function finally works. Screen readers can also read the document, which makes OCR an accessibility fix as much as a convenience.