← Guides

OCR and scanned documents

Optical character recognition is pattern matching over pixels. It has no understanding of what a document means, so its accuracy is dominated by things you control before recognition starts: resolution, contrast, straightness and layout complexity.

These guides explain what OCR does to a file, why the same engine can give a perfect result on one page and nonsense on the next, and how to prepare scans so recognition has a fair chance.