Optical Character Recognition
Optical Character Recognition (OCR) is a technology by which a software program can analyze text and convert it into a format that can be processed by a computer or other machine. Essentially, OCR technology allows the computer (or other machine) to read.
Many people believe that OCR technology is new, but its earliest origins can be traced back to the 1950s. However, that version of the technology had limitations. The form of OCR with which we are more familiar today was popularized in the 1990s with scanners.
OCR is a must-have to maximize the capabilities of contract management software. Converting all of a company’s contracts into machine-readable form supports a number of important features. For example, once your contracts are digitized, OCR technology provides the following benefits:
- Now that all contracts are digitized and stored together, they can be searched instantaneously using Google-type keyword searches.
- Custom reports can be printed or exported with relevant data located in searches.
- The contracts can be managed from any computer or other device connected to the internet.
- Unlimited users can work with the same contracts at the same time. Management makes the decision on who has access. Moreover, different types of permissions can be granted to different types of users.
Frequently Asked Questions
Why can’t I search text inside my scanned contracts?
A scanned contract is just an image until something reads it. Your computer sees pixels, not words, so keyword search finds nothing. Running the file through OCR converts those pixels into machine-readable text and makes the document searchable. In ContractSafe, uploaded PDFs and scans are processed automatically, so scanned paper contracts turn up in searches alongside everything else you’ve stored.
How accurate is OCR on old scanned documents?
Accuracy depends mostly on the source. Clean, straight scans at reasonable resolution read very well, while faxed copies, faded carbon paper, skewed pages, and handwriting read poorly. Small or unusual fonts also cause errors. If results look rough, rescanning at a higher resolution usually helps more than any software setting. Reviewing key terms by eye is smart when the source is questionable.
Does OCR change the original contract file?
No. OCR adds a text layer alongside the image rather than altering the page you see. The scanned appearance stays intact, signatures and all, so the file still works as your record copy. What changes is that the text underneath becomes selectable and searchable. That’s why a properly OCR’d PDF looks identical to the original but behaves like a text document.