Optical Character Recognition is required when sensitive information exists as pixels rather than directly extractable text. A scanned document or image-based PDF may visibly contain names, account numbers, identification numbers, or other protected information, but normal detection rules cannot evaluate those characters until they have been converted into machine-readable text. OCR performs that conversion and passes the extracted text to the DLP detection process, where keywords, regular expressions, Data Identifiers, or other policy rules can inspect it. Keyword and regular-expression matching may operate after OCR processing, but neither technique performs the image-to-text conversion. Exact Data Matching also requires readable and normalized values before comparing content with an index. Therefore, Optical Character Recognition, option B, is the correct detection method.
================
Contribute your Thoughts:
Chosen Answer:
This is a voting comment (?). You can switch to a simple comment. It is better to Upvote an existing comment if you don't have anything to add.
Submit