Erasing information on a scanned document: why a white rectangle is not enough
Before sending a copy of a document, we often have a piece of information to remove: an account number on a bank details form, an address, a signature, a social security number. On a scanned document, the temptation is to place a white or black rectangle over it with PDF software. On screen, the information disappears. In the file, it is still there. We checked this on a scanned letter, and compared it with the real erasing that FileXvert's PDF editor now offers.
A scan contains no text, only an image
A PDF created by a word processor describes its letters as characters. A PDF from a scanner or a scanning app contains something else entirely: a photograph of the page, stored as an image. The letters you read are dark pixels on a light background.
Placing a rectangle on that page simply adds a drawing on top of the image. The image itself is not modified: it remains embedded in the file, complete. Any image extraction tool, many of them free, pulls it out in seconds, minus the rectangle.
Our test
We created a fictitious bank letter, scanned the way an office scanner would: an A4 page at 200 dots per inch (1654 × 2339 pixels), with slight grain and a greyish background. It carries a fictitious “IBAN” line. We saved it in the two most common forms: a JPEG image, as most scanning apps produce, and a black-and-white image compressed with CCITT Group 4, as many scanners set for documents produce.
On each version, we hid the IBAN line in two ways: by drawing a rectangle in the background colour, and with FileXvert's “Erase an area” tool. We then extracted the image contained in each file and counted the dark pixels, the ink, in the band of the IBAN line.
| Version | PDF size | IBAN ink pixels in the file |
|---|---|---|
| JPEG scan, original | 131.6 KB | 9,973 |
| JPEG scan + rectangle on top | 132.2 KB | 9,973 (unchanged) |
| JPEG scan + “Erase an area” | 94.9 KB | 0 |
| Black-and-white CCITT scan, original | 7.5 KB | 10,395 |
| CCITT scan + rectangle on top | 8.1 KB | 10,395 (unchanged) |
| CCITT scan + “Erase an area” | 12.3 KB | 0 |
With the rectangle, the image in the file contains exactly as much ink as before: the IBAN is intact, merely covered. With FileXvert's erasing, not a single ink pixel remains in the area.
What “Erase an area” does
You draw the area to erase by hand on the page. When the PDF is created, FileXvert finds the images drawn under that area, taking their position, scale and rotation into account. It then repaints, in the image itself, the affected pixels with the dominant colour around the area, that is the paper background. Any text the page also contains as characters is removed from the same area.
The new image replaces the old one, and the old one is deleted from the file: it does not linger as a hidden object. The rest of the page is untouched. On the black-and-white version, every pixel outside the area is identical to the original. On the JPEG version, the image is re-encoded by the browser: the largest difference outside the area is 5 levels out of 255, invisible to the eye, and the amount of ink outside the area changes by less than 0.1%.
The size changes a little, one way or the other. The JPEG re-encoded by the browser is lighter here than the scanner's, and the black-and-white version, stored with a more general compression than CCITT, goes from 7.5 to 12.3 KB. On our test machine, the operation took under one and a half seconds for the JPEG page, and under a third of a second for the black-and-white page.
Supported formats and limits
Erasing works with the image formats found in the vast majority of PDFs: JPEG, JPEG 2000, CCITT and JBIG2 black-and-white images, and uncompressed or losslessly compressed images. An image's transparency mask is handled too, as are images placed inside a sub-object of the page.
A few cases are only hidden, not erased from the image: very small images written directly into the page content, and unusual formats. For highly sensitive information, the “Hide” tool offers an absolute guarantee: it turns the whole page into an image, blocks included, at the cost of text that is no longer selectable.
Check before sending
Whatever the method, a few quick checks are essential:
- Draw the area with a small margin around the text: a leftover descender or accent can be enough to guess a digit.
- Open the resulting PDF and zoom right in on the erased area: it must be uniform.
- If the scan has a layer of recognised text (OCR), search (Ctrl + F) for the erased information: it must no longer be found.
- Think about metadata: the editor can clear it at the same time, so that no author or software name remains in the file.
The PDF editor, and therefore the “Erase an area” tool, is reserved for Premium subscribers. Like every FileXvert feature, it works in your browser: the document is not sent to any server, which matters precisely when it contains information you want to protect.