Security
Why can blacked-out text in a PDF still be read, and how do I check mine?
Because a black box drawn on a PDF is just another shape on top of the page: the text underneath is still stored in the file, so anyone can copy, search or extract it unless a redaction tool actually deleted it.
It keeps happening, to law firms, companies and public bodies alike: a document is released with names and figures neatly blacked out, and within hours someone has copied the hidden text into a message. Nobody cracked anything. The text was simply still there.
Here is how that happens, and how to be sure it has not happened to the file you are about to send.
What a failed redaction actually is
A PDF page is built from instructions: draw these characters here, this image there. A black rectangle is one more instruction, drawn after the text. On screen, the rectangle wins. In the file, both exist.
Anything that reads the file rather than looking at it ignores the rectangle completely: copy and paste, the search box, a screen reader, a text-extraction tool, a search engine indexing the document.
The six ways it usually goes wrong
- A shape or a highlighter. The rectangle or highlight tool in a PDF reader adds a comment on top of the page. It is not even part of the page, so it can be moved or deleted, and the text below is untouched.
- Black shading in Word or Excel. Black text on a black background, or a filled shape over a cell, looks redacted. Exported to PDF, the text is exported with it.
- Covering the picture of a scan but not its text. Many scanned PDFs have an invisible text layer from OCR, lying exactly over the words in the image. Painting over the image leaves that layer, word for word.
- Marked but never applied. Proper redaction tools usually work in two steps: mark, then apply. A file saved between the two has boxes that look finished and remove nothing.
- A clean page in a leaky file. The page is fine, but the document properties say "Author: J. Smith", a bookmark repeats the hidden heading, a comment quotes the figure, or an earlier version of the file is still saved inside it.
- Cropped, not removed. Cropping a page hides what is outside the crop. It does not delete it, and another viewer, or a tool that ignores the crop, shows it again.
Does redacting in Adobe Acrobat really remove the text?
When you use Acrobat's dedicated Redact tool and apply the redactions, it is designed to remove the content, and it does. The leaks blamed on Acrobat almost always come from something else: a rectangle or highlight from the comment tools, a black box added while editing the page, or redactions marked and then saved without being applied. Acrobat also offers a separate step for removing hidden information such as metadata, and it is easy to skip.
The same is true of any tool. What matters is whether the content was deleted, not which program drew the box.
Five checks before you send a redacted PDF
Run these on the exact file you are about to send, not on the working copy.
- Search for each redacted word. Open the file in an ordinary PDF viewer and press Ctrl+F (Cmd+F on a Mac). Search for every name and number you removed. A single match means it is still in the file.
- Copy everything and paste it somewhere plain. Select all, copy, and paste into a plain text editor. Read the lines around each box. Hidden text often appears in the middle of an otherwise normal sentence.
- Extract the text with a second tool. PDF to Text pulls out every word the file contains. Search the result the same way.
- Look outside the pages. Open the document properties (usually File, then Properties), the bookmarks panel, the comments panel and the list of attachments. Names hide in all four.
- Open it in a second viewer. A browser's built-in viewer and a desktop reader do not always show the same things. What one hides, the other may display.
These checks are good at catching text. They are weaker at pictures: a photo or signature painted over with a shape can still be under it, and that is harder to test by hand. For images, the safer path is a tool that overwrites the pixels in the first place.
Fixing a PDF that was blacked out the wrong way
If you still control the file, you can repair it. Open it in Redact PDF and search for the words that should be hidden. If the search finds them under the old boxes, that is your proof the text survived. Mark the matches, draw boxes over any images, and apply. The words are removed this time, and the result screen confirms that no text is left under any mark.
If the old boxes were drawn as comments, tick Comments and markup under "Also remove hidden information" to clear them away at the same time.
If the file has already been sent or published, a new copy does not change the copies other people hold. Tell whoever in your organisation deals with that kind of disclosure.
What KovaPDF checks for you
After you apply, the finished file is read again by a different PDF reader from the one that wrote it. It checks that no text sits under any mark, and it reports every place a searched word still appears: another page, document properties, bookmarks, comments, form fields or the names of attached files. A page that cannot be cleaned with certainty is turned into an image with the boxes burnt in, rather than handed back leaking.
For the full walk-through, see how to redact a PDF permanently.
Common questions
Can someone remove a black box from a PDF?
If the box is a comment or a shape drawn on top, yes: it can be deleted, or the text beneath copied around it. If the content was properly redacted, removing the box reveals nothing, because there is nothing left underneath.
Does flattening a PDF remove the text under a black box?
No. Flattening merges the box into the page, so it can no longer be moved, but the text instructions beneath it are still there, and copying or extracting the text still finds them.
If I print the blacked-out page and scan it, is it redacted?
The scan is a picture with no text layer, so the words are gone from the new file. It is a crude method: the whole document loses its searchable text and some quality, and a marker pen on paper can sometimes show through on a strong scan. Check that the boxes are fully opaque on the scan.
Can I use Redact PDF just to check a file?
Yes. Open the file and search for the words that should be hidden. Search reads the file's text, not the picture on screen, so if it finds a word under a black box, that word is still in the document.