Security
How to redact personal data from a PDF before sharing it (GDPR and subject access requests)
Search the document for every name, email address, phone number, account number and ID belonging to people the recipient should not see, redact each match so it is deleted rather than covered, clear the document properties, and check the finished file before sending it.
A common job, and an uncomfortable one: you have to hand over a document, most of it can go, and parts of it are about people who are not the recipient. A reply to someone asking for their own data that also mentions colleagues. Emails disclosed in a dispute. Minutes being published. A file for a contractor that happens to include a customer list.
The answer each time is redaction. The trap each time is doing it in a way that only looks finished.
Why data protection law cares how you redact
Under the GDPR, and the UK GDPR that kept the same principles, people can ask an organisation for a copy of the personal data it holds about them. This is usually called a subject access request. The documents that answer it often contain other people's personal data too: colleagues, witnesses, other customers. That data generally has to be protected unless there is a good reason to disclose it, and redaction is the usual way of doing that.
The same logic runs through data minimisation: share what the recipient needs, not everything the document happens to contain.
Brazil's LGPD and many other data protection laws are built on the same two ideas: people can see their own data, and everyone else's is protected. Court filings in many countries also have their own rules about which personal details, such as account numbers or dates of birth, must be removed from documents placed on the public record. The rules of the court you file with decide exactly which.
Whatever the law, one point is the same everywhere. A black box that can be copied away has not protected anyone. In practice it is a disclosure.
What counts as personal data in an ordinary document
More than names. Go through the document with this list:
- Names, including initials, nicknames, usernames, signatures, email sign-offs and "cc" lines.
- Contact details: email addresses, phone numbers, home addresses.
- ID numbers: national ID, passport, tax, employee and customer numbers.
- Financial details: card numbers, bank account numbers, IBANs.
- Dates of birth and other dates tied to a person.
- Images: photographs, signatures, handwritten notes.
- Indirect identifiers: a job title that only one person holds, a vehicle registration, a description that points at one person.
- Things that are not on the page: the author name in the document properties, the names of reviewers attached to comments, the file name itself.
Finding it all: search beats reading
Reading every line and hoping to catch every mention works for two pages. For fifty it fails, quietly. Redact PDF searches every page at once, and has built-in patterns for the kinds of data that follow a format:
| What you are looking for | How to find it in Redact PDF |
|---|---|
| Email addresses | The Email addresses pattern |
| Phone numbers | The Phone numbers pattern. It picks numbers written like phone numbers or sitting next to a phone label, and skips invoice numbers and codes, so check any number it left out |
| Card and bank account numbers | Card and account numbers, and IBANs for international bank numbers |
| Dates of birth | Dates finds every date. Untick the ones that are not birthdays |
| US Social Security numbers | ID numbers (000-00-0000). Other national ID formats: search for each number itself |
| Web addresses | Web addresses, useful for personal profile links |
| Names | Search each name and each variant. Tick Whole words so "Ann" does not match inside "Annual" |
| Signatures, photos, handwriting | Draw a box over them |
Every search result appears in a list, with a snippet of the text around it, for you to review and tick. Nothing is marked until you press Mark, and nothing is removed until you apply.
Scanned documents need one extra step
A scanned letter is a picture, and search cannot read a picture. When you open a scanned PDF, the tool lists the pages with no text and offers to recognise it. After that, the patterns and name searches work on those pages as well.
Recognition is not perfect, and a misread name is a name the search will not find. On scanned pages, read each page yourself as well, and draw boxes over anything the search missed. If a scan is too grey to read well, see how to make a speckled scan readable.
Step by step
- Open the PDF in [Redact PDF](/redact-pdf). Your original on your device is never changed.
- Run each pattern search, then each name. Tick the matches that belong to someone the recipient should not see and press Mark.
- Page through for images. Signatures, photos and handwritten notes need a drawn box.
- Keep "Document properties" ticked under "Also remove hidden information". It is ticked to begin with. Consider comments and attached files too: both often carry names.
- Apply redactions. Read the result screen. It confirms no text remains under any mark, and tells you if a searched name still appears anywhere else, such as a page you did not mark, a bookmark or a comment, with the option to remove it and apply again.
- Rename the file if needed. The redacted copy is named after the original with -redacted added. If the original file name contains a person's name, change it before sending.
Keep a record of what you withheld
It is good practice to keep the unredacted original, plus a short note of what was removed and why, somewhere only the right people can reach. If the decision is ever questioned, you can show what was withheld and on what basis. Redaction cannot be reversed in the file you sent, so the original is your only record.
Where your file goes while you work
Redaction runs on our server, so the document is uploaded over an encrypted connection. It is kept only while you work: deleted within 30 minutes, or straight away when you choose "Redact another file" or leave the page. It is never used for anything else.
If your organisation's policy says a document may not be uploaded to any outside service, the policy wins. Follow it.
Check before it leaves
Before you send the file, run the checks in why blacked-out text can still be read: search for each name in an ordinary viewer, copy the text out, and look at the document properties. Two minutes, on the exact file you are about to send.
Common questions
Do I have to redact other people's names in a subject access request?
Not automatically. The other person's rights have to be weighed against the requester's, and sometimes a name can be disclosed, for example with consent or where the person acted in a professional role that is already known. Many organisations redact by default and disclose where there is a clear reason. Your data protection lead makes that call.
Is replacing a name with initials enough?
Often not. In a small team, initials identify someone just as well as a full name. Ask whether the recipient could work out who it is from what is left, including job titles, dates and context, and redact until they could not.
Can I redact the same names in many PDFs at once?
Redact PDF works on one file at a time. If the documents are going to the same person anyway, you can combine them with Merge PDF, redact once, and split them again with Split PDF if they must be sent separately.
Does KovaPDF keep a copy of the document?
No. The upload and the redacted copy are deleted within 30 minutes, or immediately when you choose Redact another file or leave the page.