The secure way to redact a PDF is to use a true redaction tool that removes the underlying text rather than painting over it, then strip hidden metadata and confirm the result before you share the file. For sensitive documents, a local tool like FlowPDF's redact page keeps the file on your device instead of a server. Before you send anything out, run a quick selection test and check the metadata.
TL;DR:
- True redaction that excises underlying text and removes metadata is essential for securely concealing sensitive information in PDFs.
- Overlay redactions leave recoverable text underneath, risking accidental disclosure if not thoroughly checked and verified.
- Rasterization offers an extra layer of security by turning the page into an image, but it compromises searchability and clarity.
- Verifying redactions involves selection, search, and metadata checks across multiple PDF viewers to ensure no recoverable traces remain.
- When redacting short, unique identifiers, widening the black box beyond the text helps prevent glyph-position leakage and accidental leaks.
Table of Contents
- Overlay vs. True Redaction vs. Rasterization: Which Should You Use?
- Redacting a PDF Step by Step: Prepare, Redact, Sanitize, Verify
- How Do You Confirm a Redaction Actually Holds Up?
- Why Even Careful Redactions Can Leak Information
- Redaction Is Not the Same as Legal Anonymization
- A Practical Default for Redacting Sensitive Documents
- FlowPDF: Redact and Sanitize Without Uploading Your File
- FAQ
- Sources
Overlay vs. True Redaction vs. Rasterization: Which Should You Use?
Most PDF editors offer a "redact" button, but not every button does the same thing. Knowing the difference decides whether your sensitive data actually disappears or just gets painted over.
A visual overlay places a black box on top of text. It looks redacted on screen, but the original text objects often remain in the file's content stream. Anyone who selects the area and copies it, or exports the page to plain text, can recover the words underneath. This is the most common redaction mistake, and it is the exact failure mode that PDF risk guidance for courts and security teams warns against when documents are shared without a final check.
True redaction, sometimes called excision, deletes the text objects and replaces them with nothing recoverable, not just a colored rectangle. This is the method to use for names, identification numbers, account details, or anything a reader should never be able to retrieve, even with effort.
Rasterization converts a page into a flattened image, removing all underlying text objects, vector data, and most metadata tied to that page. It adds a layer of security because there is no text to extract, only pixels. The tradeoff is image quality: zoomed or printed pages can look softer, and the page is no longer searchable or selectable at all, which can be inconvenient for legitimate reference use.
A simple decision rule:
- Use overlay only for low-sensitivity drafts where the information is already available elsewhere.
- Use true excision for names, numbers, and identifiers that must never be recoverable.
- Rasterize, then re-check, when a document contains scanned content or when you need an extra layer of certainty on top of excision.
Redacting a PDF Step by Step: Prepare, Redact, Sanitize, Verify
A secure redaction is a sequence, not a single click. Skipping a step is usually how sensitive data ends up in the wrong hands.
- Prepare. Duplicate the original file and work only on the copy. List the specific names, numbers, or phrases you need to remove so you can test for them later, and keep the untouched original in a separate, access-controlled location.
- Redact. Open the duplicate in a tool that excises text rather than overlaying it. Select each sensitive item, apply the redaction, confirm the change, and save the result as a new file. If your tool only offers overlays, rasterize the page first and then apply the overlay on top of the flattened image.
- Sanitize. Strip metadata, comments, hidden layers, attachments, form field data, and embedded scripts. These often live in a document's properties panel, a "clean up" or "sanitize document" menu, or an attachments pane, depending on the editor. A document can disclose more than what is visibly printed on the page, including author names, file paths, and revision history, according to PDF risk guidance used by courts and security teams.
- Verify. Try to select and copy text from the redacted areas. Search the document for the exact terms on your list. Export the file to plain text and search again. Open it in a second PDF reader to confirm the result holds up outside your original editor.
Pro Tip: Keep your list of redacted terms in a separate text file, then paste each one into the search bar of the finished PDF as a final check before sending it anywhere.
Scanned documents deserve extra attention. If a scan has gone through OCR, the image now carries an invisible, searchable text layer sitting on top of the picture. Redacting the visible image alone will not remove that hidden text layer, and a careless save can leave the original OCR text fully intact and copyable underneath your black box. The safer approach is to rasterize the page after redacting, which discards the OCR layer entirely, or to remove the OCR layer before you redact and reapply OCR afterward if you still need the document searchable.
How Do You Confirm a Redaction Actually Holds Up?
A redaction that looks finished on screen is not the same as one that is finished. A handful of quick checks separate the two.
- Selection and copy test: click and drag across the redacted area, then paste the result into a blank document. Nothing should appear.
- Search test: use the PDF reader's search function to look for the exact name, number, or phrase you redacted. A match means the redaction failed.
- Metadata and attachment check: open the document properties and attachments pane and confirm author names, custom fields, and embedded files are gone.
- Cross-reader test: open the file in a different PDF reader and repeat the selection and search tests, since some viewers render hidden text that others suppress.
Verification workflows combining selection tests, search, and metadata inspection catch the majority of common redaction tool mistakes and are practical even for non-expert users. This matters because the failure is rarely dramatic. It is usually a forgotten overlay or a metadata field nobody thought to check.
If any test turns up a trace of the original content, do not patch the existing file. Rasterize the page, enlarge the redaction box slightly beyond the original text boundary, and redact again on the flattened version.
Why Even Careful Redactions Can Leak Information
Redaction failures are not always as obvious as a selectable black box. Some of the more subtle risks come from how PDFs store text internally.
Research on glyph positions in PDF text redaction found that many common redaction tools leave information recoverable through glyph shifts and non-excising techniques, and the researchers successfully deredacted real-world documents using this method. A related analysis of the same glyph-position vulnerabilities showed that even short names and other high-value identifiers could be recovered from supposedly redacted files by exploiting how individual character positions shift around a hidden word.
- Non-excising redactions draw a box but leave the text selectable underneath, the same overlay problem described earlier.
- Glyph-position leakage can expose short, unique identifiers even when the visible text has been properly excised, because the spacing of surrounding characters hints at what used to be there.
- Rasterization closes most of this gap but is not an absolute guarantee, since residual layout clues can survive flattening, such as an approximate redaction width.
Practical mitigations include expanding the redaction box beyond the exact text boundary, rasterizing and re-checking the page, removing surrounding context that narrows down the hidden word, and converting to a monospaced font representation when the document type allows it.
Pro Tip: When redacting a short, unique identifier like a last name or a four-digit code, widen the black box well past the original letters rather than matching it exactly to the text.
For datasets with many records or highly sensitive legal material, redaction alone may not be enough, and anonymization or a legal review is the safer path.
Redaction Is Not the Same as Legal Anonymization
Redacting a PDF is a technical step: removing specific visible content from one document. Anonymization is a broader, risk-based legal concept that asks whether a person could still be identified from what remains, even indirectly, across other data sources.
PDPC's advisory guidelines on the Personal Data Protection Act.pdf) distinguish redaction from anonymization and recommend that organizations apply risk-based anonymization controls when handling personal data, rather than treating a redacted document as automatically compliant. Techniques like attribute suppression, masking, and pseudonymization are covered in PDPC's guide to basic anonymization techniques, which treats page-level redaction as one control among several rather than a complete solution on its own.
If you are preparing a court filing, a public records disclosure, or a research dataset with personal information, redaction gets the visible text out, but a compliance or legal review decides whether the remaining document meets your actual legal obligations.
A Practical Default for Redacting Sensitive Documents
Most people overestimate what a black box on a PDF actually does. My default recommendation is unglamorous but reliable: true redaction that excises the text, followed by metadata sanitization, followed by an independent verification pass, every time, regardless of how routine the document feels.

Treat short, unique identifiers with extra caution. A last name, a case number, a four-digit code: these are exactly the items that glyph-position research has shown can leak even after a technically correct redaction. When in doubt, widen the redaction area or move the data into an anonymized dataset instead of a redacted page.
Keeping the file on your own device during this process removes one entire category of risk: a server you do not control holding a copy of a document you just spent ten minutes trying to scrub clean.
— Ronald Ang
FlowPDF: Redact and Sanitize Without Uploading Your File
We built FlowPDF around the workflow described above: everything runs in your browser, nothing uploads to a server, and every feature is free with no watermark. For a document you would not want sitting on someone else's server, that local processing is the practical advantage.

Our redact PDF tool handles the excision step directly, removing the underlying text rather than drawing an overlay on top of it. From there:
- Use edit PDF to review and strip metadata, comments, and embedded objects before you export.
- Use remove pages when a whole page is too sensitive to redact piece by piece.
- Browse our full set of free PDF tools for merging, compressing, or organizing the same file afterward.
Once you have redacted and sanitized a file with us, run the verification checks from this guide, including wiping verification methods such as the selection test, the search test, and a cross-reader check, before you send it anywhere. Start with the redact PDF page and work through your document one field at a time.
This article is general information, not a substitute for advice from a qualified lawyer. Consult a qualified legal professional about your own circumstances before acting on anything here.
FAQ
How do I redact a PDF for free?
Several free tools let you redact a PDF without a subscription, including browser-based options like FlowPDF's redact tool that process the file locally. Look for a tool that explicitly removes the underlying text rather than one that only draws a black box over it.
What is redact in PDF?
Redaction in a PDF means permanently removing specific sensitive content, such as a name, number, or paragraph, so it cannot be read, searched, or copied by anyone who opens the file. A proper redaction excises the text object itself, unlike a simple visual cover that can leave the original text intact underneath.
How do I black out text in a PDF editor?
Select the text or area you want to black out, apply the redaction tool, confirm the change, and save the file as a new document rather than overwriting the original. If your editor only offers a cosmetic overlay, rasterize the page first so there is no underlying text left to recover.
Can I redact a PDF so text cannot be uncovered?
A true excision redaction combined with metadata sanitization and a verification check comes close, but research on glyph-position leakage shows that short or unique identifiers can sometimes still be inferred through subtle layout clues. For highly sensitive items, widen the redaction area, rasterize the page, and consider anonymization instead of redaction alone.
Sources
- Story Beyond the Eye: Glyph Positions Break PDF Text Redaction (PETS 2023)
- Story Beyond the Eye: Glyph Positions Break PDF Text Redaction (arXiv)
