If your PDF already has selectable text, copy it or export it straight to TXT or DOCX. If it's a scan or a photographed page, you need OCR to turn the image into readable text. For anything sensitive, a local-in-browser tool keeps the file on your device instead of sending it to a server.
TL;DR:
- Batch OCR tools can process up to 100 files at once, making large-scale extraction faster but requiring pre-testing for varying image quality.
- Exporting selectable text PDFs to DOCX or TXT preserves structure better than copy-paste, especially for complex layouts like multi-column documents.
- For sensitive files, local-in-browser OCR software ensures that documents remain on your device, avoiding privacy concerns associated with online tools.
- Running OCR on poor-quality scans or incorrectly configured language settings leads to errors and garbled text, emphasizing the need for proper preprocessing.
- Separating mixed content PDFs into pages with selectable text and scanned images improves accuracy, as each can be processed with the appropriate method.
Table of Contents
- Quick methods at a glance: which approach to pick now
- Extract text from selectable PDFs: precise steps for desktop and browser
- Extract text from scanned or image PDFs using OCR: step-by-step workflow
- Preprocessing and quality tips that materially improve OCR accuracy
- Output formats and structured extraction: OCR versus KIE
- Privacy and security: keep sensitive PDFs local and compliant
- How to batch extract text from multiple PDFs efficiently
- Common pitfalls and troubleshooting tips during extraction
- How to handle PDFs with mixed content: selectable text and scanned images
- How to extract text from password-protected or encrypted PDFs
- Author perspective: when FlowPDF is the right choice
- Try FlowPDF: extract or convert text without uploading a file
- Sources
- FAQ
Quick methods at a glance: which approach to pick now
Most PDFs fall into one of four buckets, and each one points to a different fix.
- Selectable text PDFs: copy-paste or export directly, no OCR needed.
- Scanned or photographed pages: run OCR, online or local, to generate a text layer.
- Batches of many files: use a tool built for bulk processing instead of doing one at a time.
- Structured fields like names, dates, or invoice totals: Key Information Extraction (KIE) beats plain OCR because it returns validated values instead of a wall of text.
Copy-paste is fastest when it works, but it can mangle spacing in multi-column layouts. Exporting to DOCX or TXT keeps more structure intact. Online OCR tools are convenient but require an upload, which matters if the document is private. Desktop OCR software and mobile scanner apps give you more control over image quality before recognition runs. For high volumes or compliance-heavy work, government-grade platforms or bulk APIs process many files at once and return structured output rather than loose text.
Extract text from selectable PDFs: precise steps for desktop and browser
Check first: open the file and try selecting a line of text with your cursor. If it highlights like normal text, it's selectable. If nothing highlights or you get a fuzzy image outline, it's a scan and you'll need OCR instead.
- Open the PDF in your browser or a PDF reader.
- Select the text you need with your cursor, or use "Select All" for the whole document.
- Copy it, then paste into a plain text editor to strip formatting, or into a word processor to keep more layout.
- If you need the whole file as a clean document, use "Save As" or "Export" to TXT or DOCX instead of copying by hand.
- For content you plan to reuse in a wiki, README, or blog post, convert the PDF to Markdown to keep headings and lists intact.
Pro Tip: When copy-paste produces broken line breaks, export to DOCX instead. Word processors handle paragraph structure better than a clipboard paste.
Two things commonly block this path. Some PDFs use embedded fonts that display fine but don't map cleanly to standard characters, so copied text turns into garbled symbols. Others are password-protected, which stops copying until the file is unlocked. Both are fixable, and we cover encrypted files separately below.
Extract text from scanned or image PDFs using OCR: step-by-step workflow
OCR works in three stages: it looks at the image, recognizes shapes as characters, and outputs the result as text. No selectable layer exists in a scan, so this is the only way in.
Using a typical online OCR tool:
- Upload the scanned PDF to the tool.
- Select the document's language so the recognition engine matches the right character set.
- Run OCR and wait for processing.
- Download the result, usually as TXT, DOCX, or a searchable PDF.
Running OCR locally in the browser:
- Open the scanned file in a local-in-browser tool.
- Run the built-in text extraction, which processes the file using your device's own resources.
- Export the recognized text without the file ever leaving your machine.
Local processing skips the upload step entirely, which matters for anything you'd rather not send to a third-party server: contracts, medical records, financial statements. It's also often quicker for a single file since there's no upload or download wait.
Desktop OCR software and mobile scanner apps are worth choosing when you're the one capturing the image. A mobile scanning app lets you frame the page and check for glare and shadow reflection before you commit, which produces a cleaner input than a phone photo dropped straight into an OCR tool later.
Pro Tip: If OCR output has scattered typos in a document with consistent formatting, like a form or invoice, re-run it as Key Information Extraction instead of full-text OCR. KIE validates individual fields rather than guessing at every character.
Preprocessing and quality tips that materially improve OCR accuracy
Clean input produces clean output, and the biggest accuracy gains come before OCR ever runs.
- Scan at 300 DPI or higher. Lower resolutions blur small characters into unrecognizable shapes.
- Avoid heavy JPEG-style compression, which introduces artifacts OCR engines read as noise.
- Deskew crooked pages and crop tightly to the page edges before running recognition.
- Increase contrast and remove background shading so text stands out clearly from the page.
- Re-scan blurred or photographed pages rather than trying to fix them after the fact. Software correction helps, but it can't recover detail that was never captured.
- For multi-column layouts and tables, check the output carefully. These structures are the most likely to scramble reading order during recognition.
Preprocessing like this, deskewing, normalization, and text enhancement, is part of why government OCR pipelines pair image cleanup with recognition instead of running raw scans straight through, according to documentation from Singapore's GovTech.
Output formats and structured extraction: OCR versus KIE
Not every OCR output is equally usable. Markdown output can preserve headings and simple tables, which is useful if you're reusing content in a document or a website. Plain TXT strips almost everything but the words themselves, and DOCX sits in between, often keeping paragraph breaks but losing complex table structure.
- Choose plain OCR when you need the full body of text and don't care about individual data points.
- Choose Key Information Extraction (KIE) when you need specific fields, like a name, a date, or an invoice amount, returned as clean, validated values rather than raw text you have to search through yourself.
- AISAY's own documentation recommends KIE over plain OCR for many field-based use cases, since it couples recognition with language understanding to find values even when labels vary across documents.
- Expect to manually clean up tables and multi-column sections regardless of which method you use.
Privacy and security: keep sensitive PDFs local and compliant
Uploading a document to an online tool means it leaves your device, and you're trusting that service's retention policy and access controls. Some tools keep files temporarily, some delete them right after processing, and few make that policy easy to find.
- Read the tool's retention policy before uploading anything sensitive, if one is published at all.
- Prefer local-in-browser processing for personal, financial, or legal documents.
- Use a local tool for quick, single-file jobs where speed and privacy both matter.
- Reserve online platforms for cases where you specifically need their scale or structured output.
Local processing skips the server entirely. FlowPDF explains its local-in-browser model here: files are read and processed on your own device, so nothing is transmitted for a basic extraction task.
For high-volume or regulated workloads, government-backed platforms are built differently. GovTech's AISAY documentation states that documents are not stored after processing and that portal access requires gov.sg or edu.sg credentials, which suits compliance-heavy environments better than a general-purpose online converter.
How to batch extract text from multiple PDFs efficiently
Doing ten files one at a time wastes time you don't need to spend. If every file is selectable text, most PDF readers let you open several documents and export each to TXT or DOCX in sequence, which is manageable for a handful of files but tedious past that.
For scanned documents, batch OCR tools process multiple files in one run instead of requiring you to upload and download each individually. Some platforms cap how many files you can process in one batch, so check that limit before queuing a large folder. AISAY, for example, can process up to 100 files in a single batch and returns structured output as JSON or a downloadable spreadsheet, which suits anyone processing a large folder of invoices or forms at once.
Before running a full batch, test the workflow on two or three representative files first. Scanned documents vary in quality even within the same folder, and a setting that works for a clean scan might fail on a photographed page mixed into the same batch. Once you confirm the settings hold up, queue the rest.
If your files are a mix of selectable and scanned pages, separate them first. Running OCR on a PDF that already has a text layer wastes processing time and can occasionally produce worse results than just exporting the existing text directly.
Naming your output files consistently, matching each text file to its source PDF, saves real time later when you're trying to match extracted content back to the original document, especially once you're past a dozen files.

Common pitfalls and troubleshooting tips during extraction
Garbled characters after copy-pasting usually mean the PDF uses embedded or non-standard fonts that don't map cleanly to text encoding. Exporting to DOCX instead of copying by hand often resolves this, since the export process handles font mapping differently than a clipboard copy.
Missing sections after OCR often trace back to poor image quality rather than a software failure. Blurred, low-resolution, or poorly lit scans lose fine detail that recognition engines need, so check the source image before assuming the tool is at fault.
Scrambled reading order, where paragraphs come out in the wrong sequence, is common with multi-column layouts and complex tables. OCR engines read left to right by default and can jump between columns mid-sentence. Reviewing and manually reordering the output is often faster than trying to fix this through settings alone.
Wrong language settings produce recognizable but heavily garbled text, since the engine is matching characters against the wrong alphabet or character set. Double-check the language selection before running OCR, particularly with documents that mix languages.
If a batch job returns results for some files but not others, check for encrypted or corrupted files hiding in the folder. These commonly fail silently rather than throwing a clear error, so isolating and testing the specific file that failed usually reveals the problem.
How to handle PDFs with mixed content: selectable text and scanned images
Plenty of real-world PDFs are not purely one type. A contract might have typed clauses on most pages and a scanned signature page at the end. A report might include typed paragraphs alongside a scanned chart or photographed diagram.
The safest approach is to treat the document as two separate jobs rather than forcing one method across the whole file. Extract the selectable pages using copy-paste or export, since running OCR on text that's already selectable adds an unnecessary step and can introduce recognition errors where none existed. Then isolate the scanned or image-based pages and run OCR on those specifically.

If you need to isolate pages before processing, use a tool that lets you pull specific pages out of a larger document. Extracting pages first makes it easier to apply the right method to the right content instead of running the whole file through one process and hoping for the best.
Once both pieces are extracted, combine the outputs into a single document. This is manual work, but it's usually faster and more accurate than relying on a single tool to correctly detect and handle both content types automatically, since automatic detection isn't always reliable on complex layouts.
Watch for embedded images that look like text but aren't, such as a scanned stamp or logo placed on an otherwise typed page. These will not respond to copy-paste, and it's easy to assume the whole page is broken when only that one element needs OCR.
How to extract text from password-protected or encrypted PDFs
A password-protected PDF blocks copying and often blocks OCR tools from even opening the file until the password is entered. The fix depends on whether you know the password or need to remove protection you're authorized to remove.
If you have the password, most PDF readers will prompt for it on open, and once entered, the document behaves normally for copying, exporting, or OCR. If you're the rightful owner or have permission to remove the restriction, an unlock tool can strip the password so the file behaves like any other PDF for extraction purposes.
Be careful with online password-removal tools for sensitive documents. You're sending a protected file, often protected for a reason, to a third-party server, which somewhat defeats the purpose of the protection in the first place. A local-in-browser unlock tool avoids that exposure since the file and password never leave your device.
Once unlocked, treat the file according to its content type: selectable text gets copied or exported, scanned pages get OCR. Encryption itself doesn't change which extraction method you need. It just adds a step before you can use either one.
Author perspective: when FlowPDF is the right choice
Some local-in-browser tools fit the common case well: a privacy-conscious person who needs text out of a PDF quickly, without creating an account or sending a file to a server. For a single contract, form, or scanned letter, such a tool may be the right choice.
It's not the right tool for enterprise-scale KIE, hundreds of files at once, or documents that require an audit trail. For that, a dedicated batch platform or a government-grade service makes more sense.
— Ronald Ang
Try FlowPDF: extract or convert text without uploading a file
FlowPDF handles the extraction workflows this guide covers, directly in your browser, with no signup and no watermark. If you have a document with a text layer, PDF-to-Markdown output keeps your headings and lists intact for reuse elsewhere. If you only need a few pages out of a longer file, the page extraction tool pulls exactly what you need without processing the whole document.

- Convert your PDF to Markdown to preserve structure for reuse.
- Extract specific pages instead of processing an entire file.
- Browse the full set of free PDF tools for editing, merging, and more, all processed locally.
Every one of these runs on your device, so your file never leaves your browser. Start with the free PDF tools page and pick the workflow that matches your document.
FAQ
Can I extract all text from a PDF?
Yes, if the PDF has selectable text, copy or export it directly and you'll get everything on the page. If it's a scanned or image-based PDF, you need OCR to recognize and output the text, and accuracy depends on scan quality.
Can ChatGPT extract text from a PDF?
General AI chat tools can read and summarize text from a PDF you paste or upload, but they aren't dedicated OCR or extraction tools and may struggle with complex layouts or scanned images. For reliable extraction, a purpose-built OCR tool or a local-in-browser extractor gives more consistent results.
Why does it say "extracting text from a PDF" when nothing happens?
This usually means the tool is running OCR on a scanned page, which takes longer than copying selectable text and can stall on low-quality or very large images. Try a smaller file first, check your internet connection if using an online tool, or switch to a local-in-browser tool that doesn't depend on an upload.
How can I extract text from a PDF for free?
For selectable text, open the file in any PDF reader and copy or export it at no cost. For scanned PDFs, a free local-in-browser tool such as FlowPDF runs OCR without an upload, a signup, or a watermark.
