Working with scans

How do I copy text from a PDF when copying doesn't work?

Open the file in PDF to Text instead of selecting. Pages with real text are read directly, and pages that are only pictures of text are recognised with OCR. You can then copy one page, copy everything, or download a .txt file.

4 min read

You need three paragraphs out of a PDF. You drag across them and nothing highlights. Or the whole page turns blue as one block. Or it copies, and what you paste is a line of boxes and symbols.

Each of those is a different problem, and knowing which one you have tells you what will work.

What the symptom tells you

What happensWhat it meansWhat works
Nothing highlights, or the whole page selects as one blockThe page is a picture of text, usually a scanPDF to Text, which recognises it with OCR
It pastes as boxes, symbols or wrong lettersThe text is there, but its font has no usable character mapRead the picture instead: PDF to Image, then Image to Text
Your viewer says copying is not allowedThe file carries a copying restrictionSee below
It pastes with a line break after every lineNormal: a PDF stores lines, not paragraphsJoin the lines in a text editor

Getting the text out, step by step

  1. Open PDF to Text.
  2. If the document is scanned, set Document language. The default, Detect automatically, works for most files. You can also choose Bangla, English, Hindi, Arabic, Chinese (Simplified), Russian, Spanish or Dutch; each is read together with English. The setting only affects scanned pages.
  3. Choose the PDF. Every page is read; the status line shows how many are ready and how many needed OCR.
  4. Take the text out with one of the three controls below.
  • Copy, in the top corner of each page, copies just that page.
  • Copy All Pages copies the whole document.
  • Download as .txt saves everything to a text file, with each page under a "Page N" heading.

On a phone, Copy All Pages and the .txt download sit at the bottom of the screen.

There is no selecting involved, so it does not matter whether your viewer can highlight the text or not.

Where the work happens

Pages that already contain real text are read by your browser, and for those the file never leaves your device. Pages that turn out to be pictures are drawn as images and sent together to our server to be recognised, and the text comes back.

Recognised text is an estimate. On a clean scan it is usually very good; on a faint or crooked page it will have mistakes. Check names and numbers before relying on them.

When the text pastes as gibberish

This happens with some PDFs made by older software or unusual fonts. The letters look right on screen, but the file does not record which character each shape is. Copying, and any tool that reads the text layer, including PDF to Text, gets the same meaningless codes.

The way round it is to ignore the text layer and read the picture:

  1. Convert the pages you need to images with PDF to Image, choosing PNG.
  2. Put those images through Image to Text, which reads up to five images at a time.

When a scan has a little real text on it

PDF to Text decides page by page: a page with any real text is read as text; a page with none is sent for OCR. A scan with a small typed stamp on it, such as a page number or a date added by the scanner, counts as text, and you get just the stamp.

If a scanned page comes back almost empty, use the same PDF to Image and Image to Text route for that page.

When copying is "not allowed"

Some PDFs carry a flag asking viewers not to allow copying. That is a restriction on the file, not a technical barrier to reading it, and it is covered in detail in My PDF opens but won't let me print or copy.

If it is your document, or you have the right to use its text, Unlock PDF removes the protection from your copy. A PDF that needs a password just to open has to be unlocked first; PDF to Text does not ask for passwords.

About the line breaks

The text comes out one line per line, as it was laid out on the page. That is how PDFs store text. A paragraph pasted into an email arrives with a break at the end of every line.

To rejoin it, paste into a word processor and replace single line breaks with spaces. On pages with two or more columns, check the order too. The lines are read in the order the file stores them, which can run across columns rather than down them.

Common questions

Why can't I highlight any text in my PDF?

Because the page is a picture of text, usually a scan or a photo, with no text layer. PDF to Text recognises such pages with OCR automatically.

Why does copied PDF text paste as symbols or boxes?

The font in the PDF has no usable character map, so the copied codes mean nothing. Convert the page to PNG with PDF to Image and read it with Image to Text instead.

Is the extracted text exact?

For pages with real text, yes: the characters are read directly. For scanned pages it is an OCR estimate, so check names and numbers.

Can I copy just one page?

Yes. Every page has its own Copy button. Copy All Pages and Download as .txt give you the whole document.

Is my PDF uploaded?

Pages with real text are read in your browser and never leave your device. Only pages that need OCR are sent, as images, to be recognised.

Which languages can it recognise in scans?

Bangla, English, Hindi, Arabic, Chinese (Simplified), Russian, Spanish and Dutch, or it can detect the language automatically. Each is read together with English.