How to get the text out of a scanned PDF

Whether you can copy text out of a PDF depends entirely on how it was made, and it is worth thirty seconds to find out before you start retyping anything.

Extract text
Whichever tool you use here, the file stays on your device. How that works

Step by step

  1. 1Open the PDF in any reader and try to select a line of text with your cursor.
  2. 2If the text highlights, it is real text. Run it through the extract tool and you will get all of it at once.
  3. 3If nothing highlights, or the whole page selects as one block, your PDF is a picture and there is no text in it to extract.
  4. 4For a picture based scan you need optical character recognition, which is a different job. We do not do it, and we would rather say so than waste your time.

Two things are both called a PDF

A PDF exported from a word processor contains the actual characters. They can be searched, selected, copied and pulled out in full, with formatting stripped, in a second or two.

A PDF produced by a scanner or a phone camera contains photographs of pages. To a computer there is no text there at all, just an image that happens to look like writing to you. Searching it finds nothing and selecting gives you nothing, which is why people think their reader is broken.

Turning the second kind into the first requires optical character recognition, which reads the shapes and guesses the letters. It is genuinely useful and genuinely imperfect, especially with handwriting, poor scans and unusual layouts. We do not offer it today. If your document is short, retyping the part you need is often faster than fixing what an OCR pass gets wrong.

Questions

How do I know which kind I have?

Try to select a line of text. If it highlights word by word, there is real text. If nothing happens, it is an image.

Do you do OCR?

Not at the moment. We would rather tell you plainly than have you upload a document and find out.

Is my PDF uploaded when I extract text?

No. It is read inside your browser.