PDF to text

Drop in a PDF and get the text from every page. Pages that already contain text are taken as they are, and scanned pages are read with OCR, all on your own device.

Drop images here, or paste with Ctrl V

JPG, PNG, WebP, HEIC, GIF, BMP, TIFF or PDF, up to 25 MB each. Add up to 50 at once.

Read on your device. Your images aren't uploaded. How this works

No image to hand? Try a sample:

How to use it

  1. Drag your PDF onto the page or open it with Ctrl+O (⌘O on a Mac), and every page is added in order.
  2. Let Image to Text App work through the scanned pages one at a time, with progress showing which page it's reading.
  3. Open Whole document view to read and search all the pages together and fix any words to check.
  4. Download everything as one Word file or a searchable PDF, or as a ZIP of separate files.

Getting text out of a PDF can mean two quite different jobs, depending on how the PDF was made. Image to Text App handles both, and it decides page by page, so you don’t have to know in advance.

Digital PDFs and scanned PDFs

A digital PDF was created by software: exported from Word, saved from a web page, or generated by a billing system. Its pages contain real text, and you can usually select words in a PDF viewer. When Image to Text App finds a page like this, it takes the text directly, without OCR. That means no recognition mistakes, and it’s quick.

A scanned PDF is a set of pictures of paper pages, from an office scanner, a phone scanning app, or a fax turned into a file. There’s no text inside, only images that look like text. Image to Text App reads these pages with OCR, one after another, and progress shows how far it has got.

Plenty of PDFs are a mix of the two: a contract with a signed and scanned final page, or a report with a few scanned appendices. Because each page is checked separately, the digital pages come through exactly and only the scanned ones are read with OCR.

One catch: some digital PDFs have a broken text layer. You can select the text, but when you copy it you get gibberish or missing letters, usually because of how the fonts were embedded. Since that text is taken as it is, the result will show the same problem. The workaround is to take a screenshot of the page and paste it here, so it’s read as an image instead.

Getting good text out of scanned pages

The quality of the scan decides most of the result. If you’re scanning the paper yourself:

  • Scan at a resolution meant for documents. 300 dpi is a common choice for text, and much lower settings make small print hard to read.
  • Choose grayscale or black and white for plain text documents. Color adds file size without helping recognition.
  • Lay pages flat and straight. Auto improve corrects a slight tilt, and you can rotate a page that went through upside down.

Old photocopies and faxes often have speckles and faded strokes. Reduce noise helps with the speckles, and raising contrast helps with faint text. Each page can be adjusted and read again on its own without touching the rest.

Working with long documents

Every page of the PDF appears as its own page in the workspace. You can reorder pages, read one again, remove it, or copy and download it on its own.

Whole document view puts all pages together in order, with search across every page. Click a line of text and its place lights up on the page image, which is the fastest way to check a figure or a name.

Document mode joins lines into paragraphs and keeps headings and lists. There are also options to join wrapped lines and to re-join words split by a hyphen at the end of a line, which scanned books and reports are full of. Running headers, footers and page numbers will appear on every page; find and replace makes them quick to remove.

Choosing an output

  • Word (.docx) is the default here, ready for editing. See image to Word for more on editable documents.
  • Searchable PDF keeps the scan looking exactly as it did and adds an invisible, selectable text layer. It’s the right choice for archiving: the file still looks like the original, but you can search it with Ctrl+F in any PDF reader.
  • TXT, Markdown, HTML and JSON are available too, as one combined file or a ZIP of separate files.

If the PDF holds a table you need in a spreadsheet, switch that page to Table mode and see image to Excel.

Limits worth knowing

Clear printed scans are usually read very accurately. Handwritten notes in the margins, filled-in forms and signatures are much harder. Stamps and watermarks over text cause errors, and multi-column layouts may need some tidying where columns meet. Simple printed equations can be read in Math mode, but anything complex usually needs fixing by hand.

For more on scan quality and what to expect, read how to extract text from scanned documents.

Questions people ask

Does it work with scanned PDFs?

Yes. Scanned pages are images, so each one is read with OCR, page by page. Pages that already contain selectable text are taken directly instead.

How can I tell if my PDF is scanned or digital?

Open it in any PDF viewer and try to select a single word. If it highlights, the page has real text; if you can only draw a box over a picture, it's scanned.

Is the text from a digital PDF run through OCR?

No. If a page already contains selectable text, Image to Text App takes that text directly, so it isn't affected by recognition mistakes.

Can I make a scanned PDF searchable?

Yes. Download the result as a searchable PDF. It keeps each page's original image and adds an invisible text layer, so you can search, select and copy text in any PDF reader.

Can I get one Word document from a long PDF?

Yes. Whole document view combines every page in order and exports it as one file. You can also download a ZIP of separate files instead.

How large a PDF can I use?

Files can be up to 25 MB each. If a scan is bigger, split it into parts in a PDF tool and add the parts in order.

Is my PDF uploaded?

No. Pages are read in your browser, on your device. That matters for contracts, statements and other documents you'd rather not send anywhere.