Getting text out of a PDF can mean two quite different jobs, depending on how the PDF was made. Image to Text App handles both, and it decides page by page, so you don’t have to know in advance.
Digital PDFs and scanned PDFs
A digital PDF was created by software: exported from Word, saved from a web page, or generated by a billing system. Its pages contain real text, and you can usually select words in a PDF viewer. When Image to Text App finds a page like this, it takes the text directly, without OCR. That means no recognition mistakes, and it’s quick.
A scanned PDF is a set of pictures of paper pages, from an office scanner, a phone scanning app, or a fax turned into a file. There’s no text inside, only images that look like text. Image to Text App reads these pages with OCR, one after another, and progress shows how far it has got.
Plenty of PDFs are a mix of the two: a contract with a signed and scanned final page, or a report with a few scanned appendices. Because each page is checked separately, the digital pages come through exactly and only the scanned ones are read with OCR.
One catch: some digital PDFs have a broken text layer. You can select the text, but when you copy it you get gibberish or missing letters, usually because of how the fonts were embedded. Since that text is taken as it is, the result will show the same problem. The workaround is to take a screenshot of the page and paste it here, so it’s read as an image instead.
Getting good text out of scanned pages
The quality of the scan decides most of the result. If you’re scanning the paper yourself:
- Scan at a resolution meant for documents. 300 dpi is a common choice for text, and much lower settings make small print hard to read.
- Choose grayscale or black and white for plain text documents. Color adds file size without helping recognition.
- Lay pages flat and straight. Auto improve corrects a slight tilt, and you can rotate a page that went through upside down.
Old photocopies and faxes often have speckles and faded strokes. Reduce noise helps with the speckles, and raising contrast helps with faint text. Each page can be adjusted and read again on its own without touching the rest.
Working with long documents
Every page of the PDF appears as its own page in the workspace. You can reorder pages, read one again, remove it, or copy and download it on its own.
Whole document view puts all pages together in order, with search across every page. Click a line of text and its place lights up on the page image, which is the fastest way to check a figure or a name.
Document mode joins lines into paragraphs and keeps headings and lists. There are also options to join wrapped lines and to re-join words split by a hyphen at the end of a line, which scanned books and reports are full of. Running headers, footers and page numbers will appear on every page; find and replace makes them quick to remove.
Choosing an output
- Word (.docx) is the default here, ready for editing. See image to Word for more on editable documents.
- Searchable PDF keeps the scan looking exactly as it did and adds an invisible, selectable text layer. It’s the right choice for archiving: the file still looks like the original, but you can search it with Ctrl+F in any PDF reader.
- TXT, Markdown, HTML and JSON are available too, as one combined file or a ZIP of separate files.
If the PDF holds a table you need in a spreadsheet, switch that page to Table mode and see image to Excel.
Limits worth knowing
Clear printed scans are usually read very accurately. Handwritten notes in the margins, filled-in forms and signatures are much harder. Stamps and watermarks over text cause errors, and multi-column layouts may need some tidying where columns meet. Simple printed equations can be read in Math mode, but anything complex usually needs fixing by hand.
For more on scan quality and what to expect, read how to extract text from scanned documents.