How to pull data from a scanned PDF into Word

How to get names, numbers and dates out of a scanned PDF and into a Word document: telling scans from digital PDFs, what OCR can read, and how to review.

To pull data from a scanned PDF into a Word document, recognition first turns the picture of text back into text, then the values you need are entered into the Word file, and you review them against the original. Whether the first step is needed at all depends on which kind of PDF you have, so start by checking that.

Scanned or digital: check this first

A PDF is one of two very different things wearing the same file extension. A digital PDF was exported from software and carries a text layer: the characters are in the file. A scanned PDF is a photograph of paper wrapped in a PDF container: there is no text in it at all, only pixels arranged to look like text. Everything about your next step depends on which one you are holding.

  • Try to select a word. If the cursor selects individual words and letters, the PDF is digital. If your selection draws a box over the page image, it is a scan.
  • Search for a word you can see on the page. A digital PDF finds it; a scan finds nothing.
  • Zoom to 400 percent. Digital text stays crisp at any size; scanned text dissolves into pixels.

Why copy and paste fails on a scan

Copy and paste fails on a scan because there is nothing to copy. The file contains an image, not characters, and no amount of careful selecting changes that. Opening the PDF in Word does not rescue the situation either. Word 2013 and later can convert PDFs into editable documents, but Microsoft's own support page on opening PDFs in Word states that the conversion works best with files that are mostly text, and that the result might not look the way it did as a PDF (support.microsoft.com/en-us/office/opening-pdfs-in-word-1d1d2acc-afa0-46ef-891d-b76bcd83d9c8). A scanned page is not mostly text from the file's point of view. It is one large picture.

What recognition can read from a scanned page

Text recognition reads the pixels and reconstructs the characters. From a reasonably clean scan of a printed page it reads what an office actually needs: names with diacritics and non-Latin scripts such as Cyrillic, long identity and registration numbers, addresses, dates, amounts. Quality of the scan matters more than the software. Straight pages in even light read well; skewed, dark or low-resolution scans read worse. Handwriting, signatures and stamps remain unreliable for any recognition. The useful property of a poor scan is the direction it fails in: an unreadable value usually comes back blank instead of quietly wrong, and blank is a failure you can see.

Two routes from scan to Word

Route one is the traditional one: run recognition over the whole PDF first. Adobe Acrobat's documented Scan and OCR feature, called Recognize Text, processes each page and adds a searchable text layer to the scan (experienceleague.adobe.com, Acrobat scan and OCR tutorial). Free OCR tools do the same job with varying accuracy. Once the text layer exists, you can search the PDF, select values and paste them into your Word document one at a time. This route works, and for a single value from a single document it is quick. Its costs appear at volume: you copy each value individually, you fix the formatting each paste drags along, and every value passes through your hands, which is exactly where digits flip and values land under the wrong label.

Route two skips the intermediate copying entirely, and it is what Run Done does. You upload the Word document you already use, .doc or .docx, with no tags and no changes, and add the scanned PDF alongside any other materials: photos, other documents, pasted text, dictated details. The software reads the materials and enters the values in the right places in the document, with the formatting untouched. Values come only from the materials you supplied; a value the scan does not contain leaves its field empty rather than being invented. You review every entered value, correct anything, then download the finished Word file. A document you produce often can be saved and refilled in seconds.

Review is the step you keep

  • Check each entered value against the original scan, not against memory. The review is a comparison, not a re-read.
  • Give long numbers one careful digit-by-digit pass. That single pass replaces the retyping, the first proofread and the second proofread.
  • Treat every empty field as a question: is the value genuinely absent from the materials, or was the scan unreadable there?
Rescan before you fight a bad scan. Recognition quality follows scan quality, and thirty seconds at the scanner beats twenty minutes of correcting a read of a crooked, shadowed page. Flat page, even light, 300 dpi or a sharp phone photo filling the frame.

The short version

Check whether the PDF is digital or scanned by trying to select a word. Digital: copy directly. Scanned: the file is a picture, so recognition has to reconstruct the text, either by adding a text layer with an OCR tool and copying values out by hand, or by letting software read the scan and fill your Word document in one step. Either way, review the entered values against the original once, digits included, and treat empty fields as visible questions rather than silent failures.

See it on your own documents

Your first 50 fills are free. Upload a document you already use and the materials you have, and download the finished file.