OCR is not the finish line: from text to filled document

OCR turns a scan into text, but a filled document still needs three more steps: finding the values, placing them in your template, and verifying them.

OCR converts an image of a page into machine-readable text. That is all it does. Between that text and a finished document there are still three jobs: finding the values that matter in the text, placing each value in the right spot in your template, and verifying that what landed is correct. An office that buys a scanner with OCR has automated the first quarter of the work and kept the rest.

What does OCR actually give you?

OCR gives you characters, not meaning. It looks at the pixels of a scanned page and produces the letters and digits it recognizes, typically as plain text or a searchable PDF (ibm.com/think/topics/optical-character-recognition). The output of running OCR on a scanned contract is the contract's text as one long stream. Nothing in that stream is labelled. The software that produced it does not know which digits are a price, which are a parcel number, and which are a date.

This is worth stating plainly because the marketing around scanning blurs it. A searchable PDF is genuinely useful: you can find a name with Ctrl+F instead of leafing through paper. But searchable is not the same as filled. The distance between a text stream and a completed power of attorney is the same distance a person covers when they read a document and type values into another one. OCR moved the reading from paper to screen. The typing is still there.

Job one that remains: finding the values

A title deed contains hundreds of words. Your template needs perhaps fifteen values from it: the owner's name, the personal number, the parcel, the cadastral municipality, the surface, the basis of acquisition. Finding them means understanding labels and context. The string of thirteen digits after the name is the personal number; the similar string elsewhere is not. Two dates appear near each other and only one is the date of issue. A person does this by reading. Software can only do it with a layer above OCR that interprets the text, not just recognizes it. Without that layer, a human reads the OCR output and picks the values by hand, exactly as they would from paper.

Job two: placing the values in the document

Once found, each value has a destination: a specific gap in a specific sentence of your template. Placement sounds trivial and is where quiet errors live. The seller's number goes where the buyer's belongs. A date of birth lands in the field meant for the date of issue. The value is correct and the document is still wrong. Placement also has a formatting side: the value must arrive in the font and style of the surrounding sentence, not in whatever form the copy buffer carried. Copy-pasting from an OCR layer into Word routinely drags formatting along and leaves a visibly patched paragraph.

Job three: verifying what landed

OCR makes mistakes, and its mistakes are quieter than a typist's. A worn stamp turns an 8 into a 3. The letter O and the digit 0 trade places in a reference number. Recognition of a crumpled or low-contrast page degrades without announcing itself. So the final job is a check: every value that came off a scan gets compared against the source once before the document is used. The practical shape of that check matters. Verifying fifteen listed values against the source documents takes minutes. Proofreading a whole filled document against a whole scanned file takes much longer and misses more.

  • Compare each long number against the source image one digit at a time, and do it exactly once.
  • Prefer tools where a doubtful read produces a gap rather than a confident wrong value. A visible gap gets fixed; a plausible wrong digit gets signed.
  • Work from the list of entered values rather than re-reading the whole document.

What a complete pipeline looks like

Run Done covers the whole distance rather than the first step. You upload the Word document you already use, unchanged, and add the materials: photos, scans, PDFs, pasted text, or dictated details. The software reads the materials, finds the values, and enters each one in the right place in the document, with the surrounding formatting untouched. Values come only from the supplied materials; a value that cannot be found leaves its field empty rather than being guessed. You then review every entered value, correct anything, and download the finished Word file. The verification job stays, as it should, but it becomes a short list check instead of a retype and a proofread.

A quick test for any tool that advertises OCR or scanning: hand it a scan and your real template, and ask what happens next. If the answer is that you now have text to copy from, the tool does recognition. If the answer is a filled document with each value traceable to the source and open to correction, the tool does the job.

The short version

OCR turns an image into unlabelled text and stops. A filled document still needs the values found in that text, placed in the right spots in the template, and verified against the source. If your scanning setup ends at a searchable PDF, the typing has not disappeared; it has moved to a screen. Judge tools by whether they finish the distance, and keep one human check at the end regardless.

See it on your own documents

Your first 50 fills are free. Upload a document you already use and the materials you have, and download the finished file.