Skip to content

Guide · Document data extraction

How to extract structured data from PDFs

Manual entry, templates, learned extraction, desktop software or a cloud service: what each approach is good at, where each one breaks, how to test a tool on your own documents, and when Ademero's Paige or CapturePoint 6 fits.

Practical guide · about 10 minutes to read · updated October 2026

First, what kind of PDF do you have?

"PDF" covers three very different things, and the right approach depends on which one is on your desk. Open a typical file and try to select a line of text.

KindDigital (born digital)
How to tellText selects and copies cleanly.
What it meansThe words are already inside the file. Extraction is about finding the right values, not reading them.
KindScanned (image only)
How to tellNothing selects, or the whole page selects as one picture.
What it meansIt needs OCR first. Scan quality now matters as much as the extraction method.
KindFillable form
How to tellBoxes you can click and type into.
What it meansThe values may be stored as form fields you can export directly, with no reading at all.

Most real piles are mixed: digital invoices emailed by large vendors, scans from small ones, and the occasional phone photo saved as a PDF. Plan for the worst of them, not the best.

Decide what "structured" means for you

Before comparing tools, write down five things. They decide the answer more than any feature list.

  1. The fields. The exact values you need, for example vendor, invoice number, date, PO number and total. Fewer is better.
  2. Tables. Whether you need line items (description, quantity, unit price, amount) or only header values. Line items are much harder.
  3. Variety. How many different layouts you receive. Five senders is a different problem from five hundred.
  4. Volume and timing. Twenty a month, or two thousand a day, and whether data is needed in minutes or by month end.
  5. The destination. A spreadsheet, an accounting import, a document library, or another system through an API.

Four approaches, compared honestly

ApproachManual entry
Good atNo setup, any document, a person understands context.
Breaks whenVolume grows. Errors rise with fatigue, and it does not scale.
FitsA few dozen documents a month.
ApproachCopy, paste and PDF-to-text tools
Good atFree or cheap for digital PDFs.
Breaks whenScans, tables (columns scramble) and anything repeated daily.
FitsOne-off jobs on digital files.
ApproachTemplates and zonal OCR
Good atFixed forms that never change: your own forms, government forms.
Breaks whenA sender moves a field, adds a line or sends a new layout. Every layout is a template to keep up.
FitsFew layouts, high volume, stable forms.
ApproachLearned extraction
Good atMany layouts. Learns fields and tables from examples and corrections.
Breaks whenYou skip review entirely. Good tools tell you when they are unsure; use that.
FitsMixed documents from many senders.

Manual entry is not always wrong

If you process thirty documents a month and they go into one system, typing is honest and cheap. The case for software starts when the volume, the variety or the cost of mistakes grows, or when the same people keep retyping the same kinds of values.

Templates are reliable until they are not

Zonal OCR reads a value from a fixed box on the page. On a form you control, that works well. On invoices or freight bills from hundreds of senders it turns into a maintenance job: every new layout needs a template, and every layout change quietly breaks one.

Learned extraction needs a feedback loop

Learned tools work out where the fields are from a set of examples, then improve from the corrections people make. The two things to insist on: they must say how sure they are and why a document needs a look, and corrections must actually make the next documents better.

Desktop software or a cloud service?

Separate from how the data is extracted is where it happens. Neither is better in general; they suit different teams.

QuestionWhere documents are read
Local desktop softwareOn a PC you control
Cloud service or APIOn the provider’s servers
QuestionPaper scanning
Local desktop softwareNatural fit: scanner on the same PC
Cloud service or APIPossible, but documents are uploaded first
QuestionDocuments arriving by email or from systems
Local desktop softwareSaved to a folder, then processed
Cloud service or APINatural fit: send them straight in
QuestionGetting data into other systems
Local desktop softwareFiles and folders, or connected destinations
Cloud service or APIDownload, SFTP or webhook into your workflow
QuestionIT and privacy reviews
Local desktop softwareSimpler: images are not sent out to be read
Cloud service or APIReview the provider and its hosting
QuestionCapacity
Local desktop softwareLimited by the PC
Cloud service or APIGrows with the service

A worked example: one invoice, start to finish

Here is a fictional sample invoice as CapturePoint 6 extracts it. The fields read from the header were Northbridge Office Supply, invoice INV-2401, dated 09/04/2026, due 10/04/2026, PO-1031, total 480.00. The table has four lines.

Look at line 1. Its amount was captured as 141.00, but 4 times 35.00 is 140.00. Rather than pass that on, the line is flagged: "Line 1 does not add up. Correct it, or mark it right as printed." This is the kind of check that separates extraction you can trust from extraction you have to re-check by hand.

CapturePoint 6 review: fields and line items beside the page, with a line-item math check flagging line 1.

Once confirmed, the invoice leaves as a searchable PDF (optionally PDF/A), a data file with every captured value and a text file, named and filed by the fields you choose.

CapturePoint 6 export: the files each document can arrive as.

How to evaluate any extraction tool

Demos use clean documents. Your decision should be based on yours. Run every candidate through the same test.

An evaluation you can run in a week

  • Collect 50 to 100 real documents, including the worst: faxes, skewed scans, multi-page invoices, new senders.
  • Write down the correct value of every field you need for each one before you start. That is your answer key.
  • Measure field by field, not document by document. A document with one wrong total is a wrong document.
  • Count how often the tool says it is unsure, and how often it is wrong without saying so. The second number matters most.
  • Test line items separately: row counts, columns and whether rows add up to the total.
  • Include a stack scanned as one file and check that it is split into the right documents.
  • Time a person reviewing 25 documents. Review speed is where the real cost is.
  • Export to your real destination and confirm the next system takes it without retyping.
  • Correct a few mistakes and run similar documents again. Did it learn?

When Paige fits, and when CapturePoint 6 does

Ademero makes one of each kind, so we can be straight about which suits you. Both learn from a few examples, split stacks into documents, sort them by type, and read fields and line items.

What it is
PaigeA cloud service: send in documents, get clean data back
CapturePoint 6A Windows 10/11 (64-bit) app on your own PC
Where reading happens
PaigeOn Google Cloud
CapturePoint 6Locally on the PC, graphics card optional
How documents get in
PaigeSent to the service
CapturePoint 6TWAIN scanners, or PDF, TIFF, JPEG, PNG, BMP and GIF files
How results come out
PaigeDownload, SFTP or webhook
CapturePoint 6Folders (searchable PDF, data file, text file) or Content Central; also SharePoint/OneDrive, Google Drive, Dropbox or Nucleus One
Best for
PaigeDocuments that arrive digitally and should flow into other systems
CapturePoint 6Paper scanned in-house, and teams that want documents read on their own PCs
How to try it
PaigeStart free on paige.app
CapturePoint 6Free trial starts on first launch, no form
Paige: invoices with their fields read, ready to export.

Neither is the answer for a one-off batch of twenty digital PDFs (copy them by hand) or for a fillable form whose values you can export directly. Help for each is in the Ademero help library: help.ademero.com/paige and help.ademero.com/capturepoint.

Specific documents

PDF extraction questions

Can I extract data from a scanned PDF?

Yes, but it needs OCR (optical character recognition) first, because a scanned PDF is a picture of a page with no text inside. Capture tools run OCR as part of extraction. Scan at a resolution your tool recommends, straight and clean, and accuracy follows.

How do I get tables and line items out of a PDF?

Copying a table from a PDF viewer usually scrambles the columns. Use a tool that extracts tables as rows and columns, and check that the rows add up to the total. Line items are where most tools differ, so test them on your own longest and messiest documents.

What output format should I ask for?

Whatever the next system reads without retyping: a CSV or spreadsheet for imports, JSON for developers and webhooks, or a searchable PDF with the values attached for a document library. Ask for the original file to stay linked to the data, so anyone can check a value against the page.

Do I need a template for every layout?

With template or zonal OCR tools, yes, and each template needs fixing when a sender changes their layout. Learned extraction tools work out the fields from examples instead, which is why they suit documents from many different senders.

Next step

Try the one that fits your paper.

Documents arriving digitally? Start free with Paige. Paper on a desk or in a mailroom? Download the CapturePoint 6 free trial and try a ready-made sample job.