Skip to content
Guides

Guide · Records and scanning

Searchable PDF vs PDF/A: which to use for scanned records

Both are PDFs, and a file can be both. One is about finding text; the other is about opening the file correctly in twenty years. Here is the difference, and a simple rule for which to use.

For office managers, records managers and IT admins deciding how scanned invoices, contracts and files should be saved.

The short answer

Three kinds of scanned PDF

What is inside

Image-only PDF:
A picture of each page.
Searchable PDF:
The picture, plus an invisible text layer from OCR.
PDF/A (searchable):
The picture and text layer, saved under the PDF/A rules.

Search and copy text

Image-only PDF:
No
Searchable PDF:
Yes
PDF/A (searchable):
Yes

Indexed by Windows search and document systems

Image-only PDF:
No
Searchable PDF:
Yes
PDF/A (searchable):
Yes

Built for long-term preservation

Image-only PDF:
No
Searchable PDF:
Not specifically
PDF/A (searchable):
Yes (an ISO standard)

Can be password-protected

Image-only PDF:
Yes
Searchable PDF:
Yes
PDF/A (searchable):
No

Typical use

Image-only PDF:
Nothing, ideally
Searchable PDF:
Working files: AP, HR, day-to-day records
PDF/A (searchable):
Long-term and permanent records, archives, court and agency filings

How a searchable PDF works

When a scanned page goes through OCR, the software finds each word and writes it into the PDF as invisible text, placed exactly over the word in the picture. You still see the original scan; behind it sits text you can search, select and copy. People sometimes call this “PDF with a hidden text layer” or “image plus text”.

Two things follow from that:

  • The text is only as good as the OCR. A crooked, faint or low-resolution scan produces a text layer full of errors, and search quietly misses documents. Scan text documents at 300 dpi, straight, and check a few files.
  • “Searchable PDF” is a description, not a standard. Two searchable PDFs can be built very differently. That is fine for working files, and it is exactly the gap PDF/A fills for records.

What PDF/A adds

PDF/A is the ISO 19005 family of standards: a restricted form of PDF designed for long-term archiving. The idea is simple. Everything needed to display the document must be inside the file, and nothing may depend on outside software, servers or passwords. In practice, a PDF/A file:

  • Embeds every font it uses

    So text renders identically on any computer, now or later.
  • Defines its colors in a device-independent way

    So the page looks the same on any screen or printer.
  • Carries standard metadata

    Title, author, dates and the PDF/A version, readable by archive systems.
  • Contains no encryption or passwords

    A future reader must be able to open it.
  • Contains no JavaScript, audio, video or links to outside content

    Nothing that might stop working or change what is shown.

None of this changes how the page looks today. It changes how likely it is to look the same when someone opens it in fifteen years.

PDF/A versions and conformance levels

You will see names like PDF/A-1b or PDF/A-2u. The number is the part of the standard; the letter is the level of conformance.

PDF/A-1

Based on:
PDF 1.4
Worth knowing:
The original (2005). Very widely accepted, but no transparency or JPEG 2000 image compression.

PDF/A-2

Based on:
PDF 1.7
Worth knowing:
Adds JPEG 2000 image compression, transparency and layers. A sensible default for scanned records.

PDF/A-3

Based on:
PDF 1.7
Worth knowing:
Like PDF/A-2, but allows any file to be attached, such as the XML data in some e-invoice formats.

PDF/A-4

Based on:
PDF 2.0
Worth knowing:
The newest part. Simplified levels; support is still growing.

b (basic)

Means:
The page will display reliably.
In practice:
Enough for appearance; text may not map to real characters.

u (Unicode)

Means:
Level b, plus every piece of text maps to real characters.
In practice:
Search and copy stay reliable. A good fit for searchable scans.

a (accessible)

Means:
Level u, plus a tagged structure for screen readers.
In practice:
Usually needs born-digital documents; hard to achieve from scans.

If nobody has told you which to use, PDF/A-2u is a safe choice for scanned business documents: a modern base, reliable display and a searchable text layer that maps to real characters.

Which to use when

AP invoices and receipts in active use

Use:
Searchable PDF (PDF/A optional)
Why:
You need to find them quickly; most are kept for a fixed number of years.

HR files, contracts, deeds, board minutes

Use:
PDF/A
Why:
Kept for many years or permanently.

Records you will transfer to an archive or agency

Use:
PDF/A (the version they ask for)
Why:
Many archives and agencies list PDF/A among accepted or preferred formats.

Court filings

Use:
Whatever the court specifies
Why:
Requirements vary by court; some ask for PDF/A.

Documents that must be password-protected

Use:
Searchable PDF
Why:
PDF/A forbids encryption; or store PDF/A in a secured system instead.

E-invoices with embedded XML

Use:
PDF/A-3
Why:
It is the part of the standard that allows attached data files.

How long to keep each kind of record is a separate decision from the format. Our retention schedules give typical periods to discuss with your accountant or counsel.

Common misunderstandings

  • “PDF/A makes a file tamper-proof.” It does not. A PDF/A can be edited like any PDF. Integrity comes from where you keep it: versioning, permissions and an audit trail in a document management system.
  • “PDF/A means searchable.” Not by itself. An image-only scan can be valid PDF/A-1b and contain no text at all. Ask for OCR and a u level.
  • “Saving as PDF/A is enough.” Some tools label files PDF/A that do not pass validation. Check a sample.
  • “PDF/A files are huge.” Embedded fonts and color profiles add a little. For scans, resolution and color mode decide the size far more than PDF/A does.

How to check a file

  • Is it searchable? Open it and search for a word you can see on the page, or try to select a line of text.
  • Does it claim to be PDF/A? Adobe Acrobat and Reader show a standards notice when a file declares PDF/A, and the document properties list the version.
  • Does it really conform? Run a sample through a validator. veraPDF is the free, open-source one widely used by archives.

Getting both from your scanner with CapturePoint 6

CapturePoint 6 makes the searchable PDF part automatic. It reads every page on your own Windows PC, splits the stack into separate documents, and exports each one as a searchable PDF, together with a data file holding every captured value and a plain text file. Turn on archive-grade PDF/A (PDF/A-2u) for jobs that hold records.

CapturePoint 6 export settings: each document arrives as a searchable PDF, a data file and a text file. Archive-grade PDF/A is under More.

Files are named and filed into folders from the captured values, or sent into Content Central, where retention, legal holds, versions and the audit trail take care of the integrity side. Setting up retention there is covered in the help article on document retention policies. For the rest of the capture setup, see the AP document capture checklist.

Questions

Is PDF/A required for tax or accounting records?

Usually not. Most retention rules say how long to keep a record and that it must stay complete and readable; few name a file format. Some archives, courts and agencies do ask for PDF/A when records are filed or transferred to them, so check the requirement that applies to you.

Can a PDF/A file be searchable?

Yes. A scanned page saved as PDF/A can carry the same invisible text layer as any searchable PDF. Levels ending in u (such as PDF/A-2u) also require that the text maps to real characters, which keeps search and copy reliable.

Can I password-protect a PDF/A?

No. PDF/A does not allow encryption, because a future reader might not have the password. Control access through where the file is stored instead: folder permissions or a document management system.

Should I convert my old scans to PDF/A?

Only the ones you must keep for a long time, and keep the originals until you have checked the converted files. For everyday working files, a searchable PDF is enough.

Try it on your own documents

Scan once. Get a searchable PDF, PDF/A if you want it, and the data.

CapturePoint 6 splits the stack into documents, reads them on your own Windows PC and saves each one as a searchable PDF, with archive-grade PDF/A one switch away, plus a data file and a text file, named and filed for you.

Windows 10 and 11 (64-bit). No sign-up and no credit card; sample jobs included. Priced per scanning station, with unlimited scanning. Get pricing

CapturePoint 6 review screen: a sample invoice beside its extracted fields and line items, with line 1 flagged because 4 at 35.00 was read as 141.00