Five ways to split a stack
| Method | How it works | Prep per document | Breaks when |
|---|---|---|---|
| One scan per document | Each document is its own scan job | Start a scan each time | Never, but it is the slowest |
| Fixed page count | A new document every N pages | None | Any document has a different number of pages, or a page is missing |
| Separator sheets | A barcode sheet goes in front of each document | Insert one sheet | Someone forgets a sheet, or the sheet prints too small to read |
| Barcodes on the documents | A barcode already printed on page one (a form ID, a cover sheet) starts a document | None if the barcode is already there | Later pages carry a matching barcode too |
| Automatic | Software reads each page and decides where documents start | None | Two documents look like one continuous document; uncertain starts need a check |
Blank sheets as separators are a sixth, older trick. Avoid them if you scan two-sided: the back of every single-sided page is blank too, and blank-page removal in the scanner or software may discard your separators.
Which method fits your stack
Identical forms, every time
If every document really is the same length, such as one-page time sheets or two-page applications, page count separation is free and fast. Check the count at the end of each batch: if the total pages are not a multiple of N, something is off.
Mixed lengths, one kind of document
Invoices that run one to four pages, contracts with schedules, medical records. Separator sheets are dependable here, and so is automatic separation. If the documents already carry a barcode on page one, use it; it costs nothing.
Mixed kinds of document
The typical mailroom or accounts payable pile: invoices, statements, purchase orders, a W-9, a credit memo. Automatic separation was made for this, because the document types themselves mark where one ends and the next begins. Keep separator sheets nearby for the cases it cannot know, such as two invoices from the same vendor that should stay as one document.
Documents whose values you cannot read off the page
Photocopies, faxes and handwritten forms are hard to index. A separator sheet that carries the values for the document behind it (a claim number, an employee ID) indexes the document as it is scanned, instead of someone typing the values later.
Prepare the paper
Most separation problems start before the scanner. Five minutes of preparation per box pays for itself.
- Remove staples, clips and sticky notes. Repair or copy torn pages.
- Unfold and flatten. Put receipts and small slips on a full sheet or in a carrier sleeve.
- Face every page the same way, top edge first.
- Keep each document’s pages in order, with its first page on top.
- Insert separator sheets now if you use them, one in front of each document.
- Split the pile into loads your feeder takes comfortably, and keep the loads in order.
Scanner settings that help
These work for almost any production scanner with a document feeder.
| Setting | Recommended | Why |
|---|---|---|
| Resolution | 300 dpi (200 for clean typed pages, 400 for very small print) | Text recognition and barcode reading need enough detail; higher only adds file size |
| Color mode | Grayscale by default; color when color matters | Black and white makes the smallest files but can lose faint text, stamps and pencil |
| Sides | Two-sided (duplex) | Mixed stacks always have some printing on the back |
| Blank page removal | On, in one place only | Removes empty backs; two removals stacked can drop a nearly blank signature page |
| Double-feed detection | On | Two pages pulled together means a missing page and a wrong split |
| Deskew and auto-rotate | On if the driver offers it | Straight, upright pages read better |
How CapturePoint 6 does it
CapturePoint 6 scans from TWAIN scanners on a Windows PC or imports PDF, TIFF, JPEG, PNG, BMP and GIF files. It splits stacks into documents, recognizes each document’s type and reads its fields, all on the PC. During setup you can watch it reach the step “Finding where each document starts and ends”, with the note that barcodes and separator sheets count too.
Four choices under Document starts
In Job setup › Document starts you answer one question, “How should documents start?”, with one of four choices:
| Choice | What it does |
|---|---|
| Automatically | CapturePoint decides which pages belong together. |
| At matching barcodes | Starts a document at every barcode that matches your value or prefix. For example, DOC- matches DOC-001 and DOC-002. A new-looking page without a matching barcode stays with the page before. |
| Automatic + matching barcodes | Starts a document at every matching barcode, and also finds other starts by itself. |
| Only at separator sheets | Keeps pages together until a CapturePoint separator sheet appears. |
CapturePoint separator sheets work with every option, so you can always force a split by dropping a sheet in.
Print your own break sheets
Print break sheets makes two kinds. Plain sheets are identical and reusable and carry no values: print a stack once and keep it by the scanner. Sheets with values are each prepared for one document and carry its field values, so a photocopy or handwritten form is indexed the moment it is scanned; each is used once. Either kind works whichever way round it is fed. They print as a PDF so the barcode keeps its exact size on paper; print at actual size, not “fit to page”.
Review only the uncertain starts
Review thresholds has a bar for document starts. With it on, a document whose start CapturePoint is not sure enough about waits for a person instead of going out on a guess. You set how sure it must be; the Document starts screen links there under “Want fewer requests to review?”. Fixing a wrong split is a split or merge in the viewer. When you correct starts the same way several times, CapturePoint can offer a Pattern noticed suggestion so it splits that way by itself; nothing changes until you apply it, and Undo takes it back.
More than one feeder load
Each new scan or file selection adds to the open batch with a visible boundary before its first page, which protects work you already reviewed. If a load continued a document from the previous load, merge the two parts in review. Each separately imported PDF starts its own document.
A worked example: one accounts payable pile
Picture a morning’s mail for accounts payable: invoices of one to three pages, purchase orders, vendor statements, remittance advice, a W-9 and a credit memo, all in one pile. These are the six document types CapturePoint found in its sample accounts payable folder.
- Page count fails at the first two-page invoice, and every document after it is cut in the wrong place.
- Separator sheets work, at the cost of one sheet per document and someone inserting them.
- Automatic works with no preparation: a statement after an invoice looks different, and so does a new invoice after the last page of the previous one. The few starts it is unsure about wait for a quick check.
- Automatic plus a few sheets covers the edge cases: put a plain sheet in front of anything you know is unusual, such as an invoice with a stapled delivery note.
Once the stack is split, each document is named and filed by itself. Where to go next:
- Name and file scanned documents automatically
- Scan to SharePoint or OneDrive automatically
- Scan straight to Google Drive
Pitfalls to avoid
Barcodes on every page
Some forms print the same barcode on each page. If you separate at matching barcodes, every page becomes its own document. Match on a value or prefix that only appears on page one, or use separator sheets for those forms.
Separator sheets that will not read
Sheets printed shrunk to fit, photocopied many times, or printed low on toner are the usual causes. Reprint from the original PDF at actual size, and retire sheets that are creased or torn.
Forgetting what happens to the separator page
Some systems drop the separator page from the finished document and some keep it as page one. Check one finished document before you scan a whole box, so nobody is surprised later.
Reloading the feeder in the middle of a document
Split loads between documents, not inside one. If it does happen, join the two parts in review before export.
Trusting every split without a check
Any method can be wrong now and then. A setup that holds uncertain starts for a person is safer than one that files everything and hopes.
More help
- The CapturePoint help library has setup and troubleshooting articles.
- Running an older CapturePoint version? Read QCards: where documents start and end.