What a 98% accuracy number actually means
Published
A percentage on its own is not a claim
Every vendor selling document extraction leads with an accuracy figure. The figures cluster in the same place, somewhere in the nineties, and they are almost never comparable to each other, because the number by itself doesn’t say what was measured, on what, or by whom.
That gap is where the buying mistake happens. Two systems quoting the same percentage can behave completely differently on your documents, and nothing on either vendor’s page lets you tell which is which.
The useful version of an accuracy claim always comes with the measurement attached. On the SEC filing work we published, accuracy was measured against filings that had been verified by hand. That sentence does more work than the percentage does. It says the reference set was real output somebody checked line by line, not a sample the system was tuned on and not an internal estimate of how things seem to be going.
Ask for that sentence. If a vendor can’t produce it, the number was an impression rather than a measurement.
Coverage decides whether the number helps you
The second question is what the percentage covers, and this is where most headline figures quietly fall apart.
Take a system that reads financial filings. There are five statement types in play. The income statement is the easiest to parse and the one every tool handles well. A vendor can score 98% on that alone and print 98% on the website without technically lying.
Your analyst still has to open the PDF, because the balance sheet is the one that went sideways.
That is the whole problem with a headline number: it averages away the place the system is weak, and the weak place decides whether anyone can trust the output without checking it. A tool that is excellent on four document variants and unreliable on the fifth has not removed the manual step from your process. It has moved the manual step somewhere less predictable, which is worse, because now nobody knows which outputs need a second look.
The bar worth setting is coverage without a soft spot. On the filing project that meant holding across all five statement types rather than only the easy one, and the accuracy figure was quoted that way for exactly that reason.
Why “up to” is the more honest form
An accuracy figure quoted flat reads as stronger. It is also the version more likely to break on contact with your documents.
Extraction accuracy depends on the source material. Cleaner documents score higher. A set drawn from one issuer with a stable template scores higher than a set drawn from forty issuers who share no template at all. A number produced on a favourable set describes that set, and it will keep describing that set right up until the moment you send in something the system has never seen.
Publishing a figure as “up to 98%” is a statement about that dependency. It says the number holds under the conditions measured and is not a promise about arbitrary input. A vendor who quotes a number they can stand behind on a fresh set of documents is telling you something more useful than a vendor quoting their best day.
So ask the follow-up: what did the number look like on the worst set you tested?
Where the residual error sits
A residual share of wrong figures sounds small enough to accept when it is quoted as a percentage. Whether it actually is depends on something the percentage cannot express: what the wrong figures look like.
An error that breaks loudly costs almost nothing. A cell comes out empty, a total refuses to parse, an export fails, and somebody fixes it in a minute.
An error that looks plausible is the expensive one. A revenue figure that lands in the wrong row, is off by a factor of ten, or belongs to the prior period will pass every human glance it gets, because nobody re-checks a number that looks about right. It flows into a model, then into a comparison, then into a conclusion somebody acts on. The cost of that error has nothing to do with how large the residual share is and everything to do with what was built on top of it.
Which is why validation on the way out matters more than the last point of accuracy. Totals that don’t reconcile and figures that contradict each other should be surfaced rather than exported quietly. A system that flags its own uncertain output is more useful than a system scoring slightly higher that hands you everything with equal confidence.
What to ask before you sign
Four questions, and they are all about the same thing: making the number specific enough to be checkable.
- Measured against what set, and who verified it by hand?
- Covering which document variants, and what was the score on the weakest one?
- Does the figure describe a tuned set or a fresh one?
- What happens to the cases the system gets wrong: are they flagged, or do they land in the output looking like everything else?
A vendor who has done the work answers all four without hesitating, because they had to answer them internally before they could publish anything. A vendor who deflects is quoting a number they have not interrogated, and you will be the one who finds out what it was hiding.
For a worked example of what a measured number looks like in practice, including where it holds and what we deliberately don’t claim alongside it, see the SEC filing extraction case.