Articles / AI in your office

Why 99.5 percent accurate means one in five tables is wrong

Say a tool reads each box of a table right 99.5 times out of 100. That sounds like 1 miss in 200, which sounds fine. But you bill from the whole table, and the whole table has to be right.

By Alex Yeskolski, Founder, VuseDesk. 1,368 words, about 7 min at 200 words a minute, 7 sources, 1 table.

The math nobody shows you

One small miss rate turns big when you multiply it across a whole table.

The chance a table is fully right is the chance box one is right, times box two, times box three, and so on. By our own arithmetic, 0.995 multiplied by itself 44 times comes out to 0.802. So an 11 row by 4 column takeoff, at 99.5 percent per box, comes out clean about 80 times in 100. The other 20 times, by that same math, at least 1 box is wrong and you do not know which one.

That is the one in five in the title. It assumes each miss is its own roll of the dice, which is a shortcut, not a measurement. I come back to that below.

Sinha, Arun, Goel, Staab and Geiping wrote this as a formula in a paper called The Illusion of Diminishing Returns. They show the number of steps a model can chain before its odds drop below a target is the log of the target divided by the log of the per step accuracy. By our own arithmetic with that formula, a table becomes a coin flip at 138 boxes if each box is 99.5 percent, at 69 boxes if each box is 99 percent, and at 693 boxes if each box is 99.9 percent.

What 99.5 percent turns into

Here is the same math for the tables a contractor actually touches.

Cells in the tableExampleWhole table right at 99.5 percent per cellAt 99 percentAt 99.9 percent
10A five line estimate, item and price95.1 percent90.4 percent99.0 percent
20A ten line invoice, item and price90.5 percent81.8 percent98.0 percent
44Eleven rows by four columns, a material takeoff80.2 percent64.3 percent95.7 percent
100A change order log for one month60.6 percent36.6 percent90.5 percent
200A subcontractor pay app36.7 percent13.4 percent81.9 percent
500A year of job costs8.2 percent0.7 percent60.6 percent
Coin flip pointBoxes where the odds hit 50 percent138 cells69 cells693 cells

Basis: our own arithmetic, the per cell rate raised to the number of cells, with the coin flip row from the Sinha et al. horizon formula. Each miss is treated as independent of the others.

Read the 500 row again. A year of job costs at 99.5 percent per box comes out fully right about 8 times in 100, by that same arithmetic. That is a tool you can trust with a line, and then a person checks the page.

Real tools on real paper

Tests on real invoices and receipts show the same gap between the box and the whole.

Yashwant, Dubey, Paikray and Thulsiram ran 102 real invoices, some scanned and some digital, through two readers in a paper called Invoice Information Extraction. In that test the open source reader Docling got 63 percent of fields right overall, 58 percent on line items, and 20 percent of the invoices failed a simple check that the lines add up to the total. The commercial reader LlamaExtract got 94 percent overall and 91 percent on line items on the same 102 invoices, and 5 percent still failed the math check. Even the good one sent 1 invoice in 20 back to a person in that test.

Anvari and Athitsos tested a 7 billion parameter model on receipts in a paper called From Pixels to Pairs. With clean text it scored 0.97 on a field by field measure but only 0.77 of receipts were fully right, and with real scanner output that fell to 0.83 on fields and 0.45 of receipts fully right. On a second receipt set in the same paper the model found 99.5 percent of the labels, yet only 69 percent of label and value pairs were exactly right with clean text, and 57 percent with real scans.

Microsoft researchers Smock, Pesala and Abraham saw the same thing on table layouts in a paper called Aligning benchmark datasets for table structure recognition. Whole tables came out fully right 42 to 69 percent of the time before they cleaned the test set, and 65 to 81 percent after, while box level scores stayed high through both. They wrote that in real settings single cell errors may not be tolerable, so whole table accuracy has to be reported. Ask a vendor for that number.

Small mistakes pile up

Machines that work in steps get worse as the steps stack, and old scanning research says the fix is a human queue.

Meyerson and colleagues at Cognizant AI Lab and UT Austin put it plainly in a paper called Solving a Million-Step LLM Task with Zero Errors. With a 1 percent miss rate per step, they show a system is expected to fail by step 100 of a million step job. In their tests, models solved a puzzle called Towers of Hanoi well up to 5 or 6 disks, which is 31 to 63 moves, and then success fell to zero unless the job was broken into small pieces.

Sinha et al. found something worse. In their tests, even the best open model they tried, Qwen3-32B, fell below 50 percent accuracy within 15 turns of a chained task, even though its first step was nearly perfect. A model that sees its own earlier mistake becomes more likely to make another. Bigger models did not fix that in their tests, but reasoning mode did.

Rose Holley at the National Library of Australia measured raw character accuracy on 45 pages of old newspapers and got a spread from 71 percent to 98.02 percent, and her program called 98 to 99 percent good. She also noted published accuracy claims run from 99.8 percent, counted by word after people fixed 700,000 pages, down to 68 percent by character with no fixing. Same word, two different things, so ask what was counted and who fixed it.

What we got wrong

The headline is true at 44 boxes and too harsh below it.

A ten box table at 99.5 percent per box is fine about 95 times in 100 by the same arithmetic, so small tables are safer than the title sounds. And misses are not really independent. Holley's spread from 71 to 98.02 percent tracked page quality, so bad scans wreck a few documents while clean ones sail through.

For a stack of paper, your clean pages may be better than my math and your bad pages far worse. For a machine chaining steps, Sinha et al. show errors feed on themselves, so my math is too kind. Either way, the fix is the same.

What stays human

Someone still has to catch the wrong box, and the research says that is the whole job.

Yashwant et al. caught 5 to 20 percent of their 102 invoices with one question, do the lines add to the total. That check is a person's job, or a rule a person sets. A NIST study by Geist and Wilkinson claims error rates only fall when the tool rejects the characters it is unsure of and hands them to a person. I did not read their curves, so take that as their claim.

So set up a review queue. Every table the tool reads gets a math check, and any table that fails goes to a person before it is billed or paid. VuseDesk records payments and never moves money, so a wrong box on our side is a fix, not a loss.

Have us check your tables

Book the audit and we will go through the tables your business lives on, from takeoffs to pay apps, and show you which ones need a human check. You get the math in full, whether you buy anything or not. If you want it set up after, that is what we are here for.

Alex Yeskolski Founder, VuseDesk. He writes the software these articles describe. More about the author
Walk the tour, one week of work

Or call (252) 666-7217 and ask the 1 question this article did not answer. Email [email protected] if you would rather write it down.