POClerk · Published benchmark

Every line graded, and the grading published.

An app with no reviews yet is asking you to trust it with order entry. Rather than assert accuracy, here is the whole test: what was fed in, how each line was graded, what the current engine scores, and how two established Shopify PO apps did on the same documents.

Alternatives anonymised — we grade results, not rivals.

The current run

POClerk does not ship unless a fresh, never-seen set of purchase orders clears four gates. This is the most recent run, on corpus seed 20260721 — 436 lines across 4 real store catalogs and 13 document formats.

Lines entered correctly 99.1% gate ≥ 90% · pass
Silent errors 0.2% gate ≤ 2% · pass
Worst single catalog 98.0% gate ≥ 80% · pass
Invented products 0 gate 0 · pass

What the two headline numbers mean. Lines entered correctly counts a line as correct when the right product ended up on the order — or when POClerk correctly refused to enter it and said why. Both outcomes are right; a held line is a working outcome, not a failure.

Silent errors counts the opposite: a line entered confidently, with nothing flagged, that was wrong. These are the ones that ship the wrong goods, because nobody is looking. It is the number we optimise hardest and the one we would most like you to compare against anything else you are considering.

Where it is strong and where it is weakest

Averages hide the hard cases, so here is the same run broken out. The worst catalog matters more than the average: it is the closest thing to a guarantee about a catalog we have never seen — yours.

By catalog

Four real store catalogs, 24,114 product variants between them

Coffee roaster
100.0%
Beverages
100.0%
Vehicle parts
98.3%
Apparel
98.0%

Weakest: Apparel — size and colour variants at 98.0%. Variant matrices — the same garment in a dozen sizes and colours, abbreviated on the PO — are the hardest real problem in this job, and they are where held lines concentrate.

All 13 document formats, line by line
FormatLinesEntered correctlySilent errors
PDF — classic form14100.0%0.0%
PDF — modern layout17100.0%0.0%
PDF — dense table35100.0%0.0%
PDF — multi-page131100.0%0.0%
Excel32100.0%0.0%
Excel — order on a later sheet, price-list decoy30100.0%0.0%
CSV — ERP export30100.0%0.0%
Plain email body34100.0%0.0%
Forwarded email thread27100.0%0.0%
HTML email table3196.8%0.0%
pdf-quickbooks2295.5%4.5%
Scan / phone photo1794.1%0.0%
PDF — word-processor form1693.8%0.0%

Against two established PO apps

On 2026-07-20 we installed two established purchase-order apps from the Shopify App Store on a store loaded with the same catalogs, fed them purchase orders from this corpus by hand, and graded whatever they produced with the same A–F rubric we grade ourselves with.

They were not given identical sets, so they do not share a chart. One of the two accepts email only — it has no upload interface — so it could only be given the purchase orders that arrive as email. The other accepts uploads and was given every document in the bundle. Putting all three on one bar chart would imply a like-for-like test that did not happen, so each alternative is compared separately against our grades on exactly the purchase orders and lines that alternative was given.

Every percentage below shares one denominator: ground-truth lines. Items an app added to a draft that no line of the PO asked for are excluded from the percentage and counted separately, so nobody is measured against a different yardstick.

POClerk vs Alternative A

5 purchase orders · 61 graded lines · Email intake only — it has no upload interface, so it could only be given the POs that arrive as email.

Lines entered correctly higher is better

POClerk
96.7%
Alternative A
60.7%

Silent errors lower is better — wrong lines that looked fine

POClerk
3.3%
Alternative A
18.0%

POClerk vs Alternative B

10 purchase orders · 144 graded lines · Accepts uploads, so it was given every PO in the bundle.

Lines entered correctly higher is better

POClerk
98.6%
Alternative B
59.7%

Silent errors lower is better — wrong lines that looked fine

POClerk
1.4%
Alternative B
2.8%

Plus 2 lines added to drafts that no purchase order asked for, excluded from the percentages above.

Being fair to them. These numbers come from the corpus that was live in July 2026, and our own score on those same documents (96.7% and 98.6%) is lower than our 99.1% headline, because the head-to-head bundle was deliberately weighted toward the hardest documents — scans and cryptic part numbers. We have also improved the engine since that date, and they may well have improved theirs. Both apps were used on their free tiers, as any merchant would first meet them; one of them crashed outright on the scanned purchase order even after a retry, and those lines are graded as missed rather than excluded.

Method, in full

Synthetic orders, real catalogs

The purchase orders are generated, so no merchant's real order data is exposed by publishing results. The catalogs underneath are real store catalogs — 24,114 variants, including one with 20,000+ cryptic part numbers, which is harder than the shops we expect to serve.

Never-seen documents

Each run generates a fresh corpus from a new random seed. Nothing is tuned per catalog, and the engine has not seen the documents it is graded on. Results on a corpus we had already mined are not published as gate results.

Deliberate traps

The corpus contains lines that should not be entered: services that are not products, items that do not exist in the catalog, and genuinely ambiguous descriptions. Entering one of those confidently is graded as a silent error, the same as any other wrong line.

Graded by code, not judgement

Every line is graded automatically against a known answer, using the same scoring code for us and for the alternatives. Auto-created drafts are graded exactly like reviewed ones — a wrong line nobody flagged counts as silent whether a human saw it or not.

What this does not tell you. These figures describe this corpus, not a promise about your catalog. Product names, code conventions and how your buyers write their orders all matter, and a catalog full of near-identical variants is genuinely harder than one with distinct names. That is exactly what the free first 10 orders are for: run your own real POs through it and check every line against the original, side by side, before you pay anything.

Test it on your own POs — 10 free See a PO run first →