POClerk · Published benchmark
Every line graded, and the grading published.
An app with no reviews yet is asking you to trust it with order entry. Rather than assert accuracy, here is the whole test: what was fed in, how each line was graded, what the current engine scores, and how two established Shopify PO apps did on the same documents.
Alternatives anonymised — we grade results, not rivals.
The current run
POClerk does not ship unless a fresh, never-seen set of purchase orders clears four gates. This is the most recent
run, on corpus seed 20260721 — 436 lines across 4 real store catalogs and
13 document formats.
What the two headline numbers mean. Lines entered correctly counts a line as correct when the right product ended up on the order — or when POClerk correctly refused to enter it and said why. Both outcomes are right; a held line is a working outcome, not a failure.
Silent errors counts the opposite: a line entered confidently, with nothing flagged, that was wrong. These are the ones that ship the wrong goods, because nobody is looking. It is the number we optimise hardest and the one we would most like you to compare against anything else you are considering.
Where it is strong and where it is weakest
Averages hide the hard cases, so here is the same run broken out. The worst catalog matters more than the average: it is the closest thing to a guarantee about a catalog we have never seen — yours.
By catalog
Four real store catalogs, 24,114 product variants between them
Weakest: Apparel — size and colour variants at 98.0%. Variant matrices — the same garment in a dozen sizes and colours, abbreviated on the PO — are the hardest real problem in this job, and they are where held lines concentrate.
All 13 document formats, line by line
| Format | Lines | Entered correctly | Silent errors |
|---|---|---|---|
| PDF — classic form | 14 | 100.0% | 0.0% |
| PDF — modern layout | 17 | 100.0% | 0.0% |
| PDF — dense table | 35 | 100.0% | 0.0% |
| PDF — multi-page | 131 | 100.0% | 0.0% |
| Excel | 32 | 100.0% | 0.0% |
| Excel — order on a later sheet, price-list decoy | 30 | 100.0% | 0.0% |
| CSV — ERP export | 30 | 100.0% | 0.0% |
| Plain email body | 34 | 100.0% | 0.0% |
| Forwarded email thread | 27 | 100.0% | 0.0% |
| HTML email table | 31 | 96.8% | 0.0% |
| pdf-quickbooks | 22 | 95.5% | 4.5% |
| Scan / phone photo | 17 | 94.1% | 0.0% |
| PDF — word-processor form | 16 | 93.8% | 0.0% |
Against two established PO apps
On 2026-07-20 we installed two established purchase-order apps from the Shopify App Store on a store loaded with the same catalogs, fed them purchase orders from this corpus by hand, and graded whatever they produced with the same A–F rubric we grade ourselves with.
They were not given identical sets, so they do not share a chart. One of the two accepts email only — it has no upload interface — so it could only be given the purchase orders that arrive as email. The other accepts uploads and was given every document in the bundle. Putting all three on one bar chart would imply a like-for-like test that did not happen, so each alternative is compared separately against our grades on exactly the purchase orders and lines that alternative was given.
Every percentage below shares one denominator: ground-truth lines. Items an app added to a draft that no line of the PO asked for are excluded from the percentage and counted separately, so nobody is measured against a different yardstick.
POClerk vs Alternative A
5 purchase orders · 61 graded lines · Email intake only — it has no upload interface, so it could only be given the POs that arrive as email.
Lines entered correctly higher is better
Silent errors lower is better — wrong lines that looked fine
POClerk vs Alternative B
10 purchase orders · 144 graded lines · Accepts uploads, so it was given every PO in the bundle.
Lines entered correctly higher is better
Silent errors lower is better — wrong lines that looked fine
Plus 2 lines added to drafts that no purchase order asked for, excluded from the percentages above.
Being fair to them. These numbers come from the corpus that was live in July 2026, and our own score on those same documents (96.7% and 98.6%) is lower than our 99.1% headline, because the head-to-head bundle was deliberately weighted toward the hardest documents — scans and cryptic part numbers. We have also improved the engine since that date, and they may well have improved theirs. Both apps were used on their free tiers, as any merchant would first meet them; one of them crashed outright on the scanned purchase order even after a retry, and those lines are graded as missed rather than excluded.
Method, in full
Synthetic orders, real catalogs
The purchase orders are generated, so no merchant's real order data is exposed by publishing results. The catalogs underneath are real store catalogs — 24,114 variants, including one with 20,000+ cryptic part numbers, which is harder than the shops we expect to serve.
Never-seen documents
Each run generates a fresh corpus from a new random seed. Nothing is tuned per catalog, and the engine has not seen the documents it is graded on. Results on a corpus we had already mined are not published as gate results.
Deliberate traps
The corpus contains lines that should not be entered: services that are not products, items that do not exist in the catalog, and genuinely ambiguous descriptions. Entering one of those confidently is graded as a silent error, the same as any other wrong line.
Graded by code, not judgement
Every line is graded automatically against a known answer, using the same scoring code for us and for the alternatives. Auto-created drafts are graded exactly like reviewed ones — a wrong line nobody flagged counts as silent whether a human saw it or not.
What this does not tell you. These figures describe this corpus, not a promise about your catalog. Product names, code conventions and how your buyers write their orders all matter, and a catalog full of near-identical variants is genuinely harder than one with distinct names. That is exactly what the free first 10 orders are for: run your own real POs through it and check every line against the original, side by side, before you pay anything.