Your vendor's camera is “96% accurate.” That's 40 innocent shoppers before lunch
Your vendor's camera is “96% accurate.” That's 40 innocent shoppers before lunch
Every theft-detection camera quotes an accuracy number. It's the most honest-looking lie in the category — because at the rate real shops run at, a “96% accurate” camera is wrong about nine of every ten people it stops. There's a calculator; move the sliders.
Why "accuracy" is the number vendors lead with precisely because it hides what the camera does to your customers. There's a calculator below — move the sliders.
The short version: at the base rate a real shop runs at, a theft-detection camera a vendor can honestly call "96% accurate" still flags roughly 40 innocent shoppers before lunch, and is right about one time in ten when it points at someone. "Accuracy" is engineered to hide that. The number that actually matters — the one they never quote — is precision: when the camera accuses someone, how often is it right?
A camera-AI vendor pitches you. Somewhere on the second slide: 99% accurate. One of the big UK face-recognition firms, Facewatch, advertises 99.98%. It sounds like a solved problem.
It isn't. Accuracy is the wrong number, and at the rate real shops run at, it's the wrong number on purpose.
I've argued before that most retail cameras record without understanding. This is the flip side: even when the AI does understand, the number everyone buys it on is the wrong one.
Here's the fact the deck skips: a shop is almost entirely honest people. Even with UK retail theft at record highs — the BRC counted around 20 million customer-theft incidents last year — the person walking through your door is, overwhelmingly, not a thief. Call it one in two hundred on a bad day. That tiny base rate is where the number goes to die.
Take a camera a vendor could honestly call 96% accurate. Nobody advertises 96% — they claim higher; I'm being generous. Point it at a normal morning and watch what it does. Move the sliders: how many shoppers, how rare theft actually is, how good the model is, how often it cries wolf.
[Interactive demo: an interactive base-rate calculator showing how a high-accuracy theft-detection camera still flags mostly innocent shoppers (low precision)]
Leave it near the defaults — a busy store, one thief in two hundred, a flattering 90% catch rate, a 4% false-alarm rate — and the same system that scores 96% accurate flags about forty innocent people before lunch, next to four or five actual thieves. When that camera raises its hand, it's right about one time in ten.
The gap between "96% accurate" and "right one time in ten" isn't a bug. It's the same system, described two ways:
| The vendor's camera | Number |
| --- | --- |
| Headline "accuracy" | 96% |
| Precision — chance a flag is a real thief | ~10% |
| Innocent shoppers flagged | ~119 a day (~40 before lunch) |
| Actual thieves caught | ~14 a day |
Worked example: ~3,000 shoppers, one thief in two hundred, a flattering 90% catch rate, a 4% false-alarm rate. Change any of them in the calculator above.
Why the accuracy number lies
Accuracy asks: of everyone who came in, what fraction did the system judge correctly? When 199 of every 200 people are innocent, you get almost all of that score for free — for correctly ignoring the honest crowd, which is the easy part.
Here's the tell. A camera that flags nobody — unplugged, a brick — scores about 99.5% accurate at these rates. It never catches a thief and it still beats the "96%" system on accuracy. Any metric a brick can win is not measuring the thing you care about.
The thing you care about is precision: when the camera points at someone, how often is it right? That's what decides whether your staff are stopping shoplifters or stopping shoppers. And it's the number that quietly collapses when theft is rare — there are simply so many more innocent people to get wrong. To make even half your flags real thieves, you'd need a false-alarm rate under half a percent. Almost nothing in the wild hits that.
This isn't a hot take. It's the base-rate fallacy, a first-week statistics result. A professor ran the same arithmetic on Facewatch's 99.98% and landed in the same place: roughly one innocent person misidentified for every genuine catch, or worse.
The false positives aren't rounding errors. They're people.
A "flag" doesn't stay an abstraction. It becomes a member of staff walking over.
A 66-year-old in Cardiff was told to leave a B&M, accused of stealing about £75 of shopping, after the store's face-recognition matched him to a known offender. The CCTV showed he'd paid. A 19-year-old was searched and publicly ejected from a Home Bargains in Manchester on the same kind of match; the company later admitted it was wrong. These made the news because someone pushed back. Most don't.
And the errors don't land evenly. NIST's landmark study found many face-recognition algorithms misfire far more often on women and on darker skin — 10 to 100 times more false positives for some groups (to be fair, several of the most accurate systems showed little gap). A blended "96%" hides who actually pays for the other 4%.
"But a human reviews every alert"
This is the standard reply, and it's worth taking seriously — because it's exactly what vendors say when it goes wrong. When that Cardiff pensioner was wrongly stopped, the firm's line was that it was "human error, not a failure of the technology".
But a human in the loop doesn't fix the system's precision. It hands your staff forty alerts a morning to adjudicate, and two things happen. Either they trust the confident AI and wave it through — now you've automated the accusation — or they drown, tune it out, and stop looking by week three. Either way you've rebuilt the thing I keep complaining about: a screen full of alerts nobody reads. A camera that cries wolf forty times before lunch doesn't earn a human's attention. It burns it.
Ask the only question that matters
None of this means the tech is useless. It means "accuracy" is marketing, and you should refuse to buy on it. Three questions cut through the deck:
- What's your precision — at my store's base rate, not a lab set? If they can't answer, they've never measured how many flagged shoppers were innocent, because they have no way to.
- How many false accusations a day does that precision imply at my footfall? Put a number on the people, not just the catches.
- What does one false positive cost — the stopped customer who never comes back, the staff time, the headline?
The real fix is upstream. Stop trying to recognise a thief — a rare, low-signal needle — and detect a specific behaviour with enough signal that precision is high by construction. Fewer flags, and the ones you raise are mostly right. That's the whole reason QuantumEye reads what's happening on the shelf instead of scoring who's standing in front of it. When a false positive is a person, precision isn't a nice-to-have. It's the product.
So the next time a camera vendor opens with an accuracy number, you already know the follow-up: when your camera points at one of my customers, how often is it wrong — and who pays for it?
I build QuantumEye, and I take on the gnarly "a false positive is a person" problems through Purple Luna, my product-and-engineering consultancy. The calculator above is a worked example, not a measurement — plug in your own store's numbers.
This page requires JavaScript to view fully. The summary above is for indexing.