A single accuracy percentage describes a face recognition system about as well as a single temperature describes a country.
Vendor material tends to lead with one figure — some number of nines, measured somewhere, on somebody's faces. It is not usually dishonest. It is just not a claim you can act on, because it collapses two different errors that have opposite causes and very different costs.
Two errors, two different costs
The two errors, and who feels each one
Every identification system can be wrong in exactly two directions, and the distinction is the whole subject:
| Error | What happened | Who notices | What it costs |
|---|---|---|---|
| False accept | The wrong person was recorded as present | Nobody, usually | Payroll paid for absence; the record is silently wrong |
| False reject | The right person was not recognised | The employee, immediately | A queue, a supervisor override, an unhappy worker |
The asymmetry in that last column is the important part. A false reject is loud and self-correcting — someone complains and it gets fixed. A false accept is quiet and permanent: nobody reports having been paid for a day they did not work, so the error enters payroll and stays there.
Which means a system tuned to keep employees happy is tuned in the direction that damages your data, and it will look like it is working well while doing so. This is worth knowing before someone adjusts a threshold to stop the complaints.
They trade against each other, always
Matching works on a similarity score against a threshold. Raise the threshold and you demand a closer match: fewer wrong people admitted, more right people turned away. Lower it and the reverse. There is no setting at which both improve, which is why "accuracy" as a single number is really a statement about where somebody set a dial.
So the question to a vendor is not "how accurate is it" but "at what false accept rate is that false reject rate measured, and can we change the threshold ourselves". A vendor who cannot express performance as a pair is quoting one point on a curve without telling you which point.
Why a good-sounding rate is a daily nuisance
Per-event error rates feel small and accumulate fast. Take a hypothetical site of 400 staff, each recognised twice a day, over a 22-day month — illustrative numbers, chosen to make the arithmetic plain. That is 17,600 recognition events:
| False reject rate | Failed recognitions per month | What that feels like |
|---|---|---|
| 0.1% | about 18 | A manageable exception list |
| 1% | about 176 | Eight a day; supervisors start overriding routinely |
| 3% | about 528 | The parallel paper system is back |
A 3% failure rate reads as 97% accurate, which sounds respectable and is operationally unusable. This is the same trap as character-level accuracy in document OCR, where a per-character figure hides a per-field failure rate — the arithmetic that makes a 99% claim on cheque fields far less impressive than it appears.
Published figures do not describe your workforce
Benchmark results are measured on curated datasets under controlled capture. Your entrance is not that. Accuracy varies with lighting, camera angle, distance, motion, and how much of a face is visible — and in Gulf conditions that means direct sun at one hour and glare off pale ground at another, plus headwear, hard hats, glasses, masks and beards that are ordinary rather than exceptional.
Demographic variation is also real and well documented across the field: error rates differ by skin tone, age and gender, and a system evaluated mainly on one population can perform measurably worse on another. For a workforce as mixed as a typical Omani site, a headline figure from elsewhere tells you very little, and the only number worth trusting is one measured on your own people at your own doorway.
How to run a pilot that actually tells you something
Run it at the real entrance, at shift change, not in a meeting room. Cover the awkward hours — first light, midday glare, night — because a system that works at 10am and fails at 6am has an average that means nothing. Include everyone, especially the people hardest to enrol, since excluding them from the pilot is how you discover them in production. Keep an independent record of who actually attended so you can count both error types rather than only the complaints. Ask for performance as a pair of rates and for control of the threshold. And measure the exception queue, not just the match rate, because the queue is what determines whether people trust the system a month in.
The honest summary
Ask for two numbers instead of one, insist on control of the threshold, and measure on your own site. A vendor who answers in pairs and hands you the dial is describing a system honestly; a vendor with one number and no curve is describing a marketing position. Note also that before any of this can run on real faces in Oman there is a Ministry permit to obtain, which is worth sequencing before the pilot rather than after.
Maugood AI is built in Muscat by Muscat Tech Solutions, and we would rather be measured at your entrance than quoted at. To arrange that, get in touch.
Related posts
-
Drawing Perimeter Zones That Survive Palm Shadow and Blowing Dust
Every false alarm has a physical cause. Most of them are on this list.
26 August 2025 -
Reading the Omani Mulkiya: The Fields That Break Extraction
A fixed layout makes it tractable. Bilingual text and a plate format make it interesting.
12 August 2025 -
Revolutionizing Bank Statement Analysis with Powerful AI
Turning complex financial data into clear, actionable insights in seconds.
09 August 2025


