Ayonix Face Recognition

In short

What to agree before a vendor arrives: the population, the measurement window, how many subjects a result actually needs, both error types to record, the threshold and who may change it, and how the outcome is reported. Written to be used against Ayonix as well as by it.

Agree the failure criteria before the vendor arrives

This is the entire trick, and everything else on this page follows from it. Criteria agreed after a pilot has run are criteria that moved, because by then both parties have an interest in where they land. Criteria agreed in writing beforehand — including what constitutes failure — turn a sales exercise into a measurement.

A vendor that will not agree failure criteria in advance is offering a demonstration and calling it a pilot. That is worth knowing early, and asking for the criteria is a cheap way to find out.

Pilot sequence
  1. Agree

    Population, window, criteria, threshold ownership and reporting format — in writing, before anything is installed.

    Fails when: Criteria are left to be decided later, and later never produces a number anyone will act on.

  2. Survey and enrol

    Measure capture conditions at each point; enrol under the conditions the live system will use.

    Fails when: Enrolment happens in a bright room and the door is backlit, so the pilot measures a system nobody will deploy.

  3. Run and record

    Every comparison logged, not only the alerts. The denominator is where the information is.

    Fails when: Only matches are recorded, so no error rate can be calculated at all.

  4. Report and decide

    Both error types, the breakdown by group, throughput, and everything that went wrong.

    Fails when: The report contains only successes, which means it is a case study rather than a measurement.

Sizing

How much data a measurement actually needs

The most common way a pilot produces a meaningless number is by being too small to detect the error rate it was supposed to measure.

The principle is straightforward even if the arithmetic is not. To measure an error rate, you need enough trials that the error occurs often enough to count reliably. Observing no errors at all in fifty attempts is entirely consistent with a true rate of one in a hundred, one in five hundred, or one in fifty — the measurement cannot tell them apart.

So work backwards. Decide the smallest error rate that would change your decision. Then size the pilot so that a system performing at exactly that rate would produce enough events to be countable within the window. If that number of trials is impractical, the honest conclusion is that the pilot cannot answer that question — not that the answer is good.

This applies to both error types independently, and to each demographic group you intend to report separately. A pilot large enough to measure an aggregate is usually not large enough to measure a difference between groups, which is worth knowing before promising one.

What a pilot must be sized against, and the failure that follows from getting each wrong.
Sizing question Why it decides the size If it is not asked
What is the smallest error rate that would change our decision? Everything else follows from this. It sets the number of trials needed. The pilot runs for a convenient period and produces a number nobody can interpret.
How many distinct people, and how many passes each? Both matter. Fifty people once is not the same measurement as ten people five times. Repeat passes by the same cooperative subjects flatter the result substantially.
Do we need per-group results? A pilot sized for an aggregate is usually too small to compare groups. A demographic breakdown is promised and then cannot be reported with any confidence.
What variation must the window cover? Lighting, staffing, season, day of week. The window is defined by variation, not by a round number. Two months of identical days measures one condition very precisely and the deployment not at all.
Are we measuring the exception path too? Exception resolution time needs its own sample, and it is often the binding constraint. Throughput is reported without the exceptions that will actually determine it.

The pilot checklist

Sixteen items across four phases. Work through the first phase before a vendor is on site; it is where the value is.

  1. Write the question the pilot answers, in one sentence

    Not "evaluate face recognition". Something like "can this entrance process the 08:00 peak without adding staff, at an error rate we can defend". A pilot without a question produces an impression.

  2. Define the population that will be measured

    Who, how many distinct individuals, how many passes each, and how representative they are of the people who will actually use the system.

  3. Set the measurement window against the variation it must cover

    Busiest and quietest hours, brightest and darkest conditions, a normal week and an unusual one. Define it by variation rather than by a round number of days.

  4. Agree both error criteria, including what failure looks like

    In writing, before installation. A criterion the system cannot fail is not a criterion, and agreeing them afterwards is agreeing them under pressure.

  5. Decide who owns the threshold and record the value

    The customer sets it, informed by the vendor, with the reasoning written down. Record who may change it after go-live and what is logged when they do.

  6. Agree the reporting format up front

    Counts and denominators for both error types, breakdown by group, throughput components, capture conditions, and an account of anything that went wrong.

  7. Survey capture conditions at every point

    Pixels across the eyes, angle, lighting, compression and walking speed — measured, at the worst hour. This bounds the result and explains it afterwards.

  8. Enrol under the conditions the live system will use

    Not in a well-lit meeting room if the door is backlit. Enrolment quality is a ceiling on every later comparison and is independent of camera placement.

  9. Record every comparison, not only the alerts

    The denominator is where the information is. A log of matches cannot produce an error rate, and reconstructing one afterwards is not possible.

  10. Count both error types separately

    People missed, and people matched to the wrong identity. Either can be driven to zero at the expense of the other, so one without the other describes nothing.

  11. Record retries, exceptions and resolution times

    These usually determine throughput more than the comparison does, and they are what an operations team has to plan staffing against.

  12. Break the results down by demographic group

    Using the population that will actually use the system, and only if the pilot is sized to support it. Say so if it is not.

  13. Log operator decisions separately from system output

    And compare them. If operators essentially never disagree, review has become a signature — a finding the pilot should surface rather than conceal.

  14. Test the fallback path under load

    Not just that it exists. How long it takes, how it feels to use, and whether it can absorb the exception rate the pilot measured.

  15. Exercise the audit trail by reconstructing a transaction

    Pick a comparison from the middle of the window and reconstruct it from the log alone. An audit record that has never been used is an assumption.

  16. Report the result whatever it shows

    Including the failures, the surprises and the things that had to be changed mid-pilot. A report containing only successes is a case study, not a measurement.

Download the pilot checklist

Frequently asked questions

What is the difference between a pilot and a demonstration?

A demonstration runs on the vendor’s equipment, with the vendor’s staff, in the vendor’s lighting, and produces an impression. A pilot runs on your cameras, with your population, against acceptance criteria agreed in writing before it starts, and produces a measurement. Both are useful, but only one of them tells you anything about your building — and a vendor unwilling to agree failure criteria in advance is offering the first while calling it the second.

How many test subjects does a pilot need?

Enough that the error you care about occurs often enough to count. If you expect an error rate around one in a hundred and you test fifty people once each, you cannot distinguish one in fifty from one in five hundred — the measurement is not capable of answering the question. The practical method is to decide the smallest error rate that would change your decision, then size the test so that rate would produce enough events to count within the window.

What should a pilot measure besides accuracy?

Throughput end to end including retries and exceptions; the retry rate; the exception rate and how long exceptions take to resolve; performance broken down by demographic group; and how often operators overrule the system. That last one is almost never collected and is the only way automation bias becomes visible.

Who should set the threshold during a pilot?

The customer, informed by the vendor, with the value and the reasoning recorded. A threshold set by the vendor to make the pilot look good is a threshold that will have to be changed after go-live, at which point the pilot result no longer describes the running system. Record who may change it afterwards too.

How long should a pilot run?

Long enough to cover the variation that matters: the busiest hour and the quietest, the brightest conditions and the darkest, a normal week and an unusual one. Two weeks that include a seasonal lighting change tell you more than two months of identical days. The window should be defined by what variation it needs to capture, not by a round number.

What should the pilot report contain?

Both error types with their counts and denominators; the threshold and the reasoning; the breakdown by demographic group; the retry, exception and throughput figures; the capture conditions measured at each point; and an honest account of anything that went wrong. A report containing only successes is not a measurement.

What if the pilot fails its criteria?

Then it has done its job. A pilot that cannot fail is a demonstration with extra steps. A genuine failure usually points at capture geometry or enrolment quality rather than at the algorithm, both of which are fixable — but the fix has to be identified and re-measured rather than argued around, and that is why the criteria were agreed in writing first.