# Face recognition pilot checklist

What to agree before a vendor arrives, what to measure, and how the result
should be reported.

The single most important item is the first one in Phase 1: **agree the failure
criteria in writing before anything is installed.** Criteria agreed afterwards
are criteria that moved. A vendor unwilling to agree them in advance is offering
a demonstration and calling it a pilot.

Full guidance: https://facerec.ayonix.com/resources/face-recognition-pilot-checklist

## Phase 1 — Before anyone is on site

- [ ] Write the question the pilot answers, in one sentence. Not "evaluate face recognition" but something like "can this entrance process the 08:00 peak without adding staff, at an error rate we can defend".
- [ ] Define the population: who, how many distinct individuals, how many passes each, and how representative they are of the people who will actually use the system.
- [ ] Set the measurement window against the variation it must cover — busiest and quietest hours, brightest and darkest conditions — rather than by a round number of days.
- [ ] Agree both error criteria in writing, including what failure looks like. A criterion the system cannot fail is not a criterion.
- [ ] Decide who owns the threshold. The customer sets it, advised by the vendor, with the reasoning recorded and a named role permitted to change it later.
- [ ] Agree the reporting format: counts and denominators for both error types, breakdown by group, throughput components, capture conditions, and an account of anything that went wrong.

## Phase 2 — Preparation

- [ ] Survey capture conditions at every point: pixels across the eyes where people actually pass, angle, lighting, compression and walking speed — measured at the worst hour of the day.
- [ ] Enrol under the conditions the live system will use, not in a well-lit meeting room if the door is backlit.
- [ ] Confirm the integration endpoint the result must reach, and test that it is reachable from where the analytics will run.
- [ ] Agree what happens to captured images and templates during and after the pilot, and who deletes them.

## Phase 3 — Running it

- [ ] Record every comparison, not only the alerts. The denominator is where the information is, and it cannot be reconstructed afterwards.
- [ ] Count both error types separately: people missed, and people matched to the wrong identity.
- [ ] Record retries, exceptions and exception resolution times. These usually determine throughput more than the comparison does.
- [ ] Log operator decisions separately from system output, so the two can be compared.
- [ ] Note any change made mid-pilot, with the date. A tuned system after week one is a different system from the one measured in week one.

## Phase 4 — Reporting and deciding

- [ ] Break the results down by demographic group, using the population that will actually use the system — and say so explicitly if the pilot was not sized to support that.
- [ ] Test the fallback path under load: how long it takes, how it feels to use, and whether it absorbs the exception rate measured.
- [ ] Exercise the audit trail by reconstructing one transaction from the log alone.
- [ ] Report the result whatever it shows, including the failures, the surprises and anything that had to be changed.
- [ ] Decide against the criteria agreed in Phase 1, not against criteria revised in the light of the result.

## How much data a measurement needs

- [ ] Decide the smallest error rate that would change your decision. Everything else follows from this.
- [ ] Size the pilot so a system performing at exactly that rate would produce enough events to count within the window.
- [ ] Observing no errors at all in fifty attempts is consistent with a true rate of one in fifty, one in a hundred or one in five hundred. If the practical number of trials is too large, the honest conclusion is that the pilot cannot answer that question — not that the answer is good.
- [ ] A pilot sized to measure an aggregate is usually not large enough to measure a difference between demographic groups. Do not promise one unless the sizing supports it.

## Six questions to ask any vendor, including Ayonix

- [ ] Which algorithm identifier did you submit to an independent evaluation, and on what date?
- [ ] At what threshold is your quoted figure measured, and on which dataset?
- [ ] What is the demographic breakdown behind the aggregate you publish?
- [ ] Which presentation attack instruments has liveness been tested against, and by which laboratory?
- [ ] Will you run a pilot on our cameras, and will you report the failures as well as the matches?
- [ ] Where do templates live, and what happens to recognition when the internet connection drops?

---

Published by Ayonix, facerec.ayonix.com. Ayonix sells face recognition and
therefore has an interest in how these checklists are used — they are written
so they can be applied to Ayonix as readily as to any other supplier, and if
following one leads you to a different supplier or a different category of
system, that is a legitimate outcome.

No accuracy percentage appears in any Ayonix material, for Ayonix or any other
vendor, because a figure without its threshold, dataset, gallery size and
demographic breakdown cannot be reproduced.

Corrections: infojp@ayonix.com. Last reviewed 11 September 2026.
