In short
What to agree before a vendor arrives: the population, the measurement window, how many subjects a result actually needs, both error types to record, the threshold and who may change it, and how the outcome is reported. Written to be used against Ayonix as well as by it.
Agree the failure criteria before the vendor arrives
This is the entire trick, and everything else on this page follows from it. Criteria agreed after a pilot has run are criteria that moved, because by then both parties have an interest in where they land. Criteria agreed in writing beforehand — including what constitutes failure — turn a sales exercise into a measurement.
A vendor that will not agree failure criteria in advance is offering a demonstration and calling it a pilot. That is worth knowing early, and asking for the criteria is a cheap way to find out.
-
Agree
Population, window, criteria, threshold ownership and reporting format — in writing, before anything is installed.
Fails when: Criteria are left to be decided later, and later never produces a number anyone will act on.
-
Survey and enrol
Measure capture conditions at each point; enrol under the conditions the live system will use.
Fails when: Enrolment happens in a bright room and the door is backlit, so the pilot measures a system nobody will deploy.
-
Run and record
Every comparison logged, not only the alerts. The denominator is where the information is.
Fails when: Only matches are recorded, so no error rate can be calculated at all.
-
Report and decide
Both error types, the breakdown by group, throughput, and everything that went wrong.
Fails when: The report contains only successes, which means it is a case study rather than a measurement.
Sizing
How much data a measurement actually needs
The most common way a pilot produces a meaningless number is by being too small to detect the error rate it was supposed to measure.
The principle is straightforward even if the arithmetic is not. To measure an error rate, you need enough trials that the error occurs often enough to count reliably. Observing no errors at all in fifty attempts is entirely consistent with a true rate of one in a hundred, one in five hundred, or one in fifty — the measurement cannot tell them apart.
So work backwards. Decide the smallest error rate that would change your decision. Then size the pilot so that a system performing at exactly that rate would produce enough events to be countable within the window. If that number of trials is impractical, the honest conclusion is that the pilot cannot answer that question — not that the answer is good.
This applies to both error types independently, and to each demographic group you intend to report separately. A pilot large enough to measure an aggregate is usually not large enough to measure a difference between groups, which is worth knowing before promising one.
| Sizing question | Why it decides the size | If it is not asked |
|---|---|---|
| What is the smallest error rate that would change our decision? | Everything else follows from this. It sets the number of trials needed. | The pilot runs for a convenient period and produces a number nobody can interpret. |
| How many distinct people, and how many passes each? | Both matter. Fifty people once is not the same measurement as ten people five times. | Repeat passes by the same cooperative subjects flatter the result substantially. |
| Do we need per-group results? | A pilot sized for an aggregate is usually too small to compare groups. | A demographic breakdown is promised and then cannot be reported with any confidence. |
| What variation must the window cover? | Lighting, staffing, season, day of week. The window is defined by variation, not by a round number. | Two months of identical days measures one condition very precisely and the deployment not at all. |
| Are we measuring the exception path too? | Exception resolution time needs its own sample, and it is often the binding constraint. | Throughput is reported without the exceptions that will actually determine it. |
The pilot checklist
Sixteen items across four phases. Work through the first phase before a vendor is on site; it is where the value is.
-
Write the question the pilot answers, in one sentence
Not "evaluate face recognition". Something like "can this entrance process the 08:00 peak without adding staff, at an error rate we can defend". A pilot without a question produces an impression.
-
Define the population that will be measured
Who, how many distinct individuals, how many passes each, and how representative they are of the people who will actually use the system.
-
Set the measurement window against the variation it must cover
Busiest and quietest hours, brightest and darkest conditions, a normal week and an unusual one. Define it by variation rather than by a round number of days.
-
Agree both error criteria, including what failure looks like
In writing, before installation. A criterion the system cannot fail is not a criterion, and agreeing them afterwards is agreeing them under pressure.
-
Decide who owns the threshold and record the value
The customer sets it, informed by the vendor, with the reasoning written down. Record who may change it after go-live and what is logged when they do.
-
Agree the reporting format up front
Counts and denominators for both error types, breakdown by group, throughput components, capture conditions, and an account of anything that went wrong.
-
Survey capture conditions at every point
Pixels across the eyes, angle, lighting, compression and walking speed — measured, at the worst hour. This bounds the result and explains it afterwards.
-
Enrol under the conditions the live system will use
Not in a well-lit meeting room if the door is backlit. Enrolment quality is a ceiling on every later comparison and is independent of camera placement.
-
Record every comparison, not only the alerts
The denominator is where the information is. A log of matches cannot produce an error rate, and reconstructing one afterwards is not possible.
-
Count both error types separately
People missed, and people matched to the wrong identity. Either can be driven to zero at the expense of the other, so one without the other describes nothing.
-
Record retries, exceptions and resolution times
These usually determine throughput more than the comparison does, and they are what an operations team has to plan staffing against.
-
Break the results down by demographic group
Using the population that will actually use the system, and only if the pilot is sized to support it. Say so if it is not.
-
Log operator decisions separately from system output
And compare them. If operators essentially never disagree, review has become a signature — a finding the pilot should surface rather than conceal.
-
Test the fallback path under load
Not just that it exists. How long it takes, how it feels to use, and whether it can absorb the exception rate the pilot measured.
-
Exercise the audit trail by reconstructing a transaction
Pick a comparison from the middle of the window and reconstruct it from the log alone. An audit record that has never been used is an assumption.
-
Report the result whatever it shows
Including the failures, the surprises and the things that had to be changed mid-pilot. A report containing only successes is a case study, not a measurement.
Warning signs
What a vendor’s answers tell you
These come up in every pilot negotiation. How a supplier responds is informative well before any equipment is installed.
“Will you agree failure criteria in writing beforehand?”
Encouraging: Yes, and here is what we would propose.
Concerning: Let us run it first and see what the numbers look like.
“Will the pilot run on our cameras?”
Encouraging: Yes, and we will survey them first and tell you which ones will not work.
Concerning: We will bring our own cameras so you can see the system at its best.
“Will you report the failures?”
Encouraging: Yes — the failures are where the useful information is.
Concerning: We will focus the report on the successful matches.
“Who sets the threshold?”
Encouraging: You do, advised by us, and we will record the reasoning.
Concerning: We will tune it for you as we go.
“Can you break the results down by demographic group?”
Encouraging: Yes if the pilot is sized for it; here is what that would require.
Concerning: Our algorithm does not have that problem.
“What is your accuracy?”
Encouraging: That depends on the threshold, dataset and gallery size — let us measure it here.
Concerning: A percentage, offered immediately, with none of those attached.
Ask us these too. On the last one, the answer you will get is that this site publishes no accuracy percentage for anyone including Ayonix, and that the number you want is produced by a pilot on your cameras — which is either reassuring or frustrating depending on what you were hoping for, and is the same answer either way.
Frequently asked questions
What is the difference between a pilot and a demonstration?
A demonstration runs on the vendor’s equipment, with the vendor’s staff, in the vendor’s lighting, and produces an impression. A pilot runs on your cameras, with your population, against acceptance criteria agreed in writing before it starts, and produces a measurement. Both are useful, but only one of them tells you anything about your building — and a vendor unwilling to agree failure criteria in advance is offering the first while calling it the second.
How many test subjects does a pilot need?
Enough that the error you care about occurs often enough to count. If you expect an error rate around one in a hundred and you test fifty people once each, you cannot distinguish one in fifty from one in five hundred — the measurement is not capable of answering the question. The practical method is to decide the smallest error rate that would change your decision, then size the test so that rate would produce enough events to count within the window.
What should a pilot measure besides accuracy?
Throughput end to end including retries and exceptions; the retry rate; the exception rate and how long exceptions take to resolve; performance broken down by demographic group; and how often operators overrule the system. That last one is almost never collected and is the only way automation bias becomes visible.
Who should set the threshold during a pilot?
The customer, informed by the vendor, with the value and the reasoning recorded. A threshold set by the vendor to make the pilot look good is a threshold that will have to be changed after go-live, at which point the pilot result no longer describes the running system. Record who may change it afterwards too.
How long should a pilot run?
Long enough to cover the variation that matters: the busiest hour and the quietest, the brightest conditions and the darkest, a normal week and an unusual one. Two weeks that include a seasonal lighting change tell you more than two months of identical days. The window should be defined by what variation it needs to capture, not by a round number.
What should the pilot report contain?
Both error types with their counts and denominators; the threshold and the reasoning; the breakdown by demographic group; the retry, exception and throughput figures; the capture conditions measured at each point; and an honest account of anything that went wrong. A report containing only successes is not a measurement.
What if the pilot fails its criteria?
Then it has done its job. A pilot that cannot fail is a demonstration with extra steps. A genuine failure usually points at capture geometry or enrolment quality rather than at the algorithm, both of which are fixable — but the fix has to be identified and re-measured rather than argued around, and that is why the criteria were agreed in writing first.
Next step
Run this against us
Send the checklist back with your acceptance criteria filled in and we will tell you whether we can meet them before anyone books a visit. If the honest answer is that your capture conditions will not support the criteria, that is a more useful conversation than a pilot that was always going to fail.