Flock's vendor benchmark said 96%. In Roseville, CA, 71% of alerts were misreads. The gap is structural, and it travels beyond policing.
In Roseville, California, police bought Flock Safety's license-plate-reading cameras on a 96% accuracy claim. Over two years, the cameras misread plates in 71% of the alerts they generated, a Business Insider investigation cited by Forrester found. A dispatch supervisor told investigators she had memorized a known-bad match rather than following the department's policy to verify alerts before acting. The pattern travels: any organization that buys a third-party AI inherits the same gap between vendor testing and buyer deployment.
Flock sells automatic license plate reader (ALPR) cameras to law enforcement agencies across the United States. The cameras capture plate images and run them through an AI model that compares what it sees against hotlists of wanted vehicles. The vendor benchmark, 96% accuracy, was measured under what Flock calls "optimal conditions." The Business Insider investigation reviewed two years of alerts sent to the Roseville Police Department and found that 71% of those alerts were misreads: plate numbers the system matched to the wrong vehicle, or matched correctly to a vehicle that turned out to be unrelated to the underlying case. Forrester's analysis, which re-reports the BI findings, frames the gap as a structural problem with how third-party AI is procured rather than a Flock-specific defect.
A 25-point gap is the difference between two measurement contexts. Vendor benchmarks test the model on curated data under controlled lighting, angle, and plate-condition inputs. Production cameras read plates at night, in rain, at oblique angles, on dirty or bent plates, and sometimes through windshields. NIST's AI Risk Management Framework calls the distinction "mapping risk from measuring performance": a vendor's number describes the model's behavior on its own test set, not the system's behavior in a buyer's environment. The Roseville number describes the system in Roseville's environment, and the buyer's environment is the one that has to absorb the consequences.
The enforcement layer is where the gap becomes a public problem. Roseville had a written policy requiring officers to verify an alert before acting on it. The dispatch supervisor's admission, "It's easier at this point we have it memorized," suggests the policy existed on paper but failed in practice. In a separate Flock case in Minnesota, a family was stopped by multiple police vehicles after a faulty system match. The risk in a misread is not a false positive in a database; it is a police encounter with a person who did nothing wrong.
The mechanism is portable. Any organization that buys a third-party AI inherits three failure points: a benchmark number measured in someone else's conditions, a deployment environment that differs from those conditions, and a human operator who translates model output into action. Each point can break the chain independently, and the same three points show up in healthcare triage tools, lending models, and hiring screens. The vendor's test data was not built from the buyer's population, the deployment conditions were not the test conditions, and the operator has to decide what the model's output means in a specific, real situation.
Procurement looks different when the gap is taken seriously. Pilots run in the buyer's own environment with the buyer's own ground truth. Vendor benchmarks become a starting hypothesis rather than a deliverable. Deployment metrics are tracked from day one, and rollbacks are designed in rather than improvised. The BI investigation is a reminder that hotlist-matching AI is high-stakes in policing, but the procurement lesson applies wherever a third-party model is being trusted to act on someone's behalf. The 96% number is what the vendor measured. The number a buyer lives with is the one they measure themselves, and the responsible move is to publish that number once it exists.