The Insurance Institute for Highway Safety filtered 78% of Waymo's crash reports to match human driver thresholds, then called the manual coding 'unsustainable' at scale.
To produce its 68% figure, the Insurance Institute for Highway Safety discarded 78% of Waymo's reported incidents, keeping only the 22% that humans would actually have called in. Robotaxis must report every scrape, curb strike, and undercarriage tap. Human drivers only report crashes involving injury or more than $1,000 in property damage. The filtering step is the finding. The percentage is what survived it.
That 68% applies to one operator. Waymo is the only Level 4 (the SAE classification for full driverless operation, no human driver required) robotaxi company that voluntarily reports its vehicle miles traveled (VMT), the cumulative distance its fleet has driven. The IIHS study by Teoh, Kidd, and Riexinger compared Waymo's roughly 50 million driverless miles against the human-driving baseline in San Francisco, Phoenix, Los Angeles, and Austin, where drivers logged about 222 billion miles in the same window. The dataset has one company in it.
Phoenix showed a 76% lower crash rate per mile for Waymo. Los Angeles came in 71% lower. San Francisco showed 35% lower. Austin came in 4% higher, a result IIHS flagged as small-sample noise rather than a counter-signal. Each number is a single-operator datapoint against a four-city human baseline.
Waymo's own self-reported numbers, picked up in the WEAR TV syndication and the Automotive World summary, run higher than IIHS's filtered comparison: 82% fewer injury-causing crashes and 93% fewer pedestrian crashes with injuries. These are the operator's own figures, not the IIHS result. They use a different threshold and a different baseline, so the IIHS number is the conservative one, not the company's.
IIHS also broke out crash severity. Single-vehicle crashes came in 85% lower per mile for Waymo. Injury crashes came in 81% lower. Those figures survive the same filtering, so they are the cleanest read on what a single voluntary dataset shows. They are also the cleanest read on what one company has chosen to share.
Of 736 crashes filed by L4 operators under the NHTSA Standing General Order (SGO), the crash-reporting rule NHTSA activated in 2021 to collect Level 2 and Level 4 incident data, only 22% met the police-reportable threshold used for the human-driver comparison. IIHS researchers manually coded incident narratives to filter the robotaxi set. Lead author Eric Teoh said, in the IIHS press release, that driverless operators must report events most people would not call crashes, which distorts any naive side-by-side.
IIHS's president, David Harkey, made the bottleneck explicit in the IIHS press release: "The present data collection system isn't good enough to allow continuous monitoring of a large-scale expansion. Now we need to get the data collection system right, so that we can ensure that level of safety continues." The same release calls the manual narrative coding "unsustainable" at the fleet sizes L4 operators are scaling toward.
The structural problem is not IIHS's. The NHTSA SGO requires AV developers to file crash reports but does not require fleet size, vehicles in operation, or VMT. NHTSA itself has acknowledged that incident-rate contextualization is limited without that denominator. The agency can count incidents; it cannot compute a rate.
The same week, NHTSA Administrator Jonathan Morrison opened an enforcement action over AV interaction with first responders, citing "functional insufficiency" in how robotaxis handle emergency vehicles and flagging FMVSS (Federal Motor Vehicle Safety Standards, the federal crashworthiness and crash-avoidance rules) modernization and new AV safety guidance as the next steps. That is a separate signal from the data gap, but it points in the same direction: the federal picture of robotaxi safety is being rebuilt from voluntary pieces.
The watch item is whether NHTSA amends the SGO to require VMT, fleet size, and vehicles-in-operation reporting. Without that, the next IIHS-style comparison will face the same problem: discard most robotaxi incidents, code the rest by hand, and measure the result against a single operator's voluntary mileage disclosures. The 68% is real. The infrastructure to keep producing it is not.