The Latin American used car marketplace spends as much engineering effort on AI evaluations as on the agents themselves. That choice, not the agent deployment, is what makes the 96% / 95% numbers auditable.
Kavak rebuilt its customer operations, sales workflow, and engineering org around AI agents, and the headline metrics of 96% of customer interactions and 95% of transactions handled by software rest on a quieter bet. The Latin American used-car marketplace spends as much engineering effort on the evaluations that grade those agents as on building the agents themselves, according to show notes from a16z's podcast episode with Chief Product & AI Officer Alejandro Maza Ayala (a16z).
A car listing conversation that once ping-ponged between a sales rep and a back-office agent now closes inside a single AI workflow. Maza's argument on the episode is that the binding constraint has shifted from "can it run" to "can it be measured," and where engineering hours go is now the strategic question.
Kavak's mechanism for the rebuild is what Maza calls evals-as-product parity. Building an agent is half the job. The other half is the test suite that catches regressions, drift, and edge cases before they reach a buyer in Mexico City or São Paulo. Kavak's case, laid out in the a16z episode and the cross-posted a16z.com show notes, is that an agent without a graded eval is a demo, not infrastructure, so the company staffs the two tracks on the same footing. A LinkedIn post by an outside observer restates the customer-interaction figure as 90-95%, which lines up with the podcast's range, though neither the show notes nor the post publish the underlying measurement methodology.
The org redesign is the second half. Kavak did not just license copilots and roll them out. It retrained mechanics, executives, and salespeople to work with the agents, then restructured teams so that human review sits where the evals are weakest. AI sellers are described in the show notes as outperforming the company's human sales teams, but no public metric backs that comparison, so the claim should be read as a guest assertion rather than a benchmarked result (a16z). The smaller, more falsifiable experiment is the AI "CEO" trial: in one city, for one month, an agent running pricing and operations decisions increased profits by 50% per the show notes, a single-city, single-month result that reads as a directional signal rather than a production commitment.
To reach the 96% headline, a dealer has to put evaluation work on the same engineering backlog as agent work, retrain the people who would have answered the phone, and accept that the test suite is the product. The traditional dealer software stack (CRM, dialer, lead router) was never built around measurement parity between software and human output, and adopting agents on top of that stack tends to leave the eval side underfunded. Kavak's choice to staff evals as a peer of agent development is what makes the headline metric auditable in the first place.
Two things to track next. First, whether the 96% and 95% figures hold once the full episode transcript is read and the methodology is published; the show notes are a publisher summary, not an independently verified measurement. Second, whether other LatAm marketplaces, or US dealers importing the playbook, replicate the evals-first staffing model, which is the part hardest to copy and the part that determines whether the 96% and 95% numbers hold up next quarter.