Tether's MedPsy 4B edged Google's MedGemma 27B on U.S. medical exams while running locally. CEO Paolo Ardoino says the on device curve is about to retire the trillion dollar GPU build out.
A 4-billion-parameter medical AI just beat Google's 27-billion-parameter MedGemma on the standard U.S. medical-licensing exam set, and the smaller model runs on a phone. Tether, the company behind the USDT stablecoin used by roughly 573 million wallets, built it. The benchmark numbers are Tether's own; the deployment math is where the comparison starts to bite the cloud-AI incumbents.
Tether's QVAC MedPsy-4B scores 70.54 on a seven-benchmark closed-ended medical evaluation, edging Google's MedGemma-27B-text-it at 69.95, according to the company's release on Hugging Face. The gap widens where the harder tests live: HealthBench Hard (58.00 vs. 42.00, +16.00), the full HealthBench (74.00 vs. 65.00), and MedXpertQA (30.61 vs. 25.18). The smaller sibling, MedPsy-1.7B, averages 62.62, beating Google's MedGemma-1.5-4B-it (51.20) by more than 11 points at less than half the parameters, and matching Qwen3-4B-Thinking-2507 at roughly 2.4x the size advantage in the other direction. Every number in this paragraph is Tether-reported, run on Tether Data's own evaluation harness; independent reproduction has not yet appeared in the public packet.
MedPsy-4B produces roughly 909 output tokens where the comparable Qwen3 model needs 2,953, a 3.2x reduction, because the model is trained on a narrow medical domain instead of the whole web. The quantized Q4_K_M footprint lands at about 2.6 GB; the 1.7B version is about 1.2 GB, which is what an $80 Android phone in Lagos, Karachi, or Manila can hold in RAM. Eight benchmark suites back the score: MedQA-USMLE, MedMCQA, MMLU Health, MMLU-Pro Health, MedXpertQA, PubMedQA, AfriMedQA, and HealthBench. The weights ship under Apache 2.0 through the QVAC SDK. A clinician can run a USMLE-grade assistant on the device in her hand, with no data center in the loop and no patient note leaving the building.
That is also the unit-economics case Ardoino made on Eye on AI this week. His argument: a vendor selling a $200 annual AI subscription is delivering between $1,000 and $5,000 of inference cost per customer, and the gap is being absorbed privately while the company is venture-funded. The moment one of those vendors goes public, or has to defend its margin to a public-market investor, the bill either lands on the subscriber or on retail shareholders. Ardoino frames it as a parallel to the early USDT playbook: give the user the product for free, monetize elsewhere. The 573 million wallets already on Tether's rails are the customer base the on-device thesis is being run against.
The comparison exposes a unit-economics story. On-device inference for a narrow model is cheaper per call than the cloud for a large generalist, because the marginal cost of running a 4-billion-parameter model on a phone that is already paid for is electricity, while the marginal cost of running a 27-billion-parameter model in a data center is the next H100 cluster, the next substation, and the next ten-year depreciation schedule. The trillion-dollar GPU build-out of 2024 to 2026 was priced against a world in which every useful AI call goes through a hyperscaler. If the on-device curve crosses first, that capex is on a depreciation clock the cloud is not.
The counterforce is also named. Hyperscalers can drop inference pricing below the on-device break-even if they have to. Closed medical data and regulated clinical workflows are not solved by a phone-resident model alone. Every benchmark in Tether's release is self-reported, not independently reproduced. The reader should treat the +11.42 and the +16.00 as a vendor's claim, not as a settled fact, until a hospital IT team or an independent lab reruns the eval on the published weights. Ardoino's five-year, 90%-on-device forecast is a CEO bet, not a measured adoption rate. The phone in your pocket is now running the numbers, though, and the GPU order book is the one that should be paying attention.