The White House science office (OSTP), FDA, and the federal health IT office (ONC) are running a 30 day sprint on benchmarking and evaluation only, with deployment and reimbursement explicitly carved out, so hospital purchasing rules start to
OSTP, and ONC are convening outside experts next week in a one-month sprint to define what counts as an "evaluated" clinical AI tool. The agencies are carving deployment and reimbursement out of the agenda on purpose, and the compressed 30-day window is what turns this from another federal request-for-information exercise into a coordination test with a real deadline.
The White House Office of Science and Technology Policy (OSTP), the Food and Drug Administration, and the Office of the National Coordinator for Health IT (ONC) are jointly running the sprint, according to an invitation reviewed by STAT. The format splits into a written phase and a discussion phase, with a stated goal of producing a "consensus set of principles" by the end of the window.
Federal AI convenings typically run for months. HHS already has a stack of overlapping processes on the same question. In December 2025, ONC published a request for information on AI in clinical care. In April 2026, FDA opened a separate RFI on an AI-enabled pilot for early-phase clinical trials. The agencies updated the public on those efforts in June. Industry comment letters are already on the docket. The Federation of American Hospitals' February 2026 filing on the ASTP-ONC AI RFI runs straight through evaluation, oversight, and adoption.
The sprint is positioned as something narrower: benchmarking and evaluation only, not deployment, coverage, or reimbursement. That scope carve-out is the part of the invitation that matters most for a reader trying to read the next 12 months. It signals, in advance, that whatever ships at the end of the month will be a vocabulary exercise, not a draft regulation. The policy trade analysis of the June HHS update reads the same way: the sprint is intended to harmonize fragments of existing frameworks (FDA's Good Machine Learning Practice, ONC's transparency and decision-support rules, and the NIST AI Risk Management Framework) into a single working definition of "evaluated" that hospitals can put in a request for proposals.
The test is whether the output is tight enough for vendors to cite. If the principles define "evaluated" precisely enough that a clinical-AI supplier can point to a benchmark, an audit trail, and a published evaluation summary, the document becomes a soft floor. Hospital procurement teams adopt it as a checklist. Vendors who meet it get the next contract. Vendors who do not get filtered out before a regulator ever shows up. If the principles come out mushy, the sprint evaporates into the rest of the RFI stack, and the working definition of "evaluated" gets set in procurement language at a few large systems instead of in federal vocabulary.
The invitation has not been published, and STAT's reporting rests on a single document review. The same reporter's same-week coverage of the Kaufman/ASPR nomination delay and the Westlake/SAMHSA pick places the sprint inside an active HHS policy-and-personnel stretch, which raises the stakes for a clean sprint deliverable. A bench of academic methodologists, FDA review staff, and large-hospital informatics leads would produce a different consensus than a bench dominated by vendor compliance leads.
The next date to watch is the close of the discussion phase, when the sprint's draft principles are expected to circulate. If the deliverable names specific benchmark categories (for example, bias and fairness, generalization across sites, post-deployment monitoring) and links them to the existing GMLP and ONC transparency rules, the procurement language starts hardening. If it is a list of values without a test, the sprint has earned a paragraph in the trade press and nothing in next year's RFP templates.