A 14 expert Delphi consensus panel adapted the NIH's 7 principles for autonomous AI — and split on 4 questions it could not bridge.
A patient consents to a clinical trial. The intervention is an AI that decides on its own which treatment to recommend. Who else is watching, and what rules apply? A 14-person multidisciplinary panel has just spent six months writing down the answer, and flagged exactly where the rulebook still won't hold.
The study, published in the Journal of Medical Internet Research (DOI 10.2196/91472, PMID 42627957), is the first expert-consensus effort to port the NIH's seven principles of ethical clinical research, written in 2000 for human-subject research, to a new kind of trial subject: an autonomous AI. "Autonomous" here means the system makes clinical decisions without a human in the loop, distinguishing it from assistive tools that recommend a course a clinician then accepts or rejects.
To adapt the framework, the authors ran a modified Delphi process, a structured, multi-round expert-consensus method in which panelists answer anonymously, see aggregated results, and refine positions across rounds. Over six months, fourteen purposively selected panelists, drawn from AI, data science, ophthalmology, public policy, law, bioethics, and patient advocacy, worked through two survey rounds and a final virtual meeting. Retention was high: 12 of 14 completed round 1 (85.7%), 10 of 14 completed round 2 (71.4%), and 13 of 14 attended the closing meeting (92.9%).
The mechanism is the Delphi process. Round 1 anchored the panel in a vignette, a written scenario of an autonomous AI tool being tested in a clinical setting, and asked open-ended questions. Those free responses were thematically coded into fifteen statements, which round 2 then rated on a five-point scale. By the end of round 2, the panel had reached strong agreement (≥80%) on nine statements, moderate agreement (60–80%) on two, and remained divided on four. The closing virtual meeting synthesized the outputs into actionable recommendations for trial designers, clinicians, and ethics reviewers.
The four divisive statements are the most useful finding for any working trial designer. A 14-person panel is too small to settle a field-wide question by count alone, but the structure of their disagreement is itself a map of where the rulebook is still being argued out. The paper's framing names the central fault line: in an autonomous AI trial, performance varies across clinical settings and across stakeholders. The same model can behave differently at a teaching hospital than at a rural clinic, and a patient, a clinician, and a regulator can read the same trial result three different ways.
Those are exactly the conditions that make informed consent, liability, and post-trial monitoring hard to standardize. The fault line the paper draws, performance variation across clinical settings and stakeholders, lands on the same ground any working trial designer will recognize: where to disclose what the model will do, who is responsible when it behaves differently across sites, and how to monitor it once it ships.
For an ethics committee asked to review an autonomous AI protocol, that structure is more useful than a tidy consensus. Nine statements the panel agrees on are portable enough to lean against; four that the panel could not bridge tell a reviewer exactly which questions the field has not yet answered.
For trial designers, the takeaway is operational. Before an autonomous AI protocol clears ethics review, the protocol will have to commit to specific positions on consent, on liability, and on how the model is monitored once performance drifts between sites. The 14-person panel has done the harder half of the work. It has written down what the field agrees on. The other half is still open, and the paper marks exactly where.