EndoLIFT, a preprint AI model for surgical endoscopes, splits the surgeon's instruction from the steering so a static view no longer freezes the scope; the preprint has been tested only on lab tissue simulators and pig tissue.
Inside the colon, the view hasn't changed for a second. The surgeon says "stop and pull back," but a moment ago the same view came with "advance." A new preprint, EndoLIFT, argues the right move in that moment cannot come from the picture alone. It has to come from the instruction.
The paper calls this failure mode "intent aliasing": the same observation maps to opposite axial actions depending on which command the operator gave. Vision-language-action models for surgical robots usually fold that command into the same channel as continuous control, so a static camera frame can stall the scope or send it the wrong way.
EndoLIFT splits the problem in two. Language conditions the action as a discrete intent switch; a separate 32-dimensional trajectory latent handles continuous control. In controlled same-view instruction swaps, the model picks the correct axial direction even without the trajectory latent: language alone resolves the ambiguity. Against a matched baseline without latent conditioning, the system gains 11.1 percentage points in navigation-direction accuracy and cuts wrong-direction advances by 83%. Across 44 held-out phrasings of the same intent, it follows the instruction 82.8% of the time.
The numbers come from a preprint. The closed-loop evaluation ran on colon, lung, and stomach phantoms, plus 10 of 10 ex-vivo porcine trachea trials. No human patients. The gap from phantom tissue to the operating room is years, not months, and the clinical payoff, if any, is in procedure quality and precision, not in lives saved.