arXiv · 2609.33237
SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation
Abstract
Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Changwei Chen, Xiao Liang, Yinuo Yang, Nicole Shen, Peihan Zhang, Sara Wickenhiser, Zekai Liang, Soofiyan Atar, Michael Yip. 2026-09-27. SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation. https://arxiv.org/abs/2609.33237
Cite the original work for its findings. Save a collection to share your selection of sources.