Journal List > Ann Surg Treat Res > v.110(3) > 1516094790

Oh, Jung, and Choi: Physical AI goes to the operating room: are we ready for the Surgical Data Factory?

Abstract

The operating room remains a paradox: it is one of the most sensor-rich environments in the hospital, yet it produces largely underutilized data. While surgical artificial intelligence (AI) has achieved remarkable progress in recent years, the day-to-day practice of surgery has changed little, with most systems confined to passive decision support. This narrative review traces the evolution of surgical AI from perception to cognition to early forms of action, arguing that the next paradigm shift requires “physical AI”—systems capable of meaningful physical interaction and autonomous execution. The clinical motivation for pursuing physical AI is clear: surgical outcomes vary substantially across surgeons, access is constrained by workforce shortages, and high-quality care remains tied to the scarcity of human expertise. If reliable autonomous systems can be developed, surgery could become more standardized, scalable, and reproducible. However, a critical bottleneck persists: the scarcity of synchronized, multimodal training data. The fundamental barrier is environmental rather than algorithmic, as most operating rooms are not configured to measure surgical practice objectively. We propose reconceptualizing the operating room as a “Surgical Data Factory”—a closed-loop ecosystem designed to capture multimodal signals, structure them via consensus taxonomies linked to outcomes, and utilize them for training, validation, and monitoring. Surgeons must transition from passive users to active architects of this infrastructure. Investing in systematic data governance is the prerequisite for responsibly developing, validating, and scaling physical AI in surgery.

INTRODUCTION

While modern medicine relies heavily on data, the operating room (OR) remains a paradox: it is one of the most sensor-rich environments in the hospital yet still produces relatively underutilized data [12]. Every surgery generates streams of endoscopic video, instrument kinematics, device logs, and hemodynamic signals, but much of this information is inconsistently recorded, poorly synchronized, or discarded once the procedure ends [3]. Rather than functioning as a structured learning environment, the OR has historically operated as an isolated domain where most data are not retained or organized in a form amenable to analysis [4]. Consequently, expertise remains largely confined to individual surgeons, hindering scalable standardization [56].
Over the past decade, surgical artificial intelligence (AI) has progressed through distinct technical stages—from perception systems that detect instruments and segment anatomy, to cognitive models that interpret surgical context and anticipate events, and now toward physical AI driven by vision–language–action (VLA) models capable of generating control commands for robotic effectors (Fig. 1) [7891011]. Despite these advances, the day-to-day practice of surgery has changed little. Most deployed systems function as passive decision-support or documentation tools, and the step from algorithmic achievement on curated datasets to genuine transformation of how operations are conducted has rarely been realized [210]. For AI to tangibly transform surgical practice, it must cross the fundamental barrier separating passive decision support from active physical interaction.
This capability defines the emerging field of “physical AI”, where computational intelligence is embodied in robotic systems to directly manipulate the real world [12]. The motivation for pursuing physical AI is fundamentally clinical and societal: surgical outcomes vary substantially across surgeons, access is constrained by workforce shortages, and high-quality care remains tied to the time and physical presence of individual experts [56131415]. If reliable autonomous or semi-autonomous systems can be developed, surgery could become more standardized, scalable, and reproducible—partly decoupled from the scarcity of human expertise [11]. Realizing such systems, however, requires large-scale, synchronized multimodal data—vision, kinematics, and clinical intent—for which current ORs lack the infrastructure to produce systematically.
This narrative review traces the evolution of surgical AI from perception to cognition to early forms of action, examining both achievements and limitations at each stage. We then argue that the prerequisite for advancing physical AI is not algorithmic refinement alone, but the transformation of the OR into a “Surgical Data Factory” capable of generating the multimodal data and experience base required for safe, standardized tool-level autonomy.

MAIN BODY

Evolution of surgical artificial intelligence toward physical artificial intelligence

The era of perception: seeing without understanding

For much of the past decade, surgical AI has been dominated by computer vision systems designed to “perceive”the operative field [78916]. Early work focused on discriminative tasks—detecting instruments, segmenting anatomical structures, and classifying surgical phases from laparoscopic video [17,1819202122]. Convolutional neural networks trained on curated datasets enabled substantial gains in accuracy, establishing that surgical video could be treated as a useful signal for automation rather than as a discarded by-product of the case. Most of these systems, however, relied almost exclusively on the video stream, without concurrent signals such as robotic kinematics, device logs, physiologic traces, and team communication.
These perception models also remain fundamentally limited by their lack of semantic depth [23]. A segmentation algorithm may delineate the gallbladder from the liver bed with pixel-level precision, yet it has no understanding of why the dissection plane is chosen, whether the applied traction is excessive, or whether the maneuver is progressing toward a safe configuration. In this paradigm, AI functions as a high-resolution sensor rather than a participant in clinical reasoning. It highlights structures and counts tools, but it does not represent the underlying goals, trade-offs, or risk thresholds that guide expert behavior. Consequently, despite their quantitative success, most perception systems remain confined to retrospective analysis and research, rarely influencing real-time decision-making [1718192021222425].

The era of cognition: understanding beyond seeing

As perception models matured, attention shifted from what is in the field of view to what it means in the procedural context [2627]. Transformer-based architectures and vision–language models (VLMs) made it possible to align visual signals with natural-language representations, moving from fixed label sets toward open-ended semantic descriptions [282930]. General-purpose VLMs adapted to surgery (e.g., GP-VLS) and domain-specific models (e.g., SurgVLM) exemplify this trend [3132]. These models treat surgical videos not merely as collections of frames but as narratives, in which tools, structures, and actions are embedded in temporally coherent stories [31323334].
Proof-of-concept studies have demonstrated several capabilities. VLMs fine-tuned on laparoscopic data can generate structured descriptions of the operative field, identify the short-term goal of a maneuver, and respond to context-specific questions—such as whether the critical view of safety has been achieved or which region is responsible for bleeding [3132333435363738]. Models pretrained on surgical video lectures have acquired procedural knowledge without explicit frame-level annotation [3335]. Recent work has extended vision–language representations to action prediction, estimating the next likely action or proposing multi-step sequences given a language-specified goal [37].
Current models, however, remain confined to digital reasoning. They can label, describe, summarize, and recommend, but do not directly actuate instruments. Most evaluations are retrospective and single-center, limited to held-out videos or surrogate tasks such as question answering. Prospective evidence of impact on intraoperative decisions or patient outcomes in routine clinical settings remains rare and preliminary [3132333435363738].

Early steps toward action: physical artificial intelligence in the laboratory

As surgical AI systems increasingly acquire “cognitive” capabilities, the frontier shifts to how these models can support action in the physical world [11]. Unlike conventional robotic surgery, where platforms operate in a strict master–slave configuration with no autonomous decision-making, emerging physical AI approaches introduce learned policies that sit between human intent and robotic execution [3940414243]. By processing multimodal inputs—video, kinematics, and language—these models aim to output control commands directly, analogous to how residents acquire technical skills by observing and emulating expert procedures [44].
The first credible demonstrations of this action-level intelligence appeared in supervised autonomous soft-tissue suturing systems, notably the smart tissue anastomosis robot (STAR) [4041]. By combining advanced imaging with motion planning, STAR demonstrated that a robot could plan and execute suturing trajectories on deformable tissue, achieving leak pressures comparable to expert surgeons. However, the difficulty of generalizing such explicit, model-based approaches beyond narrow tasks prompted a shift toward more flexible, learning-based architectures. Inspired by large-scale sequence models in general robotics, researchers have thus begun applying transformer architectures to surgical manipulation, treating robot actions as token sequences learned directly from demonstrations [45464748]. A prime example is the Surgical Robot Transformer (SRT) family [43]. Its hierarchical variant, SRT-H, represents a significant leap toward step-level autonomy [42]. In this framework, the surgeon specifies a high-level goal (e.g., “clip and divide the cystic duct”), which the model decomposes into subtasks—localizing the target, positioning the instrument, and deploying the clip—and executes autonomously in ex vivo porcine models. This mirrors the cognitive workflow of a trainee mapping high-level intent to low-level motor primitives (Fig. 2).
Beyond dedicated surgical platforms, the scope of physical AI is expanding to collaborative and general-purpose roles through VLA models. Systems like RoboNurse demonstrate how a robotic agent can ground verbal commands in visual affordances to perform precise instrument handovers [49]. Furthermore, humanoid robots are emerging as versatile embodiments; the LapSurgie framework utilizes a humanoid to teleoperate laparoscopic instruments, and hospital-oriented humanoids have demonstrated feasibility in tasks ranging from intubation to ultrasound-guided injections [5051]. These examples illustrate a shift toward general-purpose embodiments capable of translating high-level prompts into physical action without extensive hand-crafted control logic.
Despite these technical milestones, a critical bottleneck remains: a persistent “data gap” driven by the scarcity of synchronized, multimodal, high-quality training data. To mitigate this, many groups have turned to physics-based simulations for safe reinforcement and imitation learning [525354]. While some studies show zero-shot transfer to real robots, a persistent “simulation to reality (sim-to-real) gap” limits their utility; simulations cannot fully capture deformable tissue, dynamic fluids, or true anatomical variation. Consequently, we have effectively built the “engines” for physical AI—architectures capable of mapping observations to actions—but they are currently idling due to a lack of fuel. Most policies remain trained on narrow datasets or synthetic environments, with minimal exposure to the chaotic reality of human surgery. To move from isolated prototypes to reliable clinical assistants, the field must confront this deficit by reengineering the OR into a Surgical Data Factory that systematically captures the multimodal experience required to power these advanced systems.

The fundamental bottleneck: environmental and data deficits

Despite progress from perception through cognition to action, these advances have not translated into clinical transformation [226]. The fundamental barrier is not algorithmic but environmental: most ORs are not configured to measure surgical practice objectively [3455]. Without systematic measurement, the data required to train, validate, and monitor physical AI do not exist at scale [4]. Other barriers compound this problem—sparse prospective evidence, underdeveloped regulatory frameworks, untested liability allocation, and poorly characterized human factors. These obstacles are real, but none can be systematically studied without first establishing measurement infrastructure. The data deficit is not one gap among many; it is the prerequisite that gates progress on all others. At the same time, the clinical motivations for physical AI—reducing unwarranted variation, alleviating workforce constraints, and making high-quality surgery more scalable—remain compelling; they simply cannot be meaningfully pursued until this measurement deficit is addressed [56131415].
Current surgical data suffer from fragmentation and lack of semantic depth. Signals are captured in isolation: video may be stored selectively, but kinematics, device logs, and physiologic monitoring remain in proprietary silos with incompatible formats [23]. Critically, there are few scalable methods to capture instrument movements during non-robotic procedures, where kinematics must be inferred from video rather than measured directly—a significant limitation given that most procedures worldwide remain non-robotic [565758]. Semantic depth is equally absent. No widely adopted ontology defines phases, steps, and error modes consistently across institutions. Without shared taxonomies, recordings cannot become structured, comparable measures [59]. Ground truth linking intraoperative events to outcomes is sparse [60]. Most operations leave behind only narrative notes—from the standpoint of objective measurement, the procedure is largely invisible [61].

The Surgical Data Factory and clinical roadmap

Operating room as a data factory: capture-structuring-use

Addressing these deficiencies requires reconceptualizing the OR as a Surgical Data Factory—a closed-loop ecosystem converting operative experience into structured, reusable data [124]. Three functions are essential (Fig. 3). Capture enables routine, synchronized acquisition of multimodal streams: video, kinematics, device logs, physiologic signals, and where appropriate, team audio. Structuring transforms raw streams into computable units through consensus taxonomies, common data models, and linkage to outcomes—enabling models to learn not only what was done but what it led to. Use closes the loop: datasets train and validate policies, simulation augments coverage of rare complications, and deployed systems are continuously monitored. For physical AI, this infrastructure must deliver expert demonstrations at scale, systematic exposure to corner cases, and mechanisms for detecting performance drift. Without such a factory, physical AI will remain confined to laboratory prototypes.

Clinical roadmap for physical artificial intelligence implementation

Realizing this vision is fundamentally a clinical challenge. Surgeons should move from passive users to active architects of the measurement ecosystem. First, surgeons must lead the definition of procedural taxonomies, adopting the SAGES consensus hierarchy as a baseline while expanding it to establish the new operational concepts essential for physical AI. Second, surgeons must shape the consensus on multimodal data acquisition, ensuring the inclusion of instrument kinematics and physiologic signals alongside video to build a comprehensive data foundation. Third, to ensure that the pursuit of robotic autonomy does not compromise surgical autonomy, surgeons must strictly define operational boundaries, specifying the conditions permitting AI action and the criteria for mandatory human takeover. Fourth, evaluation frameworks must prioritize patient-centered outcomes over technical metrics, establishing staged prospective validation as a prerequisite for deployment. Fifth, the surgical community must advocate for interoperability and governance structures that enable data sharing, explicitly positioning continuous recording as quality improvement rather than punitive surveillance. The Surgical Data Factory is a conceptual framework requiring empirical validation, and practical constraints—cost, vendor lock-in, regulatory barriers—may prove substantial.

CONCLUSION

With these prerequisites in place, physical AI can move surgical AI from interpretation to action, ultimately enabling more standardized, scalable, and reproducible surgical care. While advances in model architectures are necessary, from a surgeon’s perspective the more immediate prerequisite is to make surgery measurable through routine, synchronized capture of multimodal signals and shared taxonomies linked to outcomes. The Surgical Data Factory provides a pragmatic closed-loop approach to convert operative experience into structured, reusable evidence. Investing now in measurement and governance is therefore the most practical starting point for responsibly testing, validating, and scaling physical AI in real ORs.

ACKNOWLEDGEMENTS

We gratefully acknowledge Haeun Kim for her expert artistic contributions to the visualization of the figures presented in this article. During the preparation of this work, the authors used Google Gemini (Google LLC) in order to improve the language and readability of the manuscript. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Notes

Fund/Grant Support: This research was supported by a grant of Korean ARPA-H Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (RS-2025-25424639), the Bio&Medical Technology Development Program of the National Research Foundation (NRF) funded by the Korean government (MIST) [RS-2024-00392495], and the Future Medicine 2030 Project of the Samsung Medical Center [SMX1230771].

Conflicts of Interest: No potential conflict of interest relevant to this article was reported.

Author Contribution:

  • Conceptualization: All authors.

  • Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Visualization: NO.

  • Supervision: KHJ, GSC.

  • Writing – Original Draft: NO.

  • Writing – Review & Editing: All authors.

References

1. Maier-Hein L, Vedula SS, Speidel S, Navab N, Kikinis R, Park A, et al. Surgical data science for next-generation interventions. Nat Biomed Eng. 2017; 1:691–696. PMID: 31015666.
crossref
2. Maier-Hein L, Eisenmann M, Sarikaya D, März K, Collins T, Malpani A, et al. Surgical data science: from concepts toward clinical translation. Med Image Anal. 2022; 76:102306. PMID: 34879287.
3. Mascagni P, Padoy N. OR black box and surgical control tower: recording and streaming data and analytics to improve surgical care. J Visc Surg. 2021; 158:S18–S25. PMID: 33712411.
crossref
4. Garcia Vazquez A, Verde J, Hernandez Lara A, Mutter D, Swanstrom L. Consensus for operating room multimodal data management: identifying research priorities for data-driven surgery. Ann Surg Open. 2024; 5:e459. PMID: 39310343.
5. Birkmeyer JD, Finks JF, O'Reilly A, Oerline M, Carlin AM, Nunn AR, et al. Surgical skill and complication rates after bariatric surgery. N Engl J Med. 2013; 369:1434–1442. PMID: 24106936.
crossref
6. Stulberg JJ, Huang R, Kreutzer L, Ban K, Champagne BJ, Steele SR, et al. Association between surgeon technical skills and patient outcomes. JAMA Surg. 2020; 155:960–968. PMID: 32838425.
crossref
7. Hashimoto DA, Rosman G, Rus D, Meireles OR. Artificial intelligence in surgery: promises and perils. Ann Surg. 2018; 268:70–76. PMID: 29389679.
crossref
8. Ward TM, Mascagni P, Ban Y, Rosman G, Padoy N, Meireles O, et al. Computer vision in surgery. Surgery. 2021; 169:1253–1256. PMID: 33272610.
9. Kitaguchi D, Takeshita N, Hasegawa H, Ito M. Artificial intelligence-based computer vision in surgery: recent advances and future perspectives. Ann Gastroenterol Surg. 2022; 6:29–36. PMID: 35106412.
crossref
10. Loftus TJ, Altieri MS, Balch JA, Abbott KL, Choi J, Marwaha JS, et al. Artificial intelligence-enabled decision support in surgery: state-of-the-art and future directions. Ann Surg. 2023; 278:51–58. PMID: 36942574.
11. Schmidgall S, Opfermann JD, Kim JW, Krieger A. Will your next surgeon be a robot?: autonomy and AI in robotic surgery. Sci Robot. 2025; 10:eadt0187. PMID: 40700524.
crossref
12. Liu D, Zhang J, Dinh AD, Park E, Zhang S, Mian A, et al. Generative physical ai in vision: a survey. arXiv [Preprint]. 2025; 01. 19. DOI: 10.48550/arXiv.2501.10928.
13. IHS Markit. The complexities of physician supply and demand: projections from 2015 to 2030 [Internet]. Association of American Medical Colleges;2017. cited 2026 Jan 8. Available from: https://www.aamc.org/media/8816/download.
14. Meshesha BR, Sibhatu MK, Beshir HM, Zewude WC, Taye DB, Getachew EM, et al. Access to surgical care in Ethiopia: a cross-sectional retrospective data review. BMC Health Serv Res. 2022; 22:973. PMID: 35907955.
crossref
15. Asano S, Kunisawa S, Imanaka Y. Development of a new indicator of surgical burden using person-time-adjusted surgical volume: a cross-sectional regional analysis in Japan. BMJ Public Health. 2025; 3:e002720. PMID: 40791263.
crossref
16. Mascagni P, Alapatt D, Sestini L, Altieri MS, Madani A, Watanabe Y, et al. Computer vision in surgery: from potential to clinical value. NPJ Digit Med. 2022; 5:163. PMID: 36307544.
crossref
17. Twinanda AP, Shehata S, Mutter D, Marescaux J, de Mathelin M, Padoy N. EndoNet: a deep architecture for recognition tasks on laparoscopic videos. IEEE Trans Med Imaging. 2017; 36:86–97. PMID: 27455522.
crossref
18. Kitaguchi D, Takeshita N, Matsuzaki H, Takano H, Owada Y, Enomoto T, et al. Real-time automatic surgical phase recognition in laparoscopic sigmoidectomy using the convolutional neural network-based deep learning approach. Surg Endosc. 2020; 34:4924–4931. PMID: 31797047.
crossref
19. Kitaguchi D, Lee Y, Hayashi K, Nakajima K, Kojima S, Hasegawa H, et al. Development and validation of a model for laparoscopic colorectal surgical instrument recognition using convolutional neural network-based instance segmentation and videos of laparoscopic procedures. JAMA Netw Open. 2022; 5:e2226265. PMID: 35984660.
crossref
20. Madani A, Namazi B, Altieri MS, Hashimoto DA, Rivera AM, Pucher PH, et al. Artificial intelligence for intraoperative guidance: using semantic segmentation to identify surgical anatomy during laparoscopic cholecystectomy. Ann Surg. 2022; 276:363–369. PMID: 33196488.
21. Oh N, Kim B, Kim T, Rhu J, Kim J, Choi GS. Real-time segmentation of biliary structure in pure laparoscopic donor hepatectomy. Sci Rep. 2024; 14:22508. PMID: 39341910.
crossref
22. Oh N, Lim M, Kim B, Shin J, Park S, Rhu J, et al. AI-assisted intraoperative navigation for safe right liver mobilization in pure laparoscopic donor hepatectomy: an experimental multi-institutional validation study. Sci Rep. 2025; 15:27935. PMID: 40744949.
23. Xu Y, Vaziri-Pashkam M. Limits to visual representational correspondence between convolutional neural networks and the human brain. Nat Commun. 2021; 12:2065. PMID: 33824315.
crossref
24. Czempiel T, Paschali M, Keicher M, Simson W, Feussner H, Kim ST, et al. TeCNO: surgical phase recognition with multistage temporal convolutional networks. In : Martel AL, Abolmaesumi P, Stoyanov D, editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2020. Lecture Notes in Computer Science; Springer;2020. 12263:p. 343–352.
crossref
25. Kitaguchi D, Kosugi N, Ishikawa Y, Narihiro S, Enomoto T, Oda T, et al. A multicentre randomized controlled trial exploring the clinical usefulness of the intraoperative use of an artificial intelligence-based anatomical navigation system. Br J Surg. 2025; 112.
crossref
26. Varghese C, Harrison EM, O'Grady G, Topol EJ. Artificial intelligence in surgery. Nat Med. 2024; 30:1257–1268. PMID: 38740998.
crossref
27. Khan U, Nawaz U, Qayyum A, Ashraf S, Xie Y, Khan MH, et al. Surgical scene understanding in the era of foundation ai models: a comprehensive review. arXiv [Preprint]. 2025; 02. 16. DOI: 10.48550/arXiv.2502.14886.
28. Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. In : Meila M, Zhang T, editors. Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research (PMLR). Vol 139; PMLR;2021. p. 8748–8763.
29. Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst. 2022; 35:23716–23736.
crossref
30. Li J, Li D, Xiong C, Hoi S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In : Chaudhuri K, Jegelka S, Song L, Szepesvári C, Niu G, Sabato S, editors. Proceedings of the 39th International Conference on Machine Learning. Proceedings of Machine Learning Research (PMLR). Vol 162; PMLR;2022. p. 12888–12900.
31. Schmidgall S, Cho J, Zakka C, Hiesinger W. Gp-vls: a general-purpose vision language model for surgery. arXiv [Preprint]. 2024; 07. 27. DOI: 10.48550/arXiv.2407.19305.
32. Zeng Z, Zhuo Z, Jia X, Zhang E, Wu J, Zhang J, et al. Surgvlm: a large vision-language model and systematic evaluation benchmark for surgical intelligence. arXiv [Preprint]. 2025; 06. 03. DOI: 10.48550/arXiv.2506.02555.
33. Yuan K, Navab N, Padoy N. Procedure-aware surgical video-language pretraining with hierarchical knowledge augmentation. Adv Neural Inf Process Syst. 2024; 37:122952–122983.
34. Jeon Y, Park S, Shin J, Park K, Kim B, Oh N, Jung KH. SurGen-Net: a generative approach for surgical VQA with structured text generation. In : Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops; IEEE;2025. p. 1292–1299.
35. Yuan K, Srivastav V, Yu T, Lavanchy JL, Marescaux J, Mascagni P, et al. Learning multi-modal representations by watching hundreds of surgical video lectures. Med Image Anal. 2025; 105:103644. PMID: 40513506.
crossref
36. Stueker EH, Kolbinger FR, Saldanha OL, Digomann D, Pistorius S, Oehme F, et al. Vision-language models for automated video analysis and documentation in laparoscopic surgery: a proof-of-concept study. Int J Surg. 2025; 111:7777–7786. PMID: 40679978.
crossref
37. Xu M, Huang Z, Zhang J, Zhang X, Dou Q. Surgical action planning with large language models. In : Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer;2025. p. 563–572.
38. Bogaert W, Carl N, Kowalewski KF, Michel MS, Mottrie A, De Backer P. Bridging vision and text: applications and challenges of vision-language models in urological surgery. Eur Urol Focus. 2025; 11:18–21. PMID: 40382217.
39. Lanfranco AR, Castellanos AE, Desai JP, Meyers WC. Robotic surgery: a current perspective. Ann Surg. 2004; 239:14–21. PMID: 14685095.
40. Leonard S, Wu KL, Kim Y, Krieger A, Kim PC. Smart tissue anastomosis robot (STAR): a vision-guided robotics system for laparoscopic suturing. IEEE Trans Biomed Eng. 2014; 61:1305–1317. PMID: 24658254.
crossref
41. Saeidi H, Opfermann JD, Kam M, Raghunathan S, Leonard S, Krieger A. A confidence-based shared control strategy for the smart tissue autonomous robot (STAR). Rep U S. 2018; 2018:1268–1275. PMID: 31475075.
crossref
42. Kim JW, Chen JT, Hansen P, Shi LX, Goldenberg A, Schmidgall S, et al. SRT-H: a hierarchical framework for autonomous surgery via language-conditioned imitation learning. Sci Robot. 2025; 10:eadt5254. PMID: 40632876.
crossref
43. Kim JW, Zhao TZ, Schmidgall S, Deguet A, Kobilarov M, Finn C, et al. Surgical Robot Transformer (SRT): imitation learning for surgical tasks. In : Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the 8th Conference on Robot Learning. Proceedings of Machine Learning Research (PMLR). Vol 270; PMLR;2025. p. 130–144.
44. Schmidgall S, Kim JW, Kuntz A, Ghazi AE, Krieger A. General-purpose foundation models for increased autonomy in robot-assisted surgery. Nat Mach Intell. 2024; 6:1275–1283.
crossref
45. Brohan A, Brown N, Carbajal J, Chebotar Y, Dabis J, Finn C, et al. Rt-1: robotics transformer for real-world control at scale. arXiv [Preprint]. 2022; 12. 13. DOI: 10.48550/arXiv.2212.06817.
crossref
46. Zitkovich B, Yu T, Xu S, Xu P, Xiao T, Xia F, et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In : Tan J, Toussaint M, Darvish K, editors. Proceedings of The 7th Conference on Robot Learning. Proceedings of Machine Learning Research (PMLR). Vol 229; PMLR;2023. p. 2165–2183.
47. Kim MJ, Pertsch K, Karamcheti S, Xiao T, Balakrishna A, Nair S, et al. Openvla: an open-source vision-language-action model. arXiv [Preprint]. 2024; 06. 13. DOI: 10.48550/arXiv.2406.09246.
48. Team OM, Ghosh D, Walke H, Pertsch K, Black K, Mees O, et al. Octo: an open-source generalist robot policy. arXiv [Preprint]. 2024; 05. 20. DOI: 10.48550/arXiv.2405.12213.
crossref
49. Li S, Wang J, Dai R, Ma W, Ng WY, Hu Y, et al. RoboNurse-VLA: robotic scrub nurse system based on vision-language-action model. arXiv [Preprint]. 2024; 09. 29. DOI: 10.48550/arXiv.2409.19590.
crossref
50. Liang Z, Liang X, Atar S, Das S, Chiu Z, Zhang P, et al. LapSurgie: humanoid robots performing surgery via teleoperated handheld laparoscopy. arXiv [Preprint]. 2025; 10. 03. DOI: 10.48550/arXiv.2510.03529.
51. Atar S, Liang X, Joyce C, Richter F, Ricardo W, Goldberg C, et al. Humanoids in hospitals: a technical study of humanoid surrogates for dexterous medical interventions. arXiv [Preprint]. 2025; 03. 17. DOI: 10.48550/arXiv.2503.12725.
52. Xu J, Li B, Lu B, Liu Y-H, Dou Q, Heng P-A. SurROL: an open-source reinforcement learning centered and DVRK compatible platform for surgical robot learning. In : 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE;2021. p. 1821–1828.
53. Yang Z, Long Y, Chen K, Wei W, Dou Q. Efficient physically-based simulation of soft bodies in embodied environment for surgical robot. arXiv [Preprint]. 2024; 02. 02. DOI: 10.48550/arXiv.2402.01181.
54. Zbinden L, Nelson N, Chen JT, Chen X, Kim JW, Azizian M, et al. Cosmos-Surg-dVRK: world foundation model-based automated online evaluation of surgical robot policy learning. arXiv [Preprint]. 2025; 10. 17. DOI: 10.48550/arXiv.2510.16240.
55. Cheikh Youssef S, Haram K, Noël J, Patel V, Porter J, Dasgupta P, et al. Evolution of the digital operating room: the place of video technology in surgery. Langenbecks Arch Surg. 2023; 408:95. PMID: 36807211.
crossref
56. Nema S, Vachhani L. Surgical instrument detection and tracking technologies: automating dataset labeling for surgical skill assessment. Front Robot AI. 2022; 9:1030846. PMID: 36405072.
crossref
57. Heiliger C, Andrade D, Geister C, Winkler A, Ahmed K, Deodati A, et al. Tracking and evaluating motion skills in laparoscopy with inertial sensors. Surg Endosc. 2023; 37:5274–5284. PMID: 36976421.
crossref
58. Ferrari D, Violante T, Novelli M, Starlinger PP, Smoot RL, Reisenauer JS, et al. The death of laparoscopy. Surg Endosc. 2024; 38:2677–2688. PMID: 38519609.
crossref
59. Meireles OR, Rosman G, Altieri MS, Carin L, Hager G, Madani A, et al. SAGES consensus recommendations on an annotation framework for surgical video. Surg Endosc. 2021; 35:4918–4929. PMID: 34231065.
60. Bose R, Nwoye CI, Lazo JF, Lavanchy JL, Padoy N. Feature mixing approach for detecting intraoperative adverse events in laparoscopic Roux-en-Y gastric bypass surgery. In : Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer;2025. p. 178–188.
61. Lefter LP, Walker SR, Dewhurst F, Turner RW. An audit of operative notes: facts and ways to improve. ANZ J Surg. 2008; 78:800–802. PMID: 18844913.
crossref
Fig. 1

The evolutionary stages of surgical artificial intelligence (AI): from perception to physical action. This illustrates the progression of surgical AI across 3 distinct eras, parallel to the functional areas of the human brain. (Left) Era 1: Perception. AI focuses on computer vision tasks such as instrument detection and anatomical segmentation, serving as a high-resolution sensor without contextual understanding. (Middle) Era 2: Cognition. Vision-language models enable the system to interpret surgical context, reason about safety (e.g., Critical View of Safety), and communicate via natural language, acting as a cognitive partner. (Right) Era 3: Action (physical AI). The emerging frontier where Vision-language-action models translate high-level surgeon intent into physical robotic control commands. Note that the AI plans and executes maneuvers under human supervision, bridging the gap between digital reasoning and physical interaction.

astr-110-135-g001
Fig. 2

Parallel architectures of surgical intelligence: bridging human and physical artificial intelligence (AI). This diagram illustrates the biomimetic correspondence between the surgeon’s biological loop and the Physical AI architecture. (Left) The human surgeon relies on eyes for perception, the brain for planning based on experience, and hands for tissue manipulation. (Right) Embodied AI replicates this workflow using vision encoders for input, multimodal foundation models (vision-language-action, VLA) for reasoning, and robotic agents for kinematic control. Both entities operate within a shared surgical reality, aiming to standardize surgical outcomes through physical interaction.

astr-110-135-g002
Fig. 3

The ecosystem of the Surgical Data Factory and physical artificial intelligence (AI). The figure illustrates the closed-loop transformation of surgical practice. (Bottom) Distributed data acquisition: Real-world operating rooms function as a “Surgical Data Factory,” systematically capturing massive, multimodal datasets from human-performed surgeries. (Middle) Semantic distillation: Raw data streams are processed and distilled into role-specific AI agents (operator, assistant, anesthesiologist, scrub nurse), learning the distinct behaviors and interactions of the surgical team. (Top) Digital twin and physical AI: The accumulated intelligence constructs a high-fidelity digital Twin, enabling the training of physical AI agents. In this future paradigm, the surgeon evolves from a manual operator to an “architect,” orchestrating autonomous robotic systems within a verified, data-driven environment.

astr-110-135-g003
TOOLS
Similar articles