Abstract / Summary
The focus of artificial intelligence (AI) research is rapidly moving beyond screen-based generative models toward physical AI, an umbrella paradigm for computational systems with physical embodiments that can perceive, reason, and act directly in the real world. In robotics, physical AI is increasingly being implemented through large behavior models and vision-language-action (VLA) models, which learn continuous motor policies from large-scale demonstration datasets. This transition coincides with a structural crisis in tertiary academic medical centers: chronic shortages of residents and skilled surgical assistants that increase the cognitive and physical burden on operating surgeons and may threaten patient safety. This review examines the core engineering mechanisms and clinical applicability of physical AI-driven surgical assistant robots as a practical response to this workforce gap. The engineering foundation of these robots rests on three pillars: VLA models that integrate multimodal perception and commands within a single end-to-end network, online-learning frameworks that adapt in real time to patient-specific tissue variation, and predictive world models that anticipate surgical workflow. Unlike conventional surgeon-controlled robotic systems that only reproduce a surgeon's hand movements, these platforms function as context-aware intelligent partners, with concrete capabilities such as autonomous adaptive retraction, active suction, and autonomous instrument delivery that may reduce tissue injury and preserve surgical workflow continuity. Realizing this potential will require lower real-time control latency, adaptive fail-safe architectures, and resolution of the legal and ethical questions raised by autonomous manipulation. Collaboration among clinicians, engineers, and industry partners will be essential for addressing these challenges and strengthening competitiveness in the global medical AI industry.