If an AI model can render a convincing surgical scene, can it also predict a valid surgical sequence? A paper published in npj Digital Medicine evaluates videos generated by Veo-3 and Wan2.2 across four dimensions: appearance, instrument operation, environmental response, and surgical intent. Both models produced visually persuasive footage but received substantially lower scores for surgical logic. The findings reveal a gap between realistic video generation and dependable specialist simulation.
What Does a Realistic Scene Prove?
A surgical instrument appears sharp, tissue textures look natural, and motion is smooth. The result may resemble an authentic recording. Yet visual realism alone cannot tell us whether the instrument should move in that direction or whether the tissue responds appropriately.
The distinction matters when video generation models are described as world models. A world model is expected to represent how an environment changes and how actions lead to consequences. Rendering a plausible image and predicting a valid sequence of events are related, but they are not identical achievements.
Surgery offers a demanding test. Instrument manipulation, soft tissue deformation, fluid movement, and procedural goals interact. A movement that a general viewer might overlook can constitute a consequential procedural error.
To examine this problem, Zhen Chen and colleagues introduced the SurgVeo benchmark and the Surgical Plausibility Pyramid. Their paper was published online in npj Digital Medicine on September 26, 2026.
Four Levels of Surgical Plausibility
The researchers did not treat video quality as one overall impression. Their framework separates four levels of evaluation.
Dimension | What the evaluators assess |
|---|---|
Visual perceptual plausibility | Image clarity, stability, and the appearance of the scene |
Instrument operation plausibility | The accuracy of surgical instrument movements |
Environment feedback plausibility | How tissue and the scene respond to those movements |
Surgical intent plausibility | Whether actions serve an appropriate procedural goal |
Adapted from the Surgical Plausibility Pyramid in Chen and colleagues’ paper.
The questions become more demanding as the evaluation moves upward. Does the scene look like surgery? Is the instrument being used correctly? Does the action produce an appropriate consequence? Does the sequence make sense at this stage of the procedure?
Success on one level does not guarantee success on another. An instrument may look correct but move in the wrong direction. Tissue may look realistic yet respond incorrectly to suction. Individual movements may appear possible while the sequence fails to serve a coherent surgical goal.
Separating these dimensions allows the researchers to identify where a generated video fails. A single visual quality score could conceal those differences.
The Same Task Was Given to Two Models
SurgVeo contains 50 clips from laparoscopic hysterectomy and endoscopic pituitary surgery: 18 from the former and 32 from the latter. The source material came from the public AutoLaparo and PitVis repositories, with three independent procedure recordings used for each surgical track.
For every sample, the researchers provided a starting frame and a text prompt. The model then generated an eight-second continuation. They evaluated the commercial Veo-3 model and the open Wan2.2-I2V-A14B model without fine-tuning either on surgical videos.
Each model received two types of prompts. The baseline condition specified the type of procedure; the stage-aware condition also named the current surgical stage. Both models were assessed using the same starting frames, prompts, and evaluation process.
Four board-certified surgeons scored the outputs. Two gastrointestinal surgeons independently evaluated the laparoscopic clips, while two neurosurgeons evaluated the pituitary surgery clips. They rated each dimension on a five-point scale at the one-, three-, and eight-second marks. Each mark represented a cumulative assessment from the beginning of the generated video to that point. The paper reports inter-rater agreement in the moderate-to-excellent range.
Visual Quality and Surgical Intent Diverged
Table 1 of the journal manuscript shows how the scores differed. The following selection compares visual plausibility and surgical intent under the baseline prompt. All numbers are mean scores on a five-point scale.
Model and procedure | Visual plausibility, 1s → 8s | Surgical intent, 1s → 8s |
|---|---|---|
Veo-3 · laparoscopic hysterectomy | 3.72 → 3.56 | 3.11 → 1.61 |
Veo-3 · endoscopic pituitary surgery | 3.88 → 3.41 | 2.03 → 1.12 |
Wan2.2 · laparoscopic hysterectomy | 3.97 → 3.28 | 2.61 → 1.78 |
Wan2.2 · endoscopic pituitary surgery | 3.89 → 2.62 | 3.17 → 1.42 |
Selected mean values adapted from Table 1 of the journal manuscript.
The models did not follow identical trajectories. For Wan2.2’s endoscopic pituitary outputs, visual plausibility also fell noticeably, from 3.89 to 2.62. Across the two models and procedures, however, higher-level surgical plausibility remained substantially weaker than appearance.
The Veo-3 laparoscopic baseline condition offers a particularly clear example. Its environment feedback score declined from 3.06 at one second to 1.64 at eight seconds, while its visual score changed from 3.72 to 3.56. The imagery remained comparatively plausible even as the validity of the depicted consequences weakened.
The eight-second rating covers the entire generated sequence up to that point. It does not mean that the model suddenly failed at precisely the eighth second. Nor should a mean on a five-point scale be converted into an accuracy rate or a probability of clinical failure.
Adding the Surgical Stage Did Not Resolve the Gap
Perhaps the model failed because it lacked context about the current stage of surgery. The researchers tested that possibility by comparing baseline prompts with prompts that explicitly named the stage.
Providing this information did not produce a consistent improvement. For Veo-3 laparoscopic outputs, the eight-second instrument operation score was 1.78 under the baseline prompt and 1.69 under the stage-aware prompt. Wan2.2’s corresponding score rose from 1.81 to 2.00, but its environment feedback score in endoscopic pituitary surgery fell from 1.45 to 1.11.
These findings do not establish that detailed prompting is useless. The study compared two prompt conditions and did not exhaust the possibilities for supplying anatomical, physical, or procedural information. It does show that naming the current surgical stage alone did not reliably close the observed gap.
Most Classified Errors Concerned Surgical Logic
The researchers also examined the types of errors identified by surgeons in Veo-3’s baseline-prompt outputs. Basic visual distortions accounted for 6.2% of classified errors in laparoscopic hysterectomy and 2.8% in endoscopic pituitary surgery. More than 93% in each track involved surgical logic under the authors’ classification.
The paper illustrates several forms of failure: a fabricated surgical instrument, an inappropriate direction of movement, an implausible response to suction, and an action that serves the wrong surgical purpose. These errors might occur in footage that still appears convincing at a glance.
The percentage requires careful interpretation. It does not mean that more than 93% of all generated videos were wrong. It describes the proportion of identified and classified errors related to surgical logic. This particular error breakdown also concerns Veo-3’s baseline outputs, rather than a combined analysis of both models.
What the Results Establish—and What They Do Not
The researchers call the divergence between realistic appearance and weak surgical logic a “plausibility gap.” Their findings indicate that a model intended for specialist simulation must preserve relationships among tools, actions, consequences, and goals. Improving texture and short-range motion alone may not be enough.
At the same time, an error in generated footage cannot definitively prove that a model contains no relevant causal knowledge. The task began with a single frame and asked for an eight-second prediction. A longer preceding sequence, richer anatomical information, or training on specialist data could produce different results; this study does not determine how those alternatives would perform.
The authors acknowledge the study’s scope. It examines two models in zero-shot conditions and two kinds of surgery. The space of possible prompts remains underexplored, and the evaluation depends on specialist judgments. Its results should not be generalized to every video model or surgical setting.
The recorded continuation provides the surgeons with a professional reference, but a different continuation is not automatically invalid. The relevant question is whether an alternative action is surgically plausible, not whether the model reproduces the recording frame for frame.
The Standard for a Simulator Must Be Higher
The study’s larger contribution is to make the standard for evaluating generated video more precise. “It looks real” can be a starting observation. A claim of reliable specialist simulation also requires evidence that actions, consequences, procedural goals, and temporal consistency hold together.
The authors point to future work involving specialist training, surgical knowledge, and models designed with physical constraints in mind. These are research directions proposed in response to the observed gap, not solutions proven effective by this experiment.
The problem revealed by surgical footage leads to a broader question: when a generative model produces a sophisticated scene, which parts of that scene can we trust? SurgVeo begins to answer by measuring appearance separately from the events depicted within it. That separation makes the distance between convincing video and dependable world simulation visible.