NeFut Logo NeFut
中 Admin Login

[CS.AI] Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#algorithm #AI #Machine Learning

Surgical procedures follow a phase‑step hierarchy, yet current video‑language models are evaluated with flat metrics that ignore cross‑level coherence and error structure. We address this gap with two contributions: first, SurgHiBench, the inaugural hierarchy‑aware benchmark for surgical video understanding, comprising three tasks that measure recognition, consistency and severity across granularity levels. Second, HyperSurg, a hyperbolic model that enforces phase‑step containment via entailment cones, evaluated on four public datasets covering three procedure types. We compare a general‑purpose CLIP model, a Euclidean surgical model, and HyperSurg. The benchmark reveals that models with identical accuracy can exhibit vastly different error severity, ranging from sibling confusions within the correct phase to unrelated cross‑phase predictions. Hyperbolic geometry shifts predictions toward the correct procedural neighborhood, and the gains scale with the tree‑likeness of each dataset’s annotation hierarchy, offering a principled indicator of when hierarchy‑aware geometry is beneficial.

Review

Original Source: https://arxiv.org/abs/2609.27139

[h] Back to Home