Recovering walking function requires detecting meaningful gait changes across rehabilitation sessions, yet precise 3D measurement is still confined to specialized motion‑capture labs. Small camera arrays and body‑worn inertial sensors broaden accessibility, but reliability varies across joints and over time, allowing sensing failures to masquerade as patient improvement. To address this, we introduce GaitVista, a reliability‑aware measurement layer. It employs a lightweight gating mechanism that assigns joint‑ and frame‑specific visual contributions based on four cues: camera coverage, local visual quality, cross‑modal disagreement, and root‑motion continuity, and exposes these weights for user inspection.
On the TotalCapture benchmark we evaluate GaitVista under seven clean and degraded sensing conditions. The method reduces average full‑body and lower‑body error by 27.7% and 27.8%, respectively, achieving the lowest worst‑condition error among fusion approaches. Compared with condition‑blind baselines, the error gap shrinks from $2.76$–$5.33$ cm to $1.11$ cm, approaching a joint‑frame oracle. On MoVi, using image‑derived keypoints, GaitVista is the only deployable fusion method that improves over both unimodal streams, cutting marker‑supported error by 6.4% relative to the strongest learned fusion baseline. On TotalCapture it boosts bilateral knee‑flexion waveform accuracy by 18.9%.
Raw inertial data from five TotalCapture participants reveal location‑ and time‑varying magnetic disturbances, supporting the reliability premise of our design. Both benchmarks involve neurologically healthy participants in controlled settings and retain participant‑specific IMU calibration; thus we report progress toward accessible gait assessment rather than validated clinical deployment.
Review: GaitVista’s explicit modeling of visual reliability yields substantial error reductions in multimodal fusion, offering an interpretable pathway for low‑cost gait monitoring. Future work should validate its robustness in real‑world clinical scenarios.