NeFut Logo NeFut
Admin Login

[CS.AI] MedTVL: Harnessing Vision and Language for Medical Time Series Classification

Published at: 2026-09-02 22:00 Last updated: 2026-09-03 02:56
#AI #Machine Learning #Neural

Recent advances in multimodal learning have shown that integrating complementary information sources can improve clinical decision‑making for medical time‑series (MedTS) classification. Most existing approaches only explore bi‑modal interactions between time series and text, leaving the synergy among time series, vision, and language largely untapped. Inspired by clinical practice, which simultaneously relies on numerical assessment, visual inspection, and contextual description, we introduce MedTVL—a text‑guided dual‑pathway architecture designed for MedTS classification.\ \ MedTVL employs a convolutional temporal pathway to capture fine‑grained dynamics from raw numerical sequences, while a transformer‑based visual pathway extracts holistic morphological patterns from images generated from the time series. The cross‑modal and architectural heterogeneity of these pathways provides a comprehensive diagnostic perspective. Adaptive medical text semantics guide both pathways, reducing potential diagnostic ambiguity. A Mixture‑of‑Experts (MoE) module then dynamically routes each instance to specialized fusion experts, allowing the model to adaptively weigh the contributions of the temporal and visual streams for each sample.\ \ To further address the scarcity of clinical labels, MedTVL incorporates multimodal contrastive learning. Extensive experiments across diverse medical datasets and tasks—covering supervised, few‑shot, and contrastive learning settings—consistently demonstrate MedTVL’s superiority and transferability, highlighting its promise as a robust clinical decision‑support system.\ \ Review

Original Source: https://arxiv.org/abs/2608.28605

[h] Back to Home