In recent years, decoding perceived speech from non‑invasive brain‑computer interface (BCI) signals has attracted considerable attention. Two core challenges dominate the field: extracting neural representations that preserve rich spatiotemporal details, and achieving generalization across subjects. Existing studies typically address one of these issues in isolation, leaving a gap for a unified solution. To bridge this gap we introduce the Subject‑Invariant Cross‑Modal Decoding (SICMD) framework, which jointly leverages functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). We conduct exhaustive experiments on fusion strategies, fusion locations, encoder architectures, and input configurations. Results show that, on cross‑subject perceived speech decoding, SICMD improves Top‑1, Top‑10 and Rank‑acc by more than 10.6%, 10.1% and 1.7% over baseline methods. Moreover, training time is reduced by 88.8% and 60.5% compared to multi‑subject and intra‑subject settings, respectively. Visualization studies confirm that the model effectively captures complementary multimodal features.
Review: By tightly integrating fMRI and MEG, SICMD exploits their complementary temporal and spatial resolutions, delivering substantial gains in cross‑subject decoding accuracy while dramatically cutting computational cost, thus advancing practical BCI deployments.