Spatial transcriptomics maps gene expression on whole‑slide images while preserving morphological context, offering valuable insights for disease research and therapy design. Conventional spatial profiling, however, remains expensive and time‑consuming. Most image‑based prediction methods focus on positional embeddings or image‑level tweaks, leaving text‑driven enhancements largely unexplored. GATE‑ST addresses this gap by incorporating generated gene descriptions as auxiliary inputs to improve image‑based spatial gene expression forecasts. Gene summaries are fed into a text encoder to obtain embeddings, which are then merged with image embeddings via cross‑attention layers, aligning textual cues with morphological patterns. Benchmarking against random gene embeddings and several image‑text fusion baselines demonstrates that GATE‑ST consistently achieves higher prediction accuracy. The approach promises substantial reductions in the time and cost of spatial transcriptomic analysis in pathology imaging, highlighting the promise of text‑guided gene expression prediction.
Review: GATE‑ST’s cross‑modal attention effectively bridges textual and visual modalities, offering a compelling route to make spatial transcriptomics more accessible.