Spatio‑Temporal Video Grounding (STVG) aims to locate the spatio‑temporal tube in a video that matches a natural language query. Recent approaches achieve strong results in fully supervised, weakly supervised, and zero‑shot settings, but they usually rely on computationally heavy architectures, complex training pipelines, or multimodal large language models. Pocket‑STVG (P‑STVG) introduces a lightweight cascade architecture that tackles STVG by assembling efficient pre‑trained components instead of using a large end‑to‑end model.\ \ P‑STVG consists of three main modules: a temporal‑aware video encoder built on MobileViCLIP, a spatial encoder‑decoder derived from MDETR, and a shared aligned text encoder. Temporal localization is performed either by a lightweight 1D U‑Net or a simple thresholding strategy, allowing the same framework to operate in both weakly supervised and zero‑shot scenarios.\ \ Video representations are pre‑computed independently of the query, yielding an indexing‑friendly pipeline that enables efficient inference on large‑scale video collections. With fewer than 90 M parameters, P‑STVG matches the performance of weakly supervised methods and surpasses earlier zero‑shot approaches while consuming far less memory and compute, establishing a favorable performance‑efficiency trade‑off for STVG.\ \ Review: By leveraging modular design and pre‑computed features, Pocket‑STVG manages to reduce resource demands without sacrificing accuracy, offering a practical solution for real‑world deployment.