In this paper, we propose a language representation for multimodal data, where any observation, whether an image, video, or text, is expressed as a set of atomic propositions—simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into an interpretable space that spans fine-grained facts to high-level concepts and composes into richer expressions.
This approach brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality, enabling complex multimodal understanding, rich data curation, and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.
Blogger's Review: The language-centric framework proposed in this paper offers a fresh perspective on multimodal data, enhancing interpretability and cross-modal processing through atomic propositions. It holds significant potential for applications, particularly in fields like autonomous driving. Looking forward to seeing more research outcomes based on this framework.