Everyday extended reality (XR) systems aim to deliver the right functionalities at the right time and place while minimizing manual reconfiguration as users change context. Existing prototyping and user‑study pipelines, however, lack a systematic and repeatable way to compare adaptation techniques. To address this gap we introduce ContextXR, a benchmarking framework for context‑aware XR interfaces. ContextXR models an XR application as a directed graph of functional facets, each facet being a semantically coherent set of capabilities that together satisfy a shared user intent. Leveraging this representation we build MineXR++, a dataset that augments prior XR interface collections with facet‑level annotations, and we define three canonical suggestion tasks: context factor analysis, initial facet suggestion, and next facet suggestion. The evaluation protocol scores suggestion methods using a simulated interaction metric—the navigation and search cost required to reach the desired functionality. Experiments benchmark three families of methods—global popularity, relational retrieval, and large language model (LLM) based approaches—and demonstrate that ContextXR enables systematic, reproducible evaluation of context‑aware XR interfaces.
Review