In AI research the term "agent" lacks a single definition, making evaluation, comparison and reproducibility difficult. This survey is organized around five agentic dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior and temporal coherence. For each dimension we review how prior work has conceptualized the capability and we synthesize the metrics, benchmarks and evaluation frameworks that have been used. Environmental interaction covers perception, action space and real‑time feedback from external systems; learning and adaptation includes online learning, transfer learning and rapid response to changes; autonomy measures the degree of self‑driven decision making versus reliance on external commands; goal‑directed behavior looks at task completion, efficiency and multi‑goal trade‑offs; temporal coherence evaluates the consistency of action sequences and long‑term planning ability. The review highlights mature practices and gaps, notably the lack of a standard benchmark for temporal coherence. To address this, we introduce the Agent Compendium, a public digital repository that organizes and extends the identified evaluation methods, offering a unified query interface and downloadable experiment configurations. The compendium provides a common structure for assessing and comparing agent capabilities, supporting more reproducible research, clearer communication and systematic study of artificial agents.
Review