AgBench is a benchmark suite for agentic AI on personal devices, offering reproducible evaluation pipelines. Modern agentic AI systems largely rely on cloud‑hosted large language models for planning, tool use, and iterative execution, which raises API cost and data‑privacy concerns. Advances in on‑device compute enable local execution, yet limited CPU, memory, and storage can affect task success and latency. Existing benchmarks do not systematically capture trade‑offs among device capabilities, workloads, and deployment architectures. AgBench defines a common set of agentic workloads and measures task success, latency, cloud API cost, and data exposure under three deployment modes: local‑only, hybrid, and cloud‑only. Over 162.07 million data points were collected; results show that personal devices can handle many agent tasks locally, but local‑only execution generally yields lower success rates and longer completion times, especially as concurrency grows. Purely local execution eliminates cloud model fees and exposure of sensitive information. Hybrid execution can improve success, but its cloud cost and data exposure depend on how work is partitioned and information is shared. No single architecture dominates across success, throughput, cloud cost, and security, so deployment choices should match the target workload and device capabilities. The AgBench code and data are publicly available at https://anonymous.4open.science/r/AgBench-2777.
Review