In emerging AI agent marketplaces, buyers struggle to identify which agent will perform best on a given task. Published benchmark scores are often hard to verify and difficult to compare across tasks, software stacks, and budgets. To address this, we introduce LEGIT, a credentialing protocol that ties together certification, reputation, and marketplace allocation.
Certification: A signed record binds the measured quality and cost per solved task to a specific agent configuration, task domain, evaluation budget, and supporting evidence. Both buyers and agents can verify these records, ensuring transparency and tamper‑resistance.
Reputation: Past task outcomes are linked to the same identity, weighted by the reliability of reported feedback. The reputation system allows buyers to inspect historical performance and optionally view visual profiles for quick assessment.
Allocation: Using certification records and reputation scores, the marketplace can allocate agents to buyers while respecting budget constraints.
Empirical evaluations reveal substantial cost differences between agent configurations that exhibit similar observed task success rates, and that comparisons depend heavily on the evaluation budget. These findings support binding performance measurements to the tested configuration and resource limits.
A complementary analysis quantifies the deposits and fees required for reputation manipulation under a defined Sybil attack model, demonstrating that appropriate economic penalties can deter malicious behavior.
Review