NeFut Logo NeFut
Admin Login

[CS.AI] EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

Published at: 2026-09-21 22:00 Last updated: 2026-09-22 02:29
#AI #Machine Learning #LLM

EnterpriseVal is a use‑case‑level evaluation system that bridges the gap between public benchmarks that answer “what can the model do?” and the deployment question “on our data, under our controls, is this workflow fit, reliable, safe and worth scaling?”. The system comprises five core components: first, a formal specification of the use case that freezes the socio‑technical configuration—model, prompts, retrieval, tools, guardrails and human oversight—and defines an autonomy level and consequence tier to set evaluation intensity; second, a metric catalogue covering fidelity, utility, efficiency, reliability, assurance and oversight; third, a grading protocol that scales blinded expert judgement with calibrated LLM‑as‑judge scoring via prediction‑powered inference; fourth, a two‑tier threshold gate expressed as an executable algorithm that maps metric vectors with confidence bounds to REJECT, CONDITIONAL or SCALE decisions; fifth, a value‑and‑risk model where the reviewer catch rate is a measured parameter. A pilot across three workflows in a global bank showed that in credit‑memo drafting the best model achieved 88% citation precision and a 1.6% hallucination rate, meeting gates of 70% and 5%; in procedure transformation analyst refinement effort dropped from an estimated 27.4 hours to 2.9 hours per document. The paper separates established results, documented pilot evidence, the proposed system and open hypotheses, and outlines the experiments required for full validation.

Review: EnterpriseVal offers a practical, quantitative framework that turns generative AI promises into repeatable business decisions.

Original Source: https://arxiv.org/abs/2609.21841

[h] Back to Home