The solution
Frontier language models now produce deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. The authors argue this is substantially a measurement problem: public benchmarks answer what a model can do, while a deployment decision asks whether this workflow is fit, reliable, safe and worth scaling on the organization's own data and controls.
EnterpriseVal is a use-case-level evaluation system. It starts with a formal specification of the frozen socio-technical configuration under test, whose autonomy level and consequence tier jointly set the required evaluation intensity, plus a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight.
Grading scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference, and a two-tier threshold gate, stated as an executable algorithm, maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions.
A pilot across three workflows in a global bank documented the pattern: in credit-memo drafting, the best model reached 88% human-graded citation precision and a 1.6% hallucination rate against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document.
Why it worked
Autonomy level combined with consequence tier prices the evaluation: riskier configurations face more intensity.
Prediction-powered inference keeps expert judgement authoritative while LLM judging carries the volume.
Confidence bounds on metrics stop single-sample wins from clearing the gate.
Reviewer catch rate is a measured parameter of the value-and-risk model, not an assumption.
What can be applied
Treat GenAI deployment as a measurement problem: freeze the configuration, grade against task-specific gates with confidence bounds, and measure human oversight instead of assuming it.
Aftermath
The paper reports the three-workflow pilot in a global bank and explicitly separates established results and documented pilot evidence from the proposed system, specifying the experiments still required for full validation.
FOLLOW THE EVIDENCE