Most enterprises pick their AI vendor from a leaderboard. Two vendors post scores three points apart, procurement signs the top one, rollout goes fine. Six months in, the humans on the team have drifted to the other one.
I have watched this three times in the last year. Nobody writes a memo. They just quietly re-authenticate with the losing vendor and use it for the work that matters.
The bake-off that decided nothing
A head of engineering at a B2B SaaS firm in Rotterdam, in June, ran what she called a proper bake-off. Two coding assistants, three weeks each, matched task set, an external leaderboard scored on the side. Vendor A won by a small margin. Procurement signed. Rollout was clean.
By August her engineers had drifted. In ones and twos, on the tasks that mattered. When she asked why, the answer was some version of "the other one just gets what I mean." She could not put a number on it. Neither could they. So she trusted the benchmark, which was the only number on the deck.
Her evaluation had measured the wrong thing, and she had no signal to catch it.
What the research measured
Fukushima, Xiong and Moradi Pari at Honda Research Institute Japan and USA replayed about 101,000 play actions from three public Hanabi datasets: human-human, AI-AI, and human-AI. Hanabi is a cooperative card game where each hint carries a literal message and often an implicit one on top. Within the human-AI dataset they compared three different AI partners against a similar pool of human players:
Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38–41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures.
Two partners handing humans the same literal signal. One produced a 14.4% human failure rate. The other produced 34.4%. The gap between what the AI said and what humans understood was the property that mattered, and every AI-vs-AI benchmark is blind to it.
The domain is a card game, so the specific percentages do not transfer to enterprise AI. The mechanism does. Your benchmark stack measures how the model performs alone or against another model. Your engineers measure whether it matches how they already communicate.
Put candidates in the loop with your actual humans on a bounded task, and count where those humans get stuck. The benchmark number stays on the deck. The human-failure number decides.