Wang, Sturgis and de Kadt just ran a preregistered paired-profile experiment with 1,919 U.S. job seekers. They varied six features of an AI hiring system to see what applicants actually weight. Human involvement moved stated choice about as much as cutting wrongful rejections from 30% to 10%.
The number I keep rereading
Twenty percentage points of wrongful rejection. Roughly the weight of one staffing decision.
I had to sit with that.
The appeal, the opt-out, the independent bias audit: applicants did not weight them more when the error rate rose. Each stayed inside a preregistered equivalence bound. The authors preregistered the range in which they would accept the null and stayed inside it. That is a much stronger claim than the usual p greater than .05.
The rooms I sit in have been buying the other story for years. Better model, tighter appeal SLA, cleaner audit. Legitimacy follows.
Except the applicant on the other side of the funnel reads a different signal. What they trade for, according to their stated preferences, is a person in the seat.
Where I saw this last month
Head of Talent, mid-size B2B SaaS in London. She was walking me through their new screening stack, which filters roughly 80% of applications before a human ever reads them. The dashboard was beautiful. The appeal channel had a 0.4% activation rate. She read that as things working.
I asked when she last spoke with a rejected candidate.
She had not, in eighteen months.
The room went quiet in that specific way. Then she said the vendor's whole pitch had been the bias audit and the appeal SLA. I told her about this paper. Her face did the thing faces do when a purchase decision starts moving in real time.
She had been buying legitimacy from the machinery around the model. The paper says applicants weight that machinery the same whether the error rate is high or low. The person in the seat is what shifts them.
Two caveats before you take this to your board. The measure is stated preference, not downstream hiring outcome. The population is U.S. job seekers, not your specific pipeline. The mechanism transfers; the coefficient does not.
The direction, though, is unambiguous. It points against a decade of AI-first HR architecture your peers have already bought.
Human involvement carried more weight than any procedural feature, moving stated choice about as much as cutting wrongful rejections from 30% to 10%.
Every governance dashboard you commissioned this year is buying you less than one hire would.