Your escalation queue is a shaped diet. And diets shape the eater.
I keep hearing the queue described as a neutral pipe. AI handles the easy calls. Humans handle the hard ones. Somebody reviews the reviewer, calls it "human in the loop," files the audit. Nobody in that chain is watching what the queue does to the human over time.
The fraud lead who found her own drift
In July I was on a call with a head of fraud operations at a European payments company. Her team reviews the transactions the model flags. She was proud of their accuracy on card-not-present cases and puzzled by their slide on merchant onboarding. She had assumed the fraud patterns had shifted.
Then she pulled six months of the queue itself. The model's routing had changed. Her reviewers were now seeing roughly nine card-not-present cases for every merchant case. Six months earlier the split had been closer to even.
She sat with that for a moment. Then she said, quietly, "So the model didn't get worse. My people did. And the model is why."
She was right. The training set for the humans had shifted, and nobody was looking at it, because nobody thought of it as a training set.
What the research says
A team at the University of Trento just measured exactly this. Pesenti and colleagues ran 226 people through a binary galaxy classification task. Three conditions: a balanced deferred set, a 90/10 split one way, a 90/10 split the other way. They wanted to know whether the shape of what a reviewer sees changes how accurate the reviewer is.
It does. Symmetrically. When the smooth class was over-represented, accuracy on that class dropped from 0.72 to 0.62. When the non-smooth class was over-represented, accuracy on that class dropped from 0.76 to 0.64. The interaction was χ²(2)=278.95, p<.001. Not a wobble.
They also tested whether reviewers self-correct with more exposure. Across 150 trials, they do not: χ²=0.80, p=.671.
"Participants exposed to a highly imbalanced rejection set achieved lower classification accuracy in the majority class compared to those exposed to a more balanced set, regardless of which class constituted the majority."
The mechanism the authors propose is the test-taker's effect: a mismatch between what the reviewer expects to see and what they actually see. The queue quietly rewrites the reviewer's prior, and their calibration goes with it.
Every "flag for human review" flow in your company has this second-order effect. Your vendor dashboard does not measure it. Your compliance stack does not audit it. Your reviewers cannot feel it, because the paper says they do not self-correct across 150 trials.
The composition of what your humans see is a governance decision now. Somebody in your org needs to own it.