Half the AI advisor roadmaps I see this year justify their scope with the same sentence. Users will interrogate the model. Users will keep asking until they understand.
I stopped believing that a while ago. Now there is a paper that measures it.
The chat that never happened
A Head of Product at a B2B fintech in Amsterdam, last month, on a coaching call. She pulled up the Figma. Every AI recommendation had a chat input under it. Follow-up questions, threaded state, the whole surface. The engineering estimate was nine weeks.
I asked what fraction of users would actually use it. She said most. I asked how she knew. She said the persona interviews were clear. Users had told the researcher they would want to iterate on the model's answer.
I asked whether anyone had watched a user actually iterate on anything in the current product. There was a pause. Then I sent her the paper.
What the researchers actually found
Althaus, Houf and Schwieren ran an incentivized lab experiment at Heidelberg and Karlsruhe. N=158. Three arms: a static calculator, a one-shot AI answer, and an interactive AI the participant could re-query. Lottery choices were identical across conditions. Only the interactivity of the aid varied.
Interactivity did not move risk aversion. β̂₁ = 0.032 with a 95% CI of [-0.239, 0.302]. This is not a power problem. Their equivalence tests rule out effects larger than Δr≈0.275 at 90% confidence, and the sensitivity control (safe-option position) produced a significant -0.308, p<0.01. The design catches real behavioural effects when they exist. This one did not exist.
Then the second finding, which is the one my client stared at:
the interactive follow-up feature was almost unused, with only 2 of 49 participants ever asking a follow-up.
Two of forty-nine. In an incentivized lab where subjects were paid to think.
A different paper a week earlier looked at disclosed AI persuasion in dominated financial choices, and found the opposite: users bent toward the AI even when warned. The claim here is narrower and, for a roadmap, sharper. Adding chat did not shift the decision, and almost nobody used the chat.
If your product's business case rests on the dialogue actually happening, measure whether people ask the follow-up. Track that separately from clicks. Ship the surface if you like, but do not fund a nine-week sprint on a behaviour you have not seen in your own users.
The interactivity you paid for is a demo asset until proven otherwise.