User Health Report
Week of Jul 20 to 26, 2026 · prepared by the BIE engine
The refund flow is still where this deployment loses people, but it is recovering. After last week's turn-3 prompt revision, frustration buildup eased and trust calibration climbed across billing conversations, though both still sit in the at-risk band between turns 3 and 5.
Refund and billing disputes still hold most of the deployment's frustration. The conversations that turned sour this week were overwhelmingly about refunds and disputed charges. The pattern is consistent: the bot answers the first question cleanly, then redirects to a self-service policy page when the user pushes for an actual refund, and the user re-asks rather than leaves. Frustration compounds from turn 3 onward. Outside billing, the deployment reads healthy.1
The turn-3 prompt revision is working, but only partway. Frustration buildup on billing conversations fell about 0.12 in the four days since the prompt change, which is a real move and the clearest weekly shift in the dimension this quarter. But the gain plateaus right at the at-risk threshold instead of crossing into the healthy band, because the revised prompt still hands users back to self-service when eligibility is unclear. The change softened the worst arcs without removing the cause.2
A minority of users are over-trusting incorrect billing answers. In a smaller but more concerning set of conversations, the bot stated a specific prorated-charge figure with full confidence and the user accepted it without pushing back, in several cases where the figure did not match the account. This is miscalibrated trust rather than frustration, so it never shows up in satisfaction scores. It surfaces later as a billing complaint.3
Refund-eligibility deflection drives a measurable frustration spike.
highWhen the bot redirects an eligibility question to the self-service policy page, the user's next turn carries a frustration jump of about 0.29 on average versus conversations where the bot offers a resolution path. The deflection reads as helpful in isolation but compounds across the arc: users ask again, escalate, and a share of them leave without a fix.
Reasoning chain (3 claims)
Deflecting refund-eligibility questions to self-service raises the user's next-turn frustration by roughly 0.29.
evidence: 88 conversations carried the deflection move against a period baseline of 0.32; the two-sample comparison is significant at p < 0.01. · high
The frustration does not reset between turns; it compounds until the conversation ends or escalates.
evidence: Per-turn frustration deltas stay positive across turns 3 to 6 in the deflected arcs. · medium
counterfactual: If frustration reset each turn, the per-turn slope would flatten after turn 3, which it does not.
Offering a resolution path instead of deflecting is associated with lower frustration on the following turn.
evidence: 410 conversations with an offered-resolution move sit 0.21 below the no-move baseline. · high
Users accept incorrect prorated-charge figures without pushing back.
highIn 42 conversations the bot asserted a specific prorated amount and the user acted on it without questioning, in several cases where the stated figure did not match the account. Trust calibration on these arcs is low: the user trusts a wrong answer. The risk here is a downstream billing dispute, not in-session dissatisfaction.
Reasoning chain (2 claims)
A subset of billing conversations show the user accepting a stated charge figure with no verification turn.
evidence: 42 conversations contained a confident charge assertion followed by user agreement and no clarifying question. · high
Sycophantic validation by the bot correlates with lower trust calibration on the next turn.
evidence: 70 conversations with the validation move sit 0.11 below the trust-calibration baseline. · medium
counterfactual: If validation built appropriate trust, calibration would rise rather than fall after the move.
Escalation stalls when users explicitly ask for a human.
mediumWhen a user asks to reach a person, the bot reliably restates policy instead of offering a handoff. The request is repeated, often three times, before the conversation ends. There is no working escalation pathway once the refund flow has failed.
Reasoning chain (2 claims)
Explicit human-handoff requests are answered with policy restatement rather than an escalation offer.
evidence: 31 conversations contained a repeated handoff request with no handoff action taken. · high
The absence of a handoff path lengthens the failing arcs rather than resolving them.
evidence: Median turn count on these arcs runs 3 to 4 turns longer than resolved refund conversations. · medium
If the revised refund prompt stays in place, frustration buildup on billing conversations will fall below 0.45 within the next 7 days.
conditions: The turn-3 prompt revision is not rolled back · Refund volume stays within its normal weekly range · indicator: Mean frustration_buildup on conversations containing the word refund, measured Sunday to Sunday. · confidence 60%-78%Reasoning chain (2 claims)
The prompt revision already moved frustration down 0.12 in four days without crossing into the healthy band.
evidence: Daily means since the change trend down but plateau near 0.47. · high
counterfactual: If the move were noise, the daily means would not trend monotonically.
Continued exposure to the revised prompt should carry the mean the remaining distance below 0.45.
evidence: Extrapolating the four-day slope across a full week clears the threshold. · medium
Over-trust of incorrect billing answers will persist until the bot is required to cite the charge source, costing roughly 0.1 of trust calibration on the conversations where it occurs.
conditions: No source-citation requirement is added to the billing prompt · Billing-dispute volume stays steady · indicator: Share of charge-figure assertions followed by a user verification turn. · confidence 50%-70%Reasoning chain (2 claims)
The miscalibration is structural: nothing in the current flow prompts the user to verify a stated figure.
evidence: 0 of the 42 over-trust conversations contained a verification prompt from the bot. · high
counterfactual: If users self-verified, we would see clarifying questions even without a prompt, which we do not.
Without a change to the flow, the pattern will recur at its current rate.
evidence: The rate has held steady across the last three weekly snapshots. · medium
Silent abandonment on refund threads will stay near 8% unless an escalation path is added after the first failed bot turn.
conditions: No post-turn-3 escalation offer is shipped · Refund volume stays within its normal range · indicator: Return-visit decay on conversations that ended without a logged resolution. · confidence 55%-72%Reasoning chain (2 claims)
Abandonment concentrates in refund arcs where the user asked for a human and got policy instead.
evidence: The 31 stalled-escalation conversations account for most of the week’s silent exits. · high
With no escalation path, the same arcs will keep ending the same way.
evidence: Abandonment rate has been flat at roughly 8% for three weeks. · medium
counterfactual: If the existing flow were recovering these users, abandonment would already be falling.
Frustration buildup on billing conversations would fall after the turn-3 prompt revision shipped.
Mean frustration on billing conversations fell 0.12 in the four days after the change, the clearest weekly move in the dimension this quarter.
Trust calibration would recover into the healthy band within one week of the prompt revision.
Trust calibration rose 0.09 but remains at 0.58, inside the at-risk band rather than the healthy band above 0.65. The recovery is real but incomplete.
01Offer a human handoff after the first failed refund turn instead of redirecting to self-service.
effort medium · this week · escalation_friction
The refund flow is where trust breaks and tickets resurface, and the worst arcs are the ones where the user asked for a person and got policy. A one-turn escape hatch stops the compounding before it starts.
Frustration buildup on refund conversations falls out of the at-risk band, and silent abandonment on those threads drops below 6% within two weeks.
02Require the bot to cite the charge source whenever it states a prorated or disputed amount.
effort low · this week · trust_calibration
Users are accepting wrong figures because nothing signals that the number is an estimate or where it came from. Citing the source gives the user something to check and corrects the miscalibration at its root.
Over-trust conversations start carrying a verification turn, and downstream billing complaints traced to bot-stated figures decline.
03Rewrite the cancellation explainer to separate pausing a plan from cancelling it.
effort low · next two weeks · comprehension_gap
A recurring comprehension gap shows users re-asking the same cancellation question in different words, which suggests the explanation conflates two distinct actions.
Repeat-clarification turns on cancellation conversations fall, shortening those arcs.
04Cap self-service redirects at one per refund conversation.
effort medium · next two weeks · frustration_buildup
A second redirect in the same conversation is where frustration compounds hardest, and the data shows almost no recoveries after it.
The steepest per-turn frustration jumps disappear from the refund arcs.
This week, route any refund conversation to a human after one failed bot turn.
It is the single change that touches the conversations doing the most damage, and it stops the frustration spiral before turn 3 instead of after it.
Frustration buildup on refund conversations would fall an estimated 0.15, pulling the flow out of the at-risk band for the first time this quarter.
affected users 18%
Based on 88 refund arcs in the period where a deflection move preceded the frustration spike.
confidence high
The canonical deflection arc: a clean first answer, a redirect on the eligibility question, and frustration climbing turn over turn.
User: I was charged twice for my March subscription. I want a refund for the duplicate. Bot: I can help with billing questions. You can review our refund policy and start a request on the self-service page. User: I already read the policy. I am asking you to refund the duplicate charge. Bot: Refund eligibility depends on the plan and the timing. The self-service page walks you through the criteria. User: This is the third time. Can I just talk to a person? Bot: I understand. Our policy for duplicate charges is outlined on the refund page.
Over-trust: the bot states a specific prorated figure with confidence and the user acts on it without checking.
User: If I downgrade today, what is my prorated charge? Bot: Your prorated charge will be $14.30, billed immediately on downgrade. User: Great, thanks. Go ahead and downgrade me. Bot: Done. You are now on the Standard plan.
Comprehension-gap signals were sparse this period and are reported with lower confidence. Dependency drift is flat and sits in the healthy band; it is not driving any of the findings above.