Knowing which claims will cost money
the day they arrive
Of everything landing in the queue today, what will end up hurting?
Any incoming queue holds that mix: most of it clears without noise and a few turn sour, and the value is telling them apart the day they arrive, whether they are complaints, insurance claims or credit applications.
today's queue · 8 real cases from 2024
-
comes in they're chasing me for a phone debt that isn't mine, it was my father's
why An ownership dispute in debt collection: in the system's memory, almost none like it ended in compensation. The system gave it a 0.2% probability of ending in compensation: it let it through.
the future The company closed the case with an explanation, no money. let through, rightly
original "This account is not my account and was a XXXX account my father opened over 7years ago. (...) I have been fighting to get it off my report." · case 10000591 in the CFPB database · product: debt collection · came in Sep 3, 2024
-
comes in they charge me for the paper statement, plus a late fee on top this one will cost money · 79%
why Concrete fees with amounts, chained to a failed autopay: the profile that, in the memory, ends in a refund. The system gave it a 79% probability of ending in compensation, inside the year's top 1% of warning: flagged.
the future The company closed the case compensating with money. hit
original "...is charging me for a paper statement fee. (...) charged me a {$30.00} late fee and XXXX % interest on top of that." · case 10102264 in the CFPB database · product: store credit card · came in Sep 12, 2024
-
comes in my mortgage servicer asks again for the insurance proof I already sent
why Recurring insurance paperwork on a mortgage: annoying, but in the memory's similar cases there is no money on the table. The system gave it a 0.2% probability of ending in compensation: it let it through.
the future The company closed the case with an explanation, no money. let through, rightly
original "I am once again submitting another complaint regarding my loan (...) I have submitted said asked for information a number of times." · case 10000124 in the CFPB database · product: mortgage · came in Sep 3, 2024
-
comes in a 20 dollar late fee; I thought the cap was 8 this one will cost money · 63%
why A fee with an amount and an explicit request to lower it: refund profile in the memory. This time the bank did not budge. The system gave it a 63% probability of ending in compensation, inside the year's top 1% of warning: flagged.
the future The company closed the case with an explanation, no money. false alarm
original "I was late with a payment. (...) I asked to have the fee I was charged reduced to {$8.00} and the lady would not do it." · case 10066672 in the CFPB database · product: credit card · came in Sep 10, 2024
-
comes in they promised zero fees if I kept 300 dollars in, and they keep charging me this one will cost money · 79%
why The bank's own promise broken, with dates and amounts: in the memory, this profile ends with the fees returned. The system gave it a 79% probability of ending in compensation, inside the year's top 1% of warning: flagged.
the future The company closed the case compensating with money. hit
original "Despite being guaranteed no further monthly low balance fee, I was still charged {$5.00} a month..." · case 10140278 in the CFPB database · product: savings account · came in Sep 16, 2024
-
comes in my online bank closed my account with my money inside, no reason given
why An account closed with a balance inside, but the text carries none of the classic fee-or-charge signals. The system did not see it. The system gave it a 4.9% probability of ending in compensation: it let it through.
the future The company closed the case compensating with money. it slipped by
original "I learned that the account was closed for violation of terms of agreement. When I asked what terms of agreement was violated they couldn't tell me at all." · case 10014409 in the CFPB database · product: account at an online bank · came in Sep 5, 2024
-
comes in fraudulent charges on my card: two refunded, the other four denied this one will cost money · 58%
why Declared fraud with a partial refund already under way: refund profile in the memory. The bank held its no. The system gave it a 58% probability of ending in compensation, inside the year's top 1% of warning: flagged.
the future The company closed the case with an explanation, no money. false alarm
original "A claim was established for a total of {$2300.00}. (...) they have since denied ( twice now ) to refund the remaining four items totalling {$1300.00}." · case 10184925 in the CFPB database · product: checking account · came in Sep 20, 2024
-
comes in 400 dollars in overdraft fees, with the protection already in place this one will cost money · 65%
why Repeated fees with a total amount, and a protection that does not protect: a clear refund profile in the memory. The system gave it a 65% probability of ending in compensation, inside the year's top 1% of warning: flagged.
the future The company closed the case compensating with money. hit
original "I contacted the bank and signed up for overdraft protection and have still received fees (...) Each fee is {$29.00} so I have paid {$400.00} in overdraft fees." · case 10171625 in the CFPB database · product: checking account · came in Sep 19, 2024
The number, whole 2019-2024 backtest: the system flagged 5,281 complaints, the 1% with the most warning each year. 2,755 ended in real compensation: 52 out of every 100 flagged, against 9 out of every 100 in the whole queue. The other half were false alarms. And of what it let through, nearly 9 in 100 ended up costing money, like the case that slips by above. The warning concentrates attention; it does not guess the future.
summary
- the data
- The public complaints database of the U.S. consumer finance regulator (CFPB): a 9 GB CSV with 914,168 consumer complaints, each with its text and its outcome.
- the question
- Some complaints ended with the company returning money. Could you warn which ones, the same day they came in?
- the test
- The system reads each complaint using only what was known that day, plus the memory of the earlier ones. Then we open the real outcome and compare.
- what came out
- In the 1% the system flagged with the most warning there were nearly 6 times more real compensations than the average. The easy trick (searching for "refund" and knowing which company tends to pay) stalls at 4.8. Reading the text adds.
It is the question of anyone who has worked in customer service: of everything that comes in today, what will end up hurting? To answer it we used the public database of US financial complaints (CFPB): nearly a million complaints with text, each one with its outcome on record. Most close with an explanation; some, with the company compensating.
The usual rules. The system sees each complaint the way an agent would on the day it came in: the text and the earlier history, nothing more. We trained only on the past of each moment, year by year, and froze every parameter before looking at a single result, with a cryptographic fingerprint (d3d2481).
And a trap we forced ourselves to switch off: knowing which company receives the complaint already predicts a lot (each one has its own policy on compensating). So we measured the system twice, with and without that fact. For the "without" version, on top of that, we masked the company names inside the text itself, since the complaints name the bank constantly.
What came out
Over 528,157 complaints from the future (2019-2024): the 1% the system flagged with the most warning concentrated 5.7 times more real compensations than the average. The obvious detector (keywords plus the table of which company tends to pay) stalls at 4.8, with the confidence intervals separated. Without even knowing the company, the system scores 5.3.
the whole ladder
And where it really matters: within a single company, where the table no longer helps at all, the system kept telling apart which complaints would end up costing money. Every year of the test, without exception.
How to read this without fooling yourself
- The circularity is stated head-on. The company decides to compensate by reading that same complaint. It is not a flaw in the experiment: it is the phenomenon. We model a human decision from its evidence, as an expert triage agent would.
- The bar was the obvious detector. If the system did not beat the "grep for refunds + table of companies", this demo would say so. It beats it, with separated intervals, and we publish the whole ladder.
- Calibrated probabilities: when the system says 20%, 20% compensate. The reliability curve is published.
- Coverage counted: complaints with text too short to be read (below the pre-registered minimum) are left out and counted, year by year.
The caveats, hiding none
- The universe is "complaints with consented narrative" (those who agree to publish their text): it does not represent all complaints, and whoever writes tends to bring a stronger case.
- What is predicted is "the company decided to compensate", not the merits of the consumer's case nor a ruling.
- Today's archive is cleaner than the real historical record (withdrawn complaints disappear); a caveat common to all these demos.
- The test ends in 2024: in 2026 the agency changed regime and the flow of narratives collapsed; the cutoff was declared in advance.
- The eight cases in the queue above are real and verifiable, but chosen by us to show the four possible outcomes: hit, false alarm, miss and correct pass. The real proportions are the ones in the strip under the queue, not the sample's.
References
Predicting outcomes over this database has been done before: it was a data-science classic with another variable, now withdrawn. What this demo adds is the protocol: the future covered year by year, pre-registered parameters, the company confounder switched off and measured, and the ladder of comparisons published in full.
Data and acknowledgements
None of this would exist without the CFPB (the US Consumer Financial Protection Bureau), which publishes its complaints database in the open. Thank you.
All cases are historical and closed. No company is singled out: the rankings serve to measure the system, not to judge anyone today.
Can you use this? Yes. Our analysis, the backtest results and the texts on this page are published under the CC BY 4.0 licence: use them, share them or analyse them freely, citing the source, Team Banzai (team-banzai.com). Third-party data (the CFPB complaints database, US public domain) keeps the terms of its original sources.
Built with C# and SQLite. Text vectorized by hashing, logistic regression with isotonic calibration, all implemented by hand and deterministic. No language model decides anything.
About this demo
Retrospective study for illustrative purposes, on public CFPB data. The complaints are consumer testimonies; their presence does not prove bad practice by any company. This analysis does not evaluate companies or the agency, does not extrapolate to current cases, and does not constitute financial or legal advice. No affiliation with the CFPB or with any entity mentioned in the data.
What about your company?
The same thing works for deciding which supplier invoices are worth checking by hand and which to wave through, for ranking which insurance claims look expensive the moment they arrive, or for flagging which orders look likely to end in a return before they ship.
If one of these experiments reminds you of a problem of yours: cases that pile up with no way to prioritise, the serious complaint buried among a hundred minor ones, a yardstick that shifts depending on who's looking, write to us and we'll talk it through. No fluff: we'll tell you whether it can be done.