AI Financial Advice Accuracy: What the Saturn Study Found in 2026

2026-09-23
A new study says AI financial advice is wrong 57% of the time. Here's what that means for AI tax advice, why models hallucinate, and how to check their answers.
AI financial advice has become a default first stop for a lot of people. Ask ChatGPT how to split a paycheck, and you'll get an answer in four seconds, written with total confidence, for free. That confidence is the problem: a new study from the financial technology firm Saturn says those answers are wrong about 57% of the time.
The research, published under the title "Artificial Authority," tested how well popular chatbots handle money questions. Saturn ran 18 AI models against roughly 10,000 financial questions, covering ChatGPT, Claude, Copilot, Grok, and Gemini. On average, the models got it wrong in 57% of cases. This isn't a small failure rate. If a human advisor was wrong that often, nobody would hire them twice.
Here's what the numbers actually say, why AI is so bad at this specific kind of question, and how to use these tools without getting burned.
What the Saturn Study Measured
The headline number only makes sense once you split it up, because the error rate isn't uniform.
Simple questions did better. Complex ones fell apart. Saturn reported the error rate jumping to 88% on complex inquiries, with some individual models failing 99% of complex financial queries. That last figure is worth sitting with for a second. A model that gets 99% of hard money questions wrong isn't unreliable so much as unusable for that task.
The tier split matters just as much. Paid tools averaged a 49% error rate. Free tools averaged 63%. On the hardest questions, free models climbed to 93%. Paying for a chatbot improves the odds, which makes sense once you consider that paid tiers usually get the newer models with better reasoning. It doesn't make the advice safe. A 49% error rate is still a coin flip.
The failures weren't random noise, either. Saturn found two repeating patterns. Models missed current tax legislation updates, the kind of thing that changed last year and never made it into training. And they invented financial rules that don't exist anywhere.
The errors pile up in the place where they do the most damage. ChatGPT financial mistakes and the same failures from other major models cluster around AI tax advice, the category that changes every filing season.
That second failure mode has a name: AI hallucination, where a model generates a confident-sounding statement with no basis in any real source. Ask about a contribution limit and it won't say "I don't know." It'll give you a number. The number might be from 2019.
Why AI Gets Financial Advice Wrong
Training data has a cutoff. Tax law doesn't.
That single mismatch explains most of the damage. A model trained on text up to some point in the past knows the rules as they stood then. Standard deduction amounts move. Contribution limits move. Credits get created, expanded, and killed. Every one of those changes is invisible to a model that stopped learning on an earlier date.
Saturn's report points to a second, quieter problem: fabricated rules. When a model doesn't have the current answer, it often fills the gap with a plausible-sounding one rather than admitting the gap exists. That's not a bug someone forgot to fix, it's how these systems generate text. They predict what a confident answer would look like.
Tax questions sit right in the worst possible spot for that. They're high-stakes, they change constantly, and they vary by jurisdiction. NerdWallet's own testing found models leaning heavily on federal rules and getting thin and unreliable once state specifics entered the picture. Ask about a state-level credit in a state that changed its code two years ago and you're asking for trouble.
Here's a concrete example of how that plays out in practice, because the abstract version undersells the problem. You ask a chatbot whether you can deduct home office expenses as a freelancer. It tells you yes, then hands you a set of conditions that stopped applying two tax years ago. You follow them. The return gets flagged.
Legal questions show the same weakness. Chatbots occasionally cite cases that were never filed, in courts that never heard them. A hallucinated citation is easy to spot if you know how to check, and hard to spot if you don't.
Where AI Advice Is Actually Solid
Before you delete the app, it's worth saying what the study didn't find.
Saturn's numbers focus on accuracy of factual financial claims. A separate round of scenario testing, including tests run by NPR and CBS, found that chatbot guidance on broad planning questions was largely sound. Models consistently pointed people toward emergency funds, encouraged paying down high-interest debt before investing, and recognized the basic math of compound growth. Those are the general principles most personal finance writing has agreed on for decades, and the models reproduce them reliably.
Notice the difference. General principles hold up because they don't change. Specific numbers break because they do.
That gives you a rough line to work with. A model is a decent thinking partner when you're weighing two approaches or trying to understand why an emergency fund matters at all. It's a bad source of truth for a figure, a deadline, or a rule you're about to act on.
| Question type | AI reliability | What to do |
|---|---|---|
| Concept explanations (what is a Roth IRA) | Reasonable | Use it, then confirm on an official page |
| General strategy (debt vs. investing) | Reasonable | Treat as a starting framework |
| Current contribution limits | Weak | Check IRS.gov directly |
| State tax rules | Weak | Check your state revenue department |
| Complex multi-part planning | Poor | Bring it to a professional |
| Specific legal citations | Poor | Verify every case reference |
Which Money Questions to Trust a Chatbot With
The practical split is simpler than it sounds. Sort your question by whether the answer could change without the model knowing.
Concepts are safe territory. "How does a 401(k) match work?" asks for an explanation, not a live fact, and explanations don't expire the way a published dollar figure does once a new tax year starts. Same with definitions, basic arithmetic, and the reasoning behind common strategies. If you're trying to understand the trade-offs between two moves, a chatbot will lay them out competently.
Anything tied to a number or a date needs a second source. That includes contribution limits, income thresholds, filing deadlines, deduction amounts, and eligibility cutoffs. These are exactly the facts the Saturn study caught models getting wrong, and exactly the facts that carry penalties when you get them wrong on a filing.
Complex questions deserve the most skepticism. A question that stacks multiple conditions, like how a side business affects a deduction while also changing your bracket, is where the 88% error rate lives. Models handle one variable. Real financial situations have six.
How to Cross-Check AI Financial Advice
Verification doesn't need to be elaborate. It needs to be fast enough that you'll actually do it.
Ask for the source. Add "cite the official source for that figure" to your prompt. A model that can point you to IRS Publication 590-A is doing something useful. One that can't, or that cites something that doesn't exist, has told you what its answer is worth.
Then go to the source directly. The IRS publishes current limits and rates on its own site, and state tax agencies do the same, which means the official number is always a couple of clicks away rather than buried somewhere you'd need an advisor to reach. If a chatbot's number doesn't match what's on the official page, the official page wins, every time.
Run the same question twice. Models don't give identical answers on separate attempts, and a figure that shifts between two sessions is a figure the model doesn't actually know, which is the kind of signal that a confident paragraph of prose will never give you on its own. This takes thirty seconds and catches a surprising amount.
Check the date in the answer. If it says something like "the 2024 limit is," ask why it's giving you an old number. A model that stops and corrects itself is more trustworthy than one that doubles down.
Think of AI tools the way you'd think about a weather app. You'd still look out the window before deciding to carry an umbrella.
Run this checklist before you act on anything a model tells you about money:
- Ask for the official source behind every figure, then open that source yourself
- Open the current IRS or state tax page and compare the numbers line by line
- Ask the same question in a fresh chat and see whether the answer shifts
- Flag any question that stacks more than two conditions and treat the reply as a draft
- Bring questions about business income, property sales, or inheritance straight to a professional
Limitations and When to See a Professional
Everything above assumes an ordinary financial question. Some situations shouldn't go near a chatbot at all.
If your question involves a business entity, an inheritance, a divorce, a sale of property, or a complicated tax position, the cost of a wrong answer is measured in thousands, not embarrassment. A CPA or a fee-only fiduciary advisor charges for their time and carries liability for being wrong. A chatbot does neither.
Be careful with the paid tiers, too. A 49% error rate on paid tools is better than 63%, and it's still about as reliable as flipping a coin. Upgrading the subscription buys you better reasoning, not accurate knowledge of a rule that changed after the training cutoff.
The study's own numbers make this point bluntly. Free models hit 93% errors on the hardest questions, and the paid ones still miss almost half. Nobody sells a tool that gets it right.
The FCA found that 26% of consumers trusted general-purpose AI tools for financial advice. That's the actual risk. The advice reads exactly like good advice, and there's no signal in the text telling you which parts are wrong.
So, can you trust AI advice at all? Yes, as a tutor. Not as an authority. That distinction is the whole game here, and it's the one that keeps AI reliability from turning into a financial mistake.
One last safeguard: treat financial details the way you'd treat a password. Whatever you type into a chatbot leaves your control, and there's no reason to hand over account numbers, Social Security digits, or full tax documents when you're only asking how a rule works.
Used properly, none of this makes the tools useless. Treat a chatbot as a patient, articulate tutor who explains concepts well and shouldn't be trusted with a number you'll write on a tax form. Ask it to help you understand a decision, then download the current IRS forms and check the specifics yourself, and loop in a professional the moment real money and real deadlines are involved. That workflow keeps the useful part of AI financial advice and drops the part that costs you.
FAQ
Is AI financial advice accurate enough to rely on?
Saturn's study found an average error rate of 57% across 18 models. Accuracy is better on general concepts and worse on numbers, deadlines, and complex situations, so it works better as a starting point than a final answer.
How often do AI chatbots get financial advice wrong?
57% of the time on average, rising to 88% for complex questions. Free tools averaged 63% errors; paid tools averaged 49%.
Can AI help with tax questions at all?
It can explain how a rule works in general terms. Anything involving a current limit, threshold, or state-specific rule should be confirmed on an official tax authority page before you act on it.
Which AI models did the study test?
Saturn tested 18 models, including ChatGPT, Claude, Copilot, Grok, and Gemini, against roughly 10,000 financial questions.
What causes the errors?
Two things show up repeatedly: training data that predates current tax legislation, and hallucination, where a model invents a plausible rule instead of admitting it doesn't know.
Should you ask AI for financial advice?
Yes, with limits. Use it to understand concepts and compare approaches, verify any specific figure at the source, and bring complicated or high-stakes decisions to a licensed professional.