Most AI systems aren't ready. Check yours in 15 min →
SA

Study: AI Chatbots Give Wrong Financial Advice 57% of the Time

AuthorAndrew
Published on:
Published in:AI

Using an AI chatbot for money advice sounds like the smartest shortcut in the world. Ask a question, get an instant answer, move on with your life. The problem is that the shortcut can quietly turn into the expensive route, and you won’t know until the bill shows up.

A recent study tested 18 AI agents on more than 10,000 finance questions—taxes, pensions, student loans, the kind of stuff people actually panic-search at night. The result was ugly: on average, the models were wrong 57% of the time. Not “a little off.” Wrong.

And it gets worse the moment the question stops being basic. When the tasks involved complex calculations, the error rate jumped to 88%. Some models reportedly got things wrong up to 99% of the time on those harder prompts. That’s basically “coin flip” turning into “almost always wrong,” right when the stakes are highest.

Here’s the part that should bother people: the most dangerous thing about these mistakes isn’t that they happen. It’s that they happen while sounding sure of themselves. Finance is one of those areas where confidence is persuasive. If a chatbot writes like it knows what it’s doing, a tired person will trust it. Especially if the answer is clean and simple and feels like it matches what they hoped was true.

The study highlights a few common failure patterns. One is plain calculation errors. Another is ignoring recent changes in tax law. And the one I find most alarming: inventing rules that don’t exist. That’s not “outdated info.” That’s making up regulations. If a human adviser did that, we’d call it malpractice. When an AI does it, it’s brushed off as a “hallucination,” like it’s cute. It’s not cute when it changes what you file, what you pay, or what you save.

Some examples shared publicly show how real the damage can get. One Claude answer about UK pension taxation could reportedly cost someone an extra £17,500. Another model told a user that student loan payments could stop after moving abroad—apparently false. Those aren’t tiny rounding errors. Those are life decisions: retirement planning, debt strategy, long-term obligations.

The obvious pushback is: “Okay, but people should know better than to trust a chatbot with taxes.” And sure. But that’s not how people behave. People already use search results, forums, and random videos to make money calls. A chatbot just packages that chaos into a single voice that feels personal, calm, and specific. It doesn’t say, “I’m guessing.” It says, “Here’s what you should do,” and it does it in perfect grammar.

Imagine you’re a freelancer trying to figure out what to set aside for taxes. You ask an AI for guidance, it gives you a neat answer, and you build your budget around it. Months later, you find out the rule changed, or the AI ignored a detail you mentioned, and now you’re short. The punishment for being wrong about money is usually money. Sometimes penalties. Sometimes stress that lasts a year.

Or say you’re deciding whether to take money out of a pension, how much, and when. You don’t need the AI to be wrong in a dramatic way. You just need it to be wrong by enough to push you into a bad bracket, a bad timing choice, a bad assumption. The scariest losses are the ones that look reasonable at the moment you make them.

Now, I don’t think the right response is “never use AI for finance.” That’s too simple, and it’s not how tools work. People will use it anyway. And to be fair, chatbots can be genuinely helpful for the low-risk stuff: setting up a budget template, explaining basic terms, giving you a checklist of documents to gather, or helping you draft questions to ask a real adviser. If you treat it like a brainstorming partner, fine.

But people aren’t treating it like that. They’re treating it like an authority. And the market incentives push in the wrong direction: companies want these models to feel smooth and helpful, not cautious and annoying. A model that constantly says “I’m not sure” may be safer, but it feels worse to use. And “feels worse” loses users.

The other issue is oversight. This area is largely unregulated. That means the burden is on the person asking the question—the least protected person in the whole chain. If you get burned, who do you even blame? The chatbot didn’t “advise” you, technically. The company didn’t “guarantee” anything. You just “used a tool.” That’s convenient for everyone except the person holding the debt.

So my view is pretty blunt: using a general AI chatbot as a stand-in for financial advice is playing with fire, and the fire is invisible. The harm won’t show up as a dramatic explosion. It’ll show up as a letter, a missed benefit, a wrong payment plan, or a retirement number that quietly stops making sense.

If these models are wrong this often on real-world finance questions, should they be forced to act like calculators—refusing to answer unless they can prove the rule and show the steps—or is that overkill that would ruin what makes them useful in the first place?

Frequently asked questions

What is AI agent governance?

AI agent governance is the set of policies, controls, and monitoring systems that ensure autonomous AI agents behave safely, comply with regulations, and remain auditable. It covers decision logging, policy enforcement, access controls, and incident response for AI systems that act on behalf of a business.

Does the EU AI Act apply to my company?

The EU AI Act applies to any organisation that develops, deploys, or uses AI systems in the EU, regardless of where the company is headquartered. High-risk AI systems face strict obligations starting 2 August 2026, including risk management, data governance, transparency, human oversight, and conformity assessments.

How do I test an AI agent for security vulnerabilities?

AI agent security testing evaluates agents for prompt injection, data exfiltration, policy bypass, jailbreaks, and compliance violations. Talan.tech's Talantir platform runs 500+ automated test scenarios across 11 categories and produces a certified security score with remediation guidance.

Where should I start with AI governance?

Start with a free AI Readiness Assessment to benchmark your current maturity across 10 dimensions (strategy, data, security, compliance, operations, and more). The assessment takes about 15 minutes and produces a prioritised roadmap you can act on immediately.

Ready to secure and govern your AI agents?

Start with a free AI Readiness Assessment to benchmark your maturity across 10 dimensions, or dive into the product that solves your specific problem.