A 10,000 query UK study tested ChatGPT, Claude, Copilot, Grok and Gemini on tax, pension and loan questions. The data also tells you where the chatbot answer is safe to trust.
AI chatbots gave wrong answers to financial questions 57% of the time on average and 88% of the time on harder multi-step questions, according to a 10,000-query study by UK technology firm Saturn, surfaced this week by the Financial Times. Saturn, which sells AI testing, discloses its methodology in the report but the work has not been peer-reviewed. Some models were wrong on 99% of the hardest prompts.
The study tested 18 models from ChatGPT, Claude, Copilot, Grok and Gemini against more than 100 distinct money-related questions, repeated up to five times each. The named failure modes were calculation errors, omission of upcoming tax changes, and hallucinated rules. The single worst case flagged: a free Claude model, Anthropic's Haiku 4.5, advising a pension move that could expose a saver to a £17,500 UK tax charge (roughly $22,000 at recent exchange rates) from HM Revenue and Customs, the UK tax authority. Another Claude model "invented a rule" that graduates could stop repaying student loans by moving abroad.
The 57% figure is the headline number, but the deeper finding is a confidence asymmetry. The chatbots answer tax, pension and loan questions in the same fluent, declarative voice they use for everything else, while their actual accuracy on those rule-based topics collapses. A user cannot read the answer and know whether to trust it. On questions where the wrong answer can trigger a payment or a filing, that is a worse number than 57% sounds.
Paid models beat free models. Newer models beat older ones. Anthropic's top-tier Claude Opus 5 in reasoning mode was the best performer in the study, and it was still wrong 39% of the time. Saturn's CEO Amal Jolly put the consumer-protection frame bluntly: "Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money." The pattern held across vendors: the model brand mattered less than the tier and the recency.
The UK Financial Conduct Authority published a landmark review of AI in retail financial services this year, and the Mills Review has issued recommendations on how the industry should govern chatbot advice. The regulatory direction is the same one the data points to: away from the assumption that the model output is reliable, and toward an audit and disclosure regime that treats the chatbot answer as advice rather than a search result.
AJ Bell, a UK investment platform, has logged a rise in calls from clients about "baffling" AI-suggested actions, according to the firm's head of personal finance Sarah Coles. A separate PensionBee study flagged the same trend from a different angle: the volume of consumer finance questions now being routed through chatbots, and the cost of cleaning up the mistakes that follow. A Hacker News discussion of the Saturn study landed in the same place, with users pointing out that the training lag on current tax policy and loan rules is the structural reason a fluent answer is not the same as a correct one.
The line the data supports is short. Chatbots are usable for budgeting, for comparing two pension products, for explaining what a term like "ISA allowance" means. They are not safe to ask about anything that triggers a payment, a tax filing or a loan term. The right default is to treat the chatbot answer as a research note and the human adviser, or HMRC's own guidance, as the decision.