AI chatbots give wrong financial answers most of the time, study finds

The report follows a warning from Macabacus, a Microsoft 365 productivity platform for finance and professional services teams based in New York, which highlighted that most finance firms have sent clients an AI-generated error.

Hard questions exposed the sharpest failures

The gap between AI performance and consumer expectation widened significantly on more complex queries. On hard questions, defined as multi-part scenarios involving interacting tax rules and precise figures, accuracy fell to just 12%, with mistakes made 88% of the time, per the Saturn report. Even on the easiest questions, which covered general financial literacy topics requiring no calculation, accuracy averaged only 54%.

Free models performed notably worse than their paid-for counterparts. Saturn found that free models failed 63% of the time, compared with 49% for paid-for models. The worst-performing free model, Claude Haiku 4.5, produced wrong or substantially incomplete answers 82% of the time. The best-performing free model, ChatGPT-5.6 Luna (max), still returned incorrect or incomplete responses 56% of the time, according to the Saturn study.

The best overall performer was Claude Opus 5 (reasoning), a paid-for model from Anthropic, which achieved a pass rate of 61% although that means it still failed to answer correctly in almost four out of every 10 cases. On hard questions, even this top performer made mistakes 67% of the time, the Saturn report found.

The divergence between free and paid-for performance carries equity implications the report’s authors found troubling. Consumers least able to afford professional financial advice, the same group most likely to rely on free AI tools, are being exposed to the least accurate guidance.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *