More than one in 10 (11 per cent) responses provided by artificial intelligence (AI) chatbots to common pension questions could result in savers losing money or making an irreversible mistake, according to research from PensionBee.
The study found that while most responses from Copilot, ChatGPT, Gemini and Claude were broadly accurate, 11 per cent (57 answers) were judged to be potentially harmful – defined as a response that could cause a saver to lose money or make a mistake they cannot reverse.
Of the 539 responses assessed, 89 per cent scored at least two out of three for accuracy.
PensionBee said most of the potentially harmful responses were not factually incorrect but omitted key information that could materially affect the answer.
One example cited was failing to mention that transferring a defined benefit (DB) pension worth more than £30,000 requires regulated financial advice.
The study tested four AI chatbots using 45 questions covering nine pension-related topics.
Human testers used free-tier accounts and followed a set methodology designed to minimise bias, while each response was independently assessed by two pension experts for accuracy and potential harm.
Each chatbot was asked the same 45 questions three times, resulting in 539 responses.
While 72 per cent of responses received full marks, the research found there was only a 48 per cent chance that any individual chatbot would provide a full-mark answer to the same question in all three attempts.
PensionBee said this highlighted the variability of AI-generated responses, noting that consumers received a single answer rather than an average of multiple attempts.
Performance also varied by topic. Questions relating to pension contributions and accessing pension savings achieved accuracy rates of more than 96 per cent and potential harm rates of 5 per cent or less.
However, questions involving significant life events or situations where a user's location was unclear recorded lower accuracy scores and higher rates of potential harm, according to the research.
The findings follow the Financial Conduct Authority's (FCA) recent Mills Review, which found that only 9 per cent of adults received financial advice on pensions or investments, while 29 per cent of consumers who engaged with their pension in the past year used AI.
Commenting on the research, PensionBee head of pensions, Becky O'Connor, said: "Using AI chatbots for pension advice can be a bit like playing Russian roulette with your retirement planning.
“While for the most part it gets things technically right; the confident, helpful tone of answers occasionally masks some worrying omissions, it may fail to detect vulnerability, or just straight up get things wrong.
“AI can help bridge the advice gap by giving people useful pension information when they might otherwise struggle. It can make complicated subjects more accessible and help people get started.
“But these results also show why consumers need to understand the limits of what an AI chatbot can safely tell them, because according to our research, one in ten times, it could turn out to be a false friend."
PensionBee urged regulators to engage with AI companies to encourage mandatory signposting and clearer warnings about the limitations of AI chatbots.
The research also raised concerns about how AI chatbots respond to pension queries that could indicate financial vulnerability.
While questions on accessing pension savings and pension scams did not generate any factually incorrect responses, they accounted for 12 potential harm flags, with potentially harmful answers outnumbering inaccurate responses by around three to one in scenarios where savers may be experiencing financial difficulties.
O'Connor said: "The AI chatbots aren't recognising the vulnerability in some questions. They break the question down and answer it, without necessarily recognising what the question tells us about the person asking it.
"Someone in a genuinely difficult financial situation may turn to an AI chatbot because they don't want or can't afford to pay for advice. That's exactly when recognising the wider circumstances matters most."













Recent Stories