Back to articles
Blog Apana
Why financial advisers shouldn't use ChatGPT to build advice

Why financial advisers shouldn't use ChatGPT to build advice

Contents

More and more advisers open ChatGPT between meetings to rough out an allocation or sketch a tax structure. The reflex makes sense: the tool is fast and answers with confidence. The problem isn't the reflex, it's what the tool gives back.

A general-purpose chatbot isn't built to produce wealth advice. Not because it's incompetent, but because what it produces is neither reproducible, nor neutral, nor traceable. And those are the three qualities that separate advice from an opinion.

A study from MIT Sloan, Stanford and the NBER, published on 3 August 2026, puts a number on the risk: 952 people wrote their own questions to an AI, and the researchers measured what it recommended.

In short: a chatbot's advice changes with every attempt, depends on how the question is phrased, and can cite a wrong tax rule without flagging it. The equity-share gap between a male-labelled and a female-labelled profile reaches 1.50 points, at identical circumstances. Wording the question better changes nothing. Only a tool that structures the data and computes outside the model produces defensible advice.

Can an adviser draft an allocation with ChatGPT?

To brush up on a principle, yes. To build a recommendation they'll put in front of a client, no.

On broad principles, the model is useful and honest: it explains diversification, points to the value of an emergency fund, warns against gambling. As a thinking aid it has its place, like the other general-purpose AI assistants we compared. The limit appears the moment its answer becomes the basis for a decision. Advice carries a signature; a chatbot doesn't. Ask it the same question tomorrow and the answer changes. Rephrase the client's situation and it shifts again.

Why does the answer depend on how you ask the question?

Because the model mirrors the question, not the situation. And two clients in the same position never write it the same way.

The study proves it with a clean experiment. The researchers take questions that reveal nothing about the author's gender, then randomly insert "I am a man" or "I am a woman." Because the label is assigned by coin flip, any gap can only come from the model.

The result: between a male and a female profile, the recommended equity share shifts by 1.50 points. Wording accounts for 0.96, the gender label for 0.54.

That 0.54 point is a bias nothing signals: the model reads gender as a sign of caution and adjusts its answer. A setting, not an inevitability, because the same test on ethnicity produces no gap at all. The guardrails were tuned for race and let gender through.

Does each AI carry its own biases?

Yes. Beyond gender, every model carries leanings inherited from its training and its tuning. Three examples, all measured:

  • GPT-5.2, Gemini 3 Flash and GPT-5.6 Terra all lower the equity share when the question is labelled female, by 0.54 point at identical content.

  • GPT-4 leans towards a young, high-income investor profile, far from a firm's real client base.

  • Across 33 models, Foltyn and Olsson find an average gender gap of 1.8 points.

On top of these sit subtler quirks: over-citing the products most frequent in training, favouring round numbers, anchoring on the first figure mentioned. None shows up when you read an answer, none is fixed by rephrasing.

And the most insidious part is the two thirds of the gap that come from wording alone. A client who says "I know nothing about markets" and one who wants to "optimise my exposure" ask two different questions for the same situation. The bias isn't in the answer, it's in the question.

Compounded over a lifetime, the gap adds up: roughly 4% less wealth by age 60 for profiles described by women or by people less comfortable with finance, close to 6% for those who had never used AI. US figures; only the relative gaps carry over.

Does the risk reach wealth strategy?

Yes, and more so, because that's where a mistake is most costly.

Described in cautious terms, a situation calls up a simple answer. The same estate, framed around succession or tax, opens a richer horizon: a bare-ownership gift, an arbitrage between retirement and life-insurance wrappers. It's then the vocabulary of the case that sets the ambition of the strategy, not the client's real potential.

Add the hallucination risk, with direct consequences: an AI can build a credible structure around an expired tax rule and present it as valid. The error shows neither in tone nor in logic, only on review, often at the notary's, months later. That's what happens when the AI invents a tax rule that doesn't exist.

Is a better prompt enough to stay safe?

No, and it's the study's most useful finding.

The researchers wrote the ideal prompt: full situation laid out, assumptions stated, an instruction to act as an adviser trained in portfolio theory. Some things improve, but the core holds: the model still recommends no rebalancing of an existing portfolio, even when handed the current allocation.

The conclusion takes the pressure off the professional: the problem isn't how you go about it, it's the tool for this use. No writing skill makes reproducible what isn't. It's also what the regulator now expects, with the obligations the EU AI Act imposes from 2 August 2026.

How do you use AI without handing it the advice?

By changing its role. In a chatbot, the AI decides, and that's where it all goes wrong. In a well-designed tool, it's only a conductor: it doesn't compute, it calls a calculation engine; it doesn't invent a tax rule, it opens the statutory text; it doesn't guess a figure, it queries the client file.

Three principles are enough to neutralise the biases above. The data is entered in fields, not in a sentence: two thirds of the wording bias fall away. The calculation comes from a deterministic engine: same data, same result, defensible. The products come from a verified reference universe, ISIN and KIID, not from the model's memory.

This is the approach we follow at Apana, by way of example. Each failing of the chatbot is met by a building block:

These blocks also run without the AI and slot in alongside existing tools, a sign that reliability comes from the architecture, not the model. One point matters: information moves directly from one module to the next, through Apana's own channels, never through the AI. The model chooses which module to call, it never touches the data.

A chatbot gives you an answer, with no idea where it came from. A real advice tool gives you a line of reasoning you can check, defend and sign.

FAQ

Is an adviser allowed to use ChatGPT for their work?

Nothing forbids it for teaching or research. The issue is professional: a recommendation produced by a chatbot is neither reproducible nor traceable. Since August 2026, the EU AI Act also requires disclosing to the client any exchange with an AI.

Is the AI's gender bias proven or assumed?

Proven. By randomly inserting "I am a man" or "I am a woman" into identical questions, researchers isolate a model-specific effect of 0.54 point of equity share, one third of a 1.50-point total gap.

Does a very detailed prompt solve the problem?

Not on the essentials. A careful prompt reduces rules of thumb, but doesn't create rebalancing and doesn't erase the wording bias.

Do these US results apply outside the US?

The mechanisms, yes; the figures, no. A model that mirrors phrasing does so in every language. But the wealth gaps quoted rest on a US tax framework.

Sources

  • Taha Choukhmane (MIT Sloan, NBER), Tim de Silva (Stanford GSB), Weidong Lin, Matthew Akuzawa (MIT Sloan), AI Financial Advice: Supply, Demand, and Life Cycle Implications, working paper arXiv:2608.01607v1, 3 August 2026.

  • Richard Foltyn, Jonna Olsson, The worth of a "wo": Gender bias in financial advice from LLMs, CEPR Discussion Paper 21323, 2026.

  • Anastassia Fedyk, Ali Kakhbod, Peiyao Li, Ulrike Malmendier, AI and perception biases in investments: An experimental study, Journal of Financial Economics, accepted, 2026.

  • Matthias Rumpf, Michael Haliassos, Tetyana Kosyakova, Thomas Otter, From humans to algorithms: How financial advice differs across professionals, peers, and LLMs, SSRN working paper, 2026.

  • Laurent Calvet, John Y. Campbell, Paolo Sodini, Fight or flight? Portfolio rebalancing by individual investors, The Quarterly Journal of Economics, vol. 124, 2009.

We use cookies to improve your experience. By continuing, you agree to our cookie policy.