Back to articles
Blog Apana
Training, Fine-Tuning, RAG: Who Does What? Advisor Guide

Training, Fine-Tuning, RAG: Who Does What? Advisor Guide

“Our in-house AI, trained for finance.” That line shows up in almost every demo of a tool built for advisory firms. It sounds impressive. It’s rarely accurate.

Three words get thrown around as if they meant the same thing, training, fine-tuning, and RAG. They describe three very different operations, at very different price points, requiring very different resources, and above all, offering very different guarantees.

Learning to tell them apart takes five minutes. It changes how you listen to a vendor pitch, and it tells you which of their claims actually hold up.

In brief. Training a model from scratch costs tens of millions of dollars, and almost no one does it. Fine-tuning adjusts an existing model’s tone and reflexes, it doesn’t teach it new facts. RAG connects the model to real sources at the moment it answers, and it’s the only layer that ties an answer back to a document you can check. For a number-driven recommendation, that’s the one that matters.


[[SIMULATEUR_2]]

Table of contents

Who can actually train an AI model?

Almost no one. Globally, a small handful of players have the resources.

Picture a full education, from kindergarten through a doctorate. The model starts from nothing, knowing nothing, and reads billions of pieces of text until it learns how language works. It isn’t learning facts, it’s learning patterns, which words tend to follow which words. That’s training, or pre-training.

The scale involved shows how narrow the field is. According to Stanford’s AI Index Report 2025, training GPT-4 is estimated at roughly $78 to $100 million in compute, and training Gemini Ultra at around $191 million. OpenAI’s Sam Altman confirmed to the Wall Street Journal that GPT-4’s training cost exceeded $100 million. On top of that, thousands of graphics processors running for weeks, and specialized research teams.

The takeaway for an advisory firm is direct. When a vendor says it “trained its AI,” in the vast majority of cases what they actually mean is that they took an existing model and adapted it. That’s not dishonest, and it’s not unusual, it’s standard market practice. But the word promises far more than what was actually done.

What does fine-tuning actually change?

The model’s tone, vocabulary, and reflexes. Not the facts it knows.

Think of professional specialization. Someone has already finished their general education, and now they’re taught the codes of a specific trade, the vocabulary, the way to phrase things, the expected answers in a given situation. That’s fine-tuning, a fine adjustment of an already-existing model on a domain or a style.

The operation is far lighter than training. Where pre-training runs into the tens of millions, a fine-tuning pass often costs a few hundred to a few thousand dollars. That’s within reach of many teams, which is why it’s so common.

The point worth remembering is elsewhere. Fine-tuning makes a model more comfortable with wealth-management language, more consistent in its answer format, more aligned with a firm’s tone. It doesn’t give it access to today’s Fed funds rate, or the current contribution limit, or your client’s actual contract. A fine-tuned model still answers from memory, with the same prediction mechanics we describe in our article on how an LLM works.

In other words, it talks better. It doesn’t know more.

Why does RAG change everything for a number-driven recommendation?

Because the model stops answering from memory and goes to read the source at the moment it answers.

Picture an open-book exam. Until now, the AI was reciting. With RAG, it’s handed a set of documents and asked to build its answer from what it just read. The acronym stands for Retrieval-Augmented Generation, formalized in a 2020 research paper by Lewis and coauthors.

Concretely, the advisor’s question first triggers a search across a chosen repository, regulatory text, contract terms, up-to-date tax tables, the client’s file. The retrieved excerpts are handed to the model, which drafts its answer from them.

Two consequences matter for a firm. The answer can be current, since it depends on the repository rather than the training cutoff date. And it can be verified, since you know which document it came from. Without this layer, a numeric answer is a reconstruction from memory, delivered with the same confidence as a verified fact, a point we detailed in our piece on using ChatGPT to prepare advice.

The simulator below asks the same question to three AIs, a generalist model, a fine-tuned model, and a model connected to a real source. All three answer with equal confidence. Only one can show where its number came from.

What blind spot is left at each layer?

Each layer fixes something and leaves a blind spot of its own.

Training sets a cutoff date. The model has no idea what happened after that date, and it doesn’t know it doesn’t know. A rate that changed this summer stays at last year’s figure, stated without hesitation.

Fine-tuning can reinforce a flaw instead of fixing it. A model adjusted to be helpful and assertive becomes more fluent, and therefore more convincing, including when it’s wrong. Confidence improves faster than accuracy.

RAG shifts the question rather than solving it. The answer is only as good as the repository, how fresh it is and how wide its coverage. An outdated document produces an outdated answer, just a traceable one. And if the search pulls the wrong excerpt, the model will still draft a clean-looking answer from the wrong basis.

Which gives a rule that holds across all three layers. None of them turns the tool into a source of authority. They only move where human review needs to sit.

How should you read a vendor’s claim?

By translating whatever word they use into the actual operation behind it.

A simple framework covers most vendor meetings.

“We trained our own model” raises one question. Trained from scratch, or adapted from an existing model? The second answer is the right one almost every time, and it’s entirely fine, as long as it’s stated plainly.

“Our AI is specialized in wealth management” usually describes a fine-tuning pass, sometimes just a set of instructions given to the model. The follow-up question is simple. Does that change its tone, or does it change its sources?

“Our answers rely on a verified, up-to-date repository” is the most serious claim, because it’s checkable. Three questions confirm or disprove it. Which sources exactly, how often are they refreshed, and does the tool show the document a given number came from?

That last point is the most telling. A tool that shows its source accepts being challenged. A tool that doesn’t is asking you to take its word for it. For a recommendation that carries your professional liability, that isn’t a matter of convenience, it’s a matter of kind.

As an example, that’s the approach we follow at Apana with Advisor, where portfolio scenarios run on deterministic calculation engines and real professional and tax reference data, with the advisor remaining the one who arbitrates and signs off.

FAQ

Does a fine-tuned model know French or US tax rules better? It talks about them better, in the right vocabulary and the right format. It doesn’t thereby have up-to-date thresholds or brackets. Only a connection to an external source, via RAG, gives that guarantee.

Does RAG remove the AI’s errors? No. It sharply reduces invented answers by tying them to documents, but quality still depends on the repository and the search step. An outdated source gives an outdated answer, just a traceable one.

Do I need a domestic model for domestic data? Those are two different questions. Language and legal culture are a matter of the model and its fine-tuning. Data location is a matter of hosting and deployment, covered separately elsewhere in this series.

How do I find out what a vendor actually did? Ask where the number on screen came from. If the tool can point to a dated document, there’s a sourcing layer behind it. If it can’t, the answer is coming from the model’s memory.

Sources

  • Stanford Institute for Human-Centered AI, Artificial Intelligence Index Report 2025, chapter on frontier-model training costs (estimates developed with Epoch AI).

  • Stanford HAI, Artificial Intelligence Index Report 2024, Gemini Ultra training-cost estimate.

  • The Wall Street Journal, Sam Altman’s statements on GPT-4’s training cost.

  • Patrick Lewis, Ethan Perez, Aleksandra Piktus et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems (NeurIPS), 2020.

Want to go further? If you’re new to the topic, our foundational article on LLMs explains the prediction mechanism in a few minutes. To choose a tool on the market, see which AI tool to pick as an advisor.

Want to see it on a real file? We’ll take one of your client files and walk through the analysis live, in under 15 minutes, sources shown at every step.

We use cookies to improve your experience. By continuing, you agree to our cookie policy.