Skip to content
CitationLab
Back to Glossary
Glossary

What is a foundation model? How large AI models learn and remember

A foundation model is a large, pretrained AI model — like GPT, Gemini, Claude or Llama — trained on vast amounts of text and built to be adapted to many different tasks. It 'remembers' knowledge in its weights (parametric memory), but what it did not learn during training has to be retrieved. Here is what foundation models are, the difference between pretraining and fine-tuning, and why brand facts have to be either in the training data or retrievable.

KR
Krister Ross
Founder & CEO, CitationLab
Published 3 min read
Curious how AI talks about your brand?Run a free visibility check

A foundation model is a large AI model that is pretrained on vast amounts of data and built to be adapted to many different tasks. GPT, Gemini, Claude and Llama are all foundation models. The name comes from the fact that one such base model serves as the foundation for everything from chatbots to coding assistants and AI search — instead of building a new model from scratch for each task.

Pretraining and fine-tuning

A foundation model is created in two phases. Pretraining is the large, general training: the model reads through enormous volumes of text and learns language, patterns and facts. This is where most of the capabilities emerge, and it is costly and time-consuming. Fine-tuning is a smaller, targeted phase afterwards, where the model is specialised — to a particular tone, domain or task — using far less data. Pretraining builds the foundation; fine-tuning shapes it to a purpose.

Parametric memory versus retrieval

The knowledge a model learns during training is stored in its weights. This is called parametric memory — what the model "remembers" without looking anything up. It is powerful, but has two important limits: it has a knowledge cutoff in time (anything after training does not exist to the model), and it cannot be updated without retraining. That is why parametric memory is increasingly supplemented with retrieval: Retrieval-Augmented Generation lets the model pull in fresh, specific sources at answer time. A modern AI therefore answers from two places at once: what it learned, and what it just retrieved.

Today's foundation models

The 2026 landscape is dominated by a few players: OpenAI's GPT series, Google's Gemini, Anthropic's Claude and Meta's open Llama models. They update fast, and none is "best" at everything — the choice depends on the task. What they all share is the same two limits: a knowledge cutoff in parametric memory, and a dependence on retrieval for anything fresh and specific. It is this duality that decides how a brand gets described.

Why it matters for AI visibility

For brands, a clear consequence follows: facts about you must be either in the training data or retrievable — ideally both. Parametric memory is shaped by what is written about you online: if your brand is broadly and consistently described across many credible sources, the odds rise that the model "remembers" you correctly with no retrieval at all. If you are rarely or inconsistently described, it remembers you wrong — or not at all. And for anything fresh or specific, your content has to be retrievable in the moment. This is the core of both AEO and GEO.

How CitationLab and you can work with this

CitationLab Monitor measures how the foundation models actually describe and cite your brand across ChatGPT, Gemini, Perplexity and Google AI — that is, what has settled into parametric memory and what gets retrieved. Your job comes down to two things: build a broad, consistent presence so that facts about you become part of how the models describe you, and make your content retrievable for whatever has to be looked up. Want to see what the AI models "know" about your brand today? Start with a review of your AI visibility.

Frequently asked questions

What is a foundation model?
A foundation model is a large AI model that is pretrained on vast amounts of data and built to be adapted to many different tasks. Examples are GPT, Gemini, Claude and Llama. The name 'foundation' comes from the fact that one such base model serves as the foundation for a wide range of uses, instead of being built from scratch for each task.
What is the difference between pretraining and fine-tuning?
Pretraining is the large, general training where the model learns language and knowledge from vast amounts of data — this is where most of its capabilities emerge. Fine-tuning is a smaller, targeted adjustment afterwards that specialises the model for a particular task, tone or domain. Pretraining builds the foundation; fine-tuning shapes it to a purpose.
What is parametric memory?
Parametric memory is the knowledge the model has stored in its own weights from training — what it 'remembers' without looking anything up. It is powerful but has two limits: it has a knowledge cutoff in time, and it cannot be updated without retraining. That is why it is often supplemented with retrieval (RAG) for fresh or specific facts.
Why must brand facts be in the training data or retrievable?
Because a foundation model can only answer from two sources: what it learned in training (parametric memory), or what it retrieves in the moment (RAG). If your brand is broadly and consistently described online, the odds rise that facts end up in the training data. If it is not, your content at least has to be retrievable when the model answers — otherwise you do not exist to the model.
Key terms

Definitions used in this article

Foundation model
A foundation model is a large, pretrained AI model built to be adapted to many different tasks, such as GPT, Gemini, Claude and Llama. It serves as a shared foundation instead of being built from scratch for each task.
Pretraining
Pretraining is the large, general training phase where a foundation model learns language and knowledge from vast amounts of data. This is where most of the model's capabilities and parametric memory emerge.
Fine-tuning
Fine-tuning is a targeted adjustment of a pretrained model to specialise it for a particular task, tone or domain, using far less data than pretraining.
Share:

Hold deg oppdatert

Få fagartikler, produktnyheter og analyser rett i innboksen.