A foundation model is a large AI model that is pretrained on vast amounts of data and built to be adapted to many different tasks. GPT, Gemini, Claude and Llama are all foundation models. The name comes from the fact that one such base model serves as the foundation for everything from chatbots to coding assistants and AI search — instead of building a new model from scratch for each task.
Pretraining and fine-tuning
A foundation model is created in two phases. Pretraining is the large, general training: the model reads through enormous volumes of text and learns language, patterns and facts. This is where most of the capabilities emerge, and it is costly and time-consuming. Fine-tuning is a smaller, targeted phase afterwards, where the model is specialised — to a particular tone, domain or task — using far less data. Pretraining builds the foundation; fine-tuning shapes it to a purpose.
Parametric memory versus retrieval
The knowledge a model learns during training is stored in its weights. This is called parametric memory — what the model "remembers" without looking anything up. It is powerful, but has two important limits: it has a knowledge cutoff in time (anything after training does not exist to the model), and it cannot be updated without retraining. That is why parametric memory is increasingly supplemented with retrieval: Retrieval-Augmented Generation lets the model pull in fresh, specific sources at answer time. A modern AI therefore answers from two places at once: what it learned, and what it just retrieved.
Today's foundation models
The 2026 landscape is dominated by a few players: OpenAI's GPT series, Google's Gemini, Anthropic's Claude and Meta's open Llama models. They update fast, and none is "best" at everything — the choice depends on the task. What they all share is the same two limits: a knowledge cutoff in parametric memory, and a dependence on retrieval for anything fresh and specific. It is this duality that decides how a brand gets described.
Why it matters for AI visibility
For brands, a clear consequence follows: facts about you must be either in the training data or retrievable — ideally both. Parametric memory is shaped by what is written about you online: if your brand is broadly and consistently described across many credible sources, the odds rise that the model "remembers" you correctly with no retrieval at all. If you are rarely or inconsistently described, it remembers you wrong — or not at all. And for anything fresh or specific, your content has to be retrievable in the moment. This is the core of both AEO and GEO.
How CitationLab and you can work with this
CitationLab Monitor measures how the foundation models actually describe and cite your brand across ChatGPT, Gemini, Perplexity and Google AI — that is, what has settled into parametric memory and what gets retrieved. Your job comes down to two things: build a broad, consistent presence so that facts about you become part of how the models describe you, and make your content retrievable for whatever has to be looked up. Want to see what the AI models "know" about your brand today? Start with a review of your AI visibility.
Frequently asked questions
What is a foundation model?
What is the difference between pretraining and fine-tuning?
What is parametric memory?
Why must brand facts be in the training data or retrievable?
Definitions used in this article
- Foundation model
- A foundation model is a large, pretrained AI model built to be adapted to many different tasks, such as GPT, Gemini, Claude and Llama. It serves as a shared foundation instead of being built from scratch for each task.
- Pretraining
- Pretraining is the large, general training phase where a foundation model learns language and knowledge from vast amounts of data. This is where most of the model's capabilities and parametric memory emerge.
- Fine-tuning
- Fine-tuning is a targeted adjustment of a pretrained model to specialise it for a particular task, tone or domain, using far less data than pretraining.
