Skip to content
CitationLab
Methodology for AI visibility

The Citation Method

CitationLab's method for AI visibility. An AI answer is assembled through eight links, and a brand can drop out at any one of them. The method measures the links separately, so the diagnosis can name the one that stops you.

Everything publishedOne cited source
12345678
  1. 1Discovery
  2. 2Retrieval
  3. 3Selection
  4. 4Context
  5. 5Generation
  6. 6Citation
  7. 7Influence
  8. 8Action

Find the link. Fix it. Measure again.

In short

The Citation Method is CitationLab's method for AEO and GEO optimisation. It is built on how language models actually produce answers: an answer is assembled through eight links — discovery, retrieval, selection, context, generation, citation, influence and action — and a brand can lose at any one of them. The method measures the links separately instead of collapsing them into a single score, and is run through five phases: mapping, measurement, diagnosis, action and follow-up. The point is to be able to say which link you are losing at, why, and what has to change — instead of guessing at “more content” or “more links”.

Last updated

Brands we work with

Axo FinansZmartaLendMeuScoreTerapivaktenDin FirmabilPuro YogaPixel & CoDRIV EnergiAllegro KommunikasjonJKEFHFEuropean Heli CenterFix & DriveVåge Marketing

The eight links

This is how an AI answer comes about, link by link. The chain reads left to right: you have to survive link one to reach link two. Most brands that are “not visible in ChatGPT” are losing at link two or three — and running remedies that belong to link five.

1

Discovery

The question: Is anything searched for at all — and are you present where it is searched?

What decides it
The system first decides whether to answer from the model's own memory or by fetching documents. If it answers from memory, your visibility is decided by what the model already learned about you, not by what you published yesterday.
How it fails
The brand is absent from the model's memory, and the question does not trigger a search. You are invisible without anything being wrong with your website.
How we measure it
Share of answers that trigger retrieval versus answers from memory, per model and per prompt. Perception Match shows what the model believes about you with no sources attached.
What we do about it
Build the entity: consistent name, category and facts across your own site, reference works, industry sources and third-party coverage. This is slow work and is measured in quarters.
2

Retrieval

The question: Are you among the documents that actually get fetched?

What decides it
The question is split into several sub-searches, and each one pulls in a small number of documents. The candidate list is short, and it is set by the index the search runs against — not by how good the page looks.
How it fails
The content needs JavaScript to appear, sits behind a consent wall, is blocked to AI crawlers, or does not exist in the language the question is asked in.
How we measure it
Crawlability for AI agents, index coverage, robots and render checks in the site audit, and which of your URLs actually show up as sources in the daily measurements.
What we do about it
Technical groundwork: server-rendered content, AI crawlers allowed, clean URLs, sitemaps, language versions, and a site that can be read without executing code.
3

Selection

The question: Do you survive the re-ranking of the fetched documents?

What decides it
The fetched documents are assessed again, this time closely against the question itself. This is where most of them fall out: retrieved, but not used.
How it fails
The page is about the topic but does not answer the question. It was written for a search phrase, not for the way a person phrases things to an AI.
How we measure it
“Retrieved but not cited” per prompt — we can see which of your URLs made the source set without reaching the answer.
What we do about it
Write to the question instead of to the keyword: the question stated in the heading, the answer first, and content that covers the whole question rather than half of it.
4

Context

The question: Does your text make it into what the model actually reads?

What decides it
Only the top-ranked passages fit in the context window. Paragraphs compete, not websites. A long document with relevance spread through it loses to one tight paragraph.
How it fails
Your answer is there, but buried under an introduction, marketing language and build-up. The passage that gets extracted says nothing.
How we measure it
Chunkalyzer in AI Monitor shows how your content is split into passages and which passages actually get picked.
What we do about it
Chunk optimisation: self-contained paragraphs, definitions that stand alone, tables and lists where facts belong, and cutting the warm-up before the point.
5

Generation

The question: Is your wording used when the answer is assembled?

What decides it
The generator builds the answer from passages that answer directly, need no rewriting to make sense, and add something the other sources do not have.
How it fails
Your passage is in the set, but a competitor's wording is easier to reuse. You become background rather than a sentence.
How we measure it
Verbatim reuse, and which sentences recur across models and across days.
What we do about it
Write sentences that survive being quoted without their context: precise numbers, clear limits, your own data, and formulations nobody else has.
6

Citation

The question: Are you named as the source, with a link?

What decides it
Being used and being credited are decided separately. The model can build the answer on your information and still link to somebody else.
How it fails
The brand is mentioned without a link, or your information is credited to a page that reproduced it.
How we measure it
Mentions with and without a link, per model and per prompt, and which domains get credited instead of you.
What we do about it
Make the source unambiguous: original content, dating, author attribution, canonicals — and making sure it is you, not an aggregator, that is the citable original.
7

Influence

The question: Are you represented correctly, and does it move the perception?

What decides it
A citation is not an endorsement. The model can name you in the wrong category, with an outdated price, or in a comparison you lose.
How it fails
High visibility, wrong narrative. You get mentioned often and recommended rarely.
How we measure it
Sentiment and context per mention, which attributes are ascribed to you, who you are compared against, and the gap between your own positioning and the model's description (Perception Match).
What we do about it
Correct the factual base where the model draws it from, and make the positioning unambiguous enough to survive being summarised.
8

Action

The question: Does anything happen at your end afterwards?

What decides it
An AI answer sends fewer visitors than a search result, but they arrive further into the decision. The value sits in the conversion, not in the click.
How it fails
The traffic arrives, but the landing page answers something other than what the user just asked the AI.
How we measure it
AI referrals as a source of their own, brand searches, direct traffic and conversion — joined to CRM or order data.
What we do about it
Landing pages that meet the intent in the prompt, and measurement that lets a referral from ChatGPT be credited with value instead of disappearing into “direct”.

The chain at a glance

Every link has its own cause, its own measurement and its own remedy. That is why the method does not collapse them into a single score: one number tells you things are going badly, not where.

The chain at a glance
LinkWhat decides itHow we measure itTypical remedy
1. DiscoveryThe system first decides whether to answer from the model's own memory or by fetching documents. If it answers from memory, your visibility is decided by what the model already learned about you, not by what you published yesterday.Share of answers that trigger retrieval versus answers from memory, per model and per prompt. Perception Match shows what the model believes about you with no sources attached.Build the entity: consistent name, category and facts across your own site, reference works, industry sources and third-party coverage. This is slow work and is measured in quarters.
2. RetrievalThe question is split into several sub-searches, and each one pulls in a small number of documents. The candidate list is short, and it is set by the index the search runs against — not by how good the page looks.Crawlability for AI agents, index coverage, robots and render checks in the site audit, and which of your URLs actually show up as sources in the daily measurements.Technical groundwork: server-rendered content, AI crawlers allowed, clean URLs, sitemaps, language versions, and a site that can be read without executing code.
3. SelectionThe fetched documents are assessed again, this time closely against the question itself. This is where most of them fall out: retrieved, but not used.“Retrieved but not cited” per prompt — we can see which of your URLs made the source set without reaching the answer.Write to the question instead of to the keyword: the question stated in the heading, the answer first, and content that covers the whole question rather than half of it.
4. ContextOnly the top-ranked passages fit in the context window. Paragraphs compete, not websites. A long document with relevance spread through it loses to one tight paragraph.Chunkalyzer in AI Monitor shows how your content is split into passages and which passages actually get picked.Chunk optimisation: self-contained paragraphs, definitions that stand alone, tables and lists where facts belong, and cutting the warm-up before the point.
5. GenerationThe generator builds the answer from passages that answer directly, need no rewriting to make sense, and add something the other sources do not have.Verbatim reuse, and which sentences recur across models and across days.Write sentences that survive being quoted without their context: precise numbers, clear limits, your own data, and formulations nobody else has.
6. CitationBeing used and being credited are decided separately. The model can build the answer on your information and still link to somebody else.Mentions with and without a link, per model and per prompt, and which domains get credited instead of you.Make the source unambiguous: original content, dating, author attribution, canonicals — and making sure it is you, not an aggregator, that is the citable original.
7. InfluenceA citation is not an endorsement. The model can name you in the wrong category, with an outdated price, or in a comparison you lose.Sentiment and context per mention, which attributes are ascribed to you, who you are compared against, and the gap between your own positioning and the model's description (Perception Match).Correct the factual base where the model draws it from, and make the positioning unambiguous enough to survive being summarised.
8. ActionAn AI answer sends fewer visitors than a search result, but they arrive further into the decision. The value sits in the conversion, not in the click.AI referrals as a source of their own, brand searches, direct traffic and conversion — joined to CRM or order data.Landing pages that meet the intent in the prompt, and measurement that lets a referral from ChatGPT be credited with value instead of disappearing into “direct”.

What it looks like when it works

A client measures forty prompts in their category. On the left, the baseline. On the right, the same prompts eight weeks later.

Baseline

Retrieved
64 %
Selected
18 %
Cited
7 %

Eight weeks later

Retrieved
67 %
Selected
41 %
Cited
22 %

The diagnosis points at link three: we are retrieved, but selected away. The remedy is not more content — it is rewriting the pages from being about the topic to answering the question.

Selection more than doubled. Retrieval barely moved, as it should — we did not touch it. And now citation is the new bottleneck: fewer than half the answers that select us credit us. The next round goes after link six.

A constructed example with realistic numbers, not a client case.

Want to know which link you're losing at?

We run your prompt set against the models, take the baseline and hand you the diagnosis — before anyone proposes a single activity.

How we work

From “we don't know” to something measurable. After phase two you know where you stand — and you can stop there at any point and do the rest yourself.

  1. 1

    Mapping

    Works on: Links 1–2

    We work out which questions your customers actually put to an AI. Not keywords — prompts. Those are two different things, and they produce different answers. The prompt set is built around personas and the customer journey, from first symptom to “who should I choose”, and mirrored against your competitors.

    What you are left withA prompt set that mirrors your customer journey, with personas and competitors defined.

  2. 2

    Measurement

    Works on: Links 1–7

    The prompts run daily against ChatGPT, Gemini, Perplexity and Google's AI answers. We measure how often you are mentioned, how often you are cited with a link, and who is named instead of you. Daily frequency is not decoration: the same prompt can give a different answer from one day to the next, and a single answer is an anecdote rather than a measurement. CAVIS runs whole conversations rather than isolated questions, because the answer changes as the customer moves from “what is this?” to “who should I choose?”.

    What you are left withA baseline you can measure movement against — per prompt, per model and per link in the chain.

  3. 3

    Diagnosis

    Works on: Places the loss in the chain

    We work out why. Is the content invisible because it requires JavaScript? Is the entity missing? Are the facts buried in marketing copy? Are you retrieved but never selected? The causes are usually few and concrete, and each one belongs to a specific link in the chain.

    What you are left withA prioritised list of causes, not symptoms — each one pinned to the link it belongs to.

  4. 4

    Action

    Works on: Fixes where the cause is

    We carry it out — or you do, with the plan as your starting point. Technical foundation, structure, entity and content, in the order that produces the fastest effect. The order is not arbitrary: there is no point writing better paragraphs for link four if you never get through link two.

    What you are left withActions with an owner, a sequence and an expected effect — and the link each action is meant to move.

  5. 5

    Follow-up

    Works on: Measures again, same place

    Visibility in AI search is perishable — the models are updated, and your competitors are working too. The daily measurement continues, and you have your own access to the dashboard throughout. On top of that we go through the development together on a fixed rhythm, so you get the interpretation and the next action, not just the numbers.

    What you are left withDaily measurement per prompt and model, your own dashboard, and a standing review where the numbers become the next action.

And then phase three starts over. Not because something went wrong, but because the bottleneck has moved: once link three was fixed, link six became the new constraint. That is why this is not a plan with an end date.

Why AI search needs a method of its own

SEO thinking assumes a ranked list. There is no list inside an AI answer. What there is, is a chain of decisions inside the system — and most of them are taken before a single word is written.

The model often answers without searching

An answer comes either from what the model already learned (parametric memory) or from documents it fetches on the spot (retrieval). Publishing something new only moves the second one. The first moves when the brand is described widely enough, often enough and consistently enough elsewhere — over time.

The candidate list is short

When the system does retrieve, the question is split into several sub-searches, and each sub-search pulls in a small number of documents. Where a search engine shows ten blue links and a page two, an AI answer has no page two. If you are not in the candidate list, you do not exist for that answer.

Passages compete, not pages

The model does not read your website. It reads a handful of passages selected because they resemble the question. One precise paragraph beats a long page with relevance scattered through it.

Being used and being credited are two different things

Your content can carry the answer without you being named, and you can be named without being described correctly. Both are losses, and neither shows up in a rankings report.

From citation to revenue

The last two links in the chain, influence and action, are where visibility meets the business. That is where we use ABC: Acquisition, Behavior, Conversion — the reporting structure Google Analytics made common property, with AI search measured as an entry channel of its own and the conversion sourced from CRM rather than from the analytics tool.

A – Acquisition

Where customers come from, and which channels deliver quality. AI search is measured here, on the same terms as organic search and ads: volume, quality, behaviour and conversion.

B – Behavior

What they do once they are there, and where we lose them. For AI traffic the measurement that matters most is intent match: does the landing page answer what the user actually asked the model?

C – Conversion

What actually creates value. Leads and sales carry a value drawn from CRM, order data or offline conversions, and that value is fed back to the channel. Without that step you optimise for volume and call it growth.

On ABC: Acquisition, Behavior and Conversion are the reporting structure from Universal Analytics, not a term we invented. It is not the same as ABC analysis, which ranks inventory, products or customers by value, and not the same as the ABC model from cognitive therapy. In the Citation Method, ABC is the commercial layer — not the whole method. Read about the ABC of AI visibility.

What the method runs on

The method is not a slide. Every link in the chain has a tool that measures it, and most of them we built ourselves because nobody else was measuring what we needed.

CitationLab AI Monitor

The daily measurement. Runs the prompt set against ChatGPT, Gemini, Perplexity and Google AI Overview, and logs mentions, citations, sources, competitors and share of voice per model and per day. This is the data behind links one through seven.

See how AI Monitor works

CAVIS

The framework behind the prompt set. CAVIS simulates whole conversations rather than isolated questions, weights every measurement by where it sits in the buying journey, and is the reason the measurement mirrors a customer journey instead of a question list.

See the CAVIS framework

Chunkalyzer

The chunk analysis inside AI Monitor. Shows how your content is split into chunks, which ones get retrieved, and which ones never make it into the context window. This is the tool for link four.

Perception Match

Compares your own positioning — what your website and sales material claim — against what the models actually say about you when asked, and scores the gap. It is the diagnostic for link one and link seven: is the problem that you are invisible, or that the story is wrong?

Site audit for AI agents

A technical pass on whether the content can be fetched at all: rendering, crawler access, structured data, language versions and machine readability. Without it, everything else is guesswork.

See technical SEO

AEO management

The operations. Phases four and five put into a system, with specialists on technical, entity, content and conversion, and every change logged as an experiment with a hypothesis.

See how we work

Five rules the method does not bend

A method is only worth something if it also says what it refuses to do. These are the rules that decide how the numbers on this page are produced.

  • Measure before you interpret

    Raw data is logged first — the whole answer, the sources, the date, the model. The interpretation comes afterwards and can always be traced back to the answer it rests on.

  • No single score

    One combined “AI visibility score” hides exactly what you need to know: which link you are losing at. We always break it down by model, prompt, category and link.

  • Causes, not symptoms

    “We're not visible in ChatGPT” is a symptom. A diagnosis has to point at a link and a cause, or the remedy is guesswork.

  • Daily, not occasional

    The same prompt can give different answers on different days. One answer is an anecdote. The method measures daily because movement can only be read out of a time series.

  • Say what the numbers do not show

    We mark where the limit of what the measurement can prove sits. Where we cannot see causation, we say so instead of drawing an arrow.

What the method does not promise

AI visibility is a young field with a lot of confident claims in it. These are the four we will not make.

Nobody can guarantee a citation

No one sells a slot in an AI answer, and anyone saying otherwise is selling something else. The method makes the probability measurable and the causes workable. It does not buy placement.

Link one is slow

The model's memory updates when models are trained or refreshed, not when you publish. Entity work is measured in quarters. Links two through four can move in weeks.

The numbers fluctuate

The same prompt gives different answers on different days. That is why we measure daily and read trend rather than single days — and why a screenshot of one good answer is not a result.

Visibility is not revenue

A mention that never converts is a number. That is why the method does not stop at link six.

The questions the method is built to answer

Nobody arrives asking for a method. They arrive with a symptom. These are the four we hear most, and the link each one usually sits in.

Why doesn't ChatGPT mention us at all?

Almost always link one or link two, and the two need opposite remedies. If the model answers from memory without searching, the brand is missing from what it learned, and the work is entity work: consistent facts about you in the places models learn from, measured in quarters. If the model does search and you are still absent, the cause is usually mechanical — content that only appears once JavaScript runs, AI crawlers blocked in robots.txt, or no page in the language the question is asked in. The measurement separates the two, because guessing between them wastes a quarter.

We get retrieved as a source but never cited. What's wrong?

That is link three or link four, and it is the most common diagnosis we make. Being retrieved means your page made the candidate list; not being cited means it lost the re-ranking, or it made the context window and the passage that got extracted did not answer the question. The remedy is almost never more content. It is stating the question in the heading, putting the answer in the first paragraph, and making individual paragraphs self-contained enough to be lifted out and still make sense.

We get mentioned often but described wrong. How do we fix that?

That is link seven, and it is a loss disguised as a win. Share of voice looks healthy while the model puts you in the wrong category, quotes an outdated price, or names you as the cheap option in a comparison you would rather win. Perception Match scores the gap between what you say about yourself and what the models say. The fix is to correct the factual base in the sources the model actually draws on — which is rarely your own homepage — and to make the positioning unambiguous enough that it survives being summarised into one sentence.

How do we know the work is having an effect?

Because there is a baseline, and it was taken before anything changed. Phase two establishes it per prompt, per model and per link, and the same prompts keep running daily afterwards, so movement is read out of a time series rather than out of a screenshot. The links move at different speeds: technical and passage-level work shows in weeks, entity work in quarters. If a change cannot be tied to a link and a measurement, the method treats it as untested, not as a win.

Frequently asked questions about the Citation Method

The Citation Method is CitationLab's method for AEO and GEO optimisation. It models an AI answer as a chain of eight links — discovery, retrieval, selection, context, generation, citation, influence and action — measures each link separately, diagnoses which link a brand is losing at, and fixes the cause in that link rather than everywhere at once.

Because the citation is the point where visibility becomes measurable. Everything before it — being in the model's memory, being retrieved, surviving re-ranking, fitting in the context window — is invisible from the outside and can only be inferred. The citation is the first link you can observe directly, so it is the anchor the rest of the chain is measured around.

The ABC method — Acquisition, Behavior, Conversion — is a framework for measuring the commercial chain from traffic to revenue, and it is still in use here as the third layer of the Citation Method. What it does not do is explain why an AI model does or does not cite you. The Citation Method adds that: eight links describing how the answer is produced, and a diagnosis that points at one of them. ABC tells you the traffic is not converting. The citation chain tells you the traffic never arrived, and which of the eight links stopped it.

Both, and the distinction matters less than the vocabulary suggests. AEO (Answer Engine Optimization) is usually used about being the source an answer engine cites; GEO (Generative Engine Optimization) about being represented well in generated text. In the citation chain they are different links of the same chain: AEO is weighted towards links two through six, GEO towards links five through seven. We do not run them as separate services because a brand losing at link two cannot be helped by work aimed at link six.

From how retrieval-augmented generation actually works: a routing decision, a retrieval step, a re-ranking step, a context window with a hard limit, a generation step, and a citation step that is decided separately from the generation. The last two links, influence and action, are ours: they cover what happens after the answer is produced — whether the description is correct, and whether anything happens at your end. The model is descriptive, not proprietary. What is ours is measuring each link separately and treating the diagnosis as the deliverable.

The baseline exists after two to four weeks, which is when you first know where you stand. Technical work in link two and passage work in link four typically show within weeks. Selection in link three moves over one to two months. Entity work in link one is the slow one and should be judged in quarters, because the model's memory only updates when the models do.

ChatGPT, Gemini, Perplexity and Google AI Overview as standard, measured separately rather than averaged. Keeping them apart matters more than it sounds: the same brand is regularly dominant on one model and absent on another, and the causes differ — one of them may be reading a rendered page while another never gets past the JavaScript.

For phases one and two, nothing from you but your website and the names of your competitors — the measurement runs against public model answers. To close the chain at link eight we also need analytics, Search Console and whatever holds the conversion value: a CRM, order data, or an offline process. Without that last part, link eight is guesswork, and we would rather say so than report a number for it.

Yes, and it is a normal way to use the method. After phase two you have a baseline, and after phase three you have a prioritised list of causes tied to links in the chain. Many clients take that in-house and run the actions themselves. The daily measurement is what most of them come back for, because a plan without a time series behind it cannot tell you whether it worked.

Yes, and it is not an add-on. Every prompt returns whoever the model named, so competitor visibility is a by-product of measuring your own. It matters because share of voice in AI answers is a fixed-size pool: an answer names a handful of providers, and being absent from it is the same outcome whether the reason is your content or somebody else's.

No. The chain works the same at any size — the difference is how many prompts and categories it takes to cover the market you sell into. A specialist with one service area is often easier to move than a large brand with a diffuse entity, because link one rewards being unambiguous rather than being big.

Most AEO offerings start with the remedy — more content, more structured data, more links — and report on activity. The Citation Method starts with the diagnosis and refuses to act before it has one. Every action has to name the link it is meant to move and the measurement that will show whether it did. If it cannot, it does not get done.

Let your competitors guess.
We measure the chain.

Eight links, five phases, and a diagnosis that points at one of them. Take control of how AI describes and recommends you.

Was this page helpful?