AI citations

How AI Chooses Sources: What Gets Your Pages Cited

How AI chooses sources to cite in ChatGPT, Perplexity, Gemini and AI Overviews, plus a checklist to make your pages easier to find, read and cite.

AI assistants mostly choose sources by searching the web at question time, reading a handful of the results, and citing the pages that supplied specific facts for the answer. In practice, pages that rank for the assistant’s own sub-queries, answer the question early and state checkable details tend to be cited most.

No AI company publishes a complete recipe for how AI chooses sources, and the systems change often. But by looking at which pages get cited, repeatedly and across assistants, a few consistent patterns emerge. Those citations matter: they’re one of the few signals you get about what an assistant read before it wrote about your brand. This guide covers the patterns and what you can do about them.

Key takeaways

  • Most assistants cite pages they retrieved through a live web search, so being findable for the right queries comes first.
  • Pages that answer a specific question clearly and early tend to be cited more than pages that bury the answer.
  • Third-party pages (reviews, comparisons, community threads) are cited heavily for “best X” questions.
  • Each assistant has its own habits. Track AI citations per assistant rather than assuming they agree.
  • You influence citations by improving your own pages and by being present on the pages assistants already trust.

Two ways an assistant “knows” about you

It helps to separate two sources of knowledge:

  1. Training data. What the model learned before its cutoff. This shapes general impressions of your brand but doesn’t produce citations, and you can’t update it directly.
  2. Retrieval. Pages the assistant fetches at question time, usually through a web search. These are the pages that show up as citations.

For commercial questions, like “best CRM for a startup”, assistants with web access tend to search, especially when the question implies current information. That’s good news: retrieval is something you can work on today.

How AI chooses sources: the retrieval pipeline, simplified

In broad terms, an assistant with search does something like this:

  1. Rewrites the question into search queries. ChatGPT, for example, often runs several queries for one prompt. This is called query fan-out, and the queries it chooses decide which pages are even in the running.
  2. Gets results from a search index. Different assistants rely on different search providers or their own indexes, which is one reason their citations differ.
  3. Reads a subset of pages. Not every result gets read. Pages that load quickly and present their content as text are easier to use.
  4. Writes the answer and attributes parts of it. Pages that contributed specific facts or claims are the ones that tend to appear as citations.

Each stage filters. A page has to be found, then read, then actually used.

What tends to get a page cited

These are patterns, not guarantees. They line up with what’s publicly known about retrieval systems and with what you can observe by tracking citations over time.

It ranks for the sub-queries, not just the head term

If the assistant searches “CRM with free plan for small teams” and your page doesn’t rank for that, it’s unlikely to be read. Looking at the actual searches an assistant runs is often more useful than your keyword list.

It answers the question directly

Pages that state the answer in plain sentences near the top are easier to quote. “Our CRM is free for up to two users and includes email tracking” is more usable than three paragraphs about your mission.

It contains specific, checkable facts

Pricing, plan limits, integrations, supported regions, setup time. Concrete details give the assistant something to cite. Vague claims (“powerful”, “best-in-class”) give it nothing.

It’s structured and accessible

Clear headings, comparison tables, short lists and FAQ sections make it easier for a system to extract a specific piece of information and attribute it. Content that requires JavaScript to render, sits behind a login, or is blocked to crawlers may never be read at all.

It’s current

For questions with a time element (“in 2026”, “latest”), recently updated pages tend to be favored. A visible, accurate “last updated” date helps readers and machines alike.

It’s from a source the category already trusts

For “best X” questions, assistants lean heavily on review platforms, comparison articles, established media and community discussions. If those pages consistently get cited in your category, being covered there matters as much as your own content. Our guide on getting cited through review sites goes into detail.

Why ChatGPT citations differ from other assistants

ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews don’t share an index or a selection process. In practice you’ll see one assistant favor community threads, another favor big publishers, and another cite brand sites more readily. That’s why a page that drives your visibility in one assistant may be invisible in another. Google AI Overviews behave differently again because they sit inside a results page; see what AI Overviews mean for your traffic.

A citation-readiness checklist for your key pages

Run this on your homepage, pricing page, main product pages and any comparison pages:

  • The first paragraph says what the product is, who it’s for and what makes it different
  • Pricing and plan limits are written as text, not only in images
  • Key facts (integrations, limits, regions) are listed explicitly
  • Headings match questions people ask
  • There’s a short FAQ answering common buyer questions
  • The page renders meaningful text without client-side JavaScript
  • The page isn’t blocked in robots.txt for the crawlers you care about
  • A last-updated date is visible and accurate
  • Claims are specific and verifiable

See how AI chooses sources in your category with Citegram

Don’t guess which sources matter. Log them. Citegram is a Chrome extension that does the logging for you:

  1. Add the prompts your buyers ask. Start with your most important category and comparison questions.
  2. Run them across assistants. Citegram asks each prompt in the real ChatGPT, Gemini, Perplexity, Claude, DeepSeek, Grok and Google AI Overviews apps in background tabs, using the accounts already signed in to your browser. No API key is needed.
  3. Read the citations in rank order. Every answer is saved with its cited sources in order, plus whether and where your brand is mentioned.
  4. Check ChatGPT’s searches and Google’s AI Overviews. For ChatGPT you see the web searches it ran, which shows the sub-queries you need to rank for. For Google you see whether an AI Overview appeared and which sources it used.
  5. Review the most cited domains. Reports list the domains that come up most, so after a few runs you can sort them into your own pages, competitor pages, third-party pages that mention you and third-party pages that don’t.

That last group is your outreach list. Your own cited pages tell you which content is working.

FAQ

How does AI choose sources for its answers?

Mostly through retrieval. Assistants with web access tend to search, read a few of the results, and cite the pages that supplied specific facts. Ranking for the assistant’s sub-queries and stating facts clearly near the top of the page both appear to help.

How does ChatGPT decide which sources to cite?

When it browses, ChatGPT typically turns your question into several web searches, reads some of the results and cites the pages it drew on. OpenAI doesn’t publish the full selection logic, so the practical approach is to observe its citations and searches across repeated runs.

Can I pay to be cited by an AI assistant?

Organic citations in chatbot answers aren’t something you can buy directly. Some platforms are introducing ad formats, but those are separate from the sources an assistant cites. The reliable path is to be on the pages assistants retrieve.

Why is a competitor’s page cited for a question about my product?

Usually because it ranks for the query the assistant ran and contains specific statements about your product. Check what it says, correct inaccuracies where you can, and publish your own page that answers that question more clearly.

See which sources shape your answers

If you want to see the exact sources behind what assistants say about your brand, the free plan lets you track one prompt across all major assistants, and Pro adds unlimited prompts for $10 a month. Have a look at the plans.

Free Chrome extension

See if ChatGPT recommends your brand

Check one question for free in ChatGPT, Gemini, Perplexity, Claude and 3 more assistants.

Check my brand for free