GEO

LLM SEO: How Large Language Models Pick the Brands They Mention

LLM SEO (LLMO) explained: what models learn about brands in training, what they retrieve from the web at answer time, and how to work on both.

LLM SEO, also called LLMO (large language model optimization), is the work of influencing which brands a large language model mentions and how it describes them. A model knows about your brand through two channels: what it learned from training data before its cutoff, and what it retrieves from the web when it answers. LLM SEO means working on both, knowing that they move at very different speeds.

Most GEO advice focuses on the second channel, because that’s where citations come from. This article goes deeper on the model side: what AI vendors document about training data, knowledge cutoffs, search triggers and crawlers, and what each implies for a brand.

Key takeaways

  • LLM SEO (or LLMO) targets how the model itself mentions and describes your brand, not only which pages it cites.
  • Models know about brands through training data, which is frozen at a cutoff date, and through retrieval, which happens at answer time.
  • Vendors document separate crawlers for training and for search, so blocking a training crawler doesn’t remove you from AI search results.
  • The model usually decides whether to search. Vendor docs say it tends to search for current, changing or product-specific information.
  • Training-based knowledge changes slowly and only with new model versions. Retrieval can reflect your changes as soon as pages are crawled again.
  • An answer with no cited sources most likely came from what the model already knew, which tells you which channel to work on.

What LLM SEO means, and how it differs from GEO

The terms overlap. Our glossary entry on LLMO treats LLM SEO as a variant of generative engine optimization, and in practice most of the tactics are shared. The difference is emphasis:

  • GEO and AEO usually start from the answer page: which sources get cited, and how to be one of them.
  • LLM SEO starts from the model: what it believes about your category, which brands it associates with which needs, and when it bothers to check the web at all.

That view matters because not every answer is grounded in a search. If an assistant answers “what’s the best CRM for a 20-person startup?” from memory, no page can be cited.

Channel 1: training data

What vendors document

AI companies describe their training data in broad terms only. Anthropic’s Transparency Hub says recent Claude models were trained on a proprietary mix that includes “publicly available information from the internet”, along with public and private datasets and synthetic data. Anthropic’s models overview lists a “reliable knowledge cutoff” for each model, defined as the date through which its knowledge is “most extensive and reliable”, and a separate training data cutoff.

Vendors also name the crawlers that collect public web content for training. OpenAI’s crawler documentation says disallowing GPTBot “indicates a site’s content should not be used in training generative AI foundation models.” Anthropic’s crawler help article describes ClaudeBot as collecting web content that could contribute to training. Google’s crawler documentation describes Google-Extended as the control for whether content may be used for training future Gemini models.

What it implies for brands

  • It’s a snapshot. Anything that happened after the cutoff, such as a new product line, a price change or a rebrand, isn’t in the model’s own knowledge.
  • You can’t edit it. There is no form to submit corrections to a model’s training data. Changes arrive, if at all, with a future model version.
  • Consistency likely helps. A model learns associations from many documents, so a brand described the same way across many independent pages has better odds of being learned that way. That’s an inference, not a documented rule.
  • Old content lingers. Outdated reviews or descriptions can keep shaping how a model frames you long after you’ve changed. Our guide on fixing wrong AI information about your brand covers what you can do.

Channel 2: retrieval at answer time

What vendors document

Retrieval, also called grounding, is when the assistant searches the web while answering and uses the results. Vendor developer docs describe who decides to search and when:

  • Claude. Anthropic’s web search tool documentation says Claude “determines when to search based on the prompt.” It searches when a request depends on information that is current, changing or outside its training data, including “information about specific organizations, people, or products that might have changed”, and answers directly for stable knowledge.
  • Gemini. Google’s Grounding with Google Search documentation says the model “analyzes the prompt and determines if a Google Search can improve the answer”, then generates one or more search queries.
  • Google AI Overviews and AI Mode. Google’s AI features documentation says both may use query fan-out, running several related searches.
  • ChatGPT and Perplexity. Their crawler docs describe dedicated search crawlers: OAI-SearchBot for ChatGPT’s search features, and PerplexityBot, which Perplexity’s bot documentation says is designed “to surface and link websites in search results” and is not used to crawl content for foundation models.

These are mostly developer docs. Consumer apps can behave differently, so read them as a description of the mechanism, not a guarantee.

What it implies for brands

Ranking for the model’s sub-queries, stating clear facts on your pages and being covered on third-party pages are what count here. How AI assistants choose sources and query fan-out explained cover that pipeline.

How the two channels interact

The model decides whether to search and writes the queries itself. It’s reasonable to expect that what it already knows about a category shapes those queries and how it frames the results. A model that strongly associates “startup CRM” with two brands may still name them first while citing pages that list five.

Training data Retrieval
When it updates With a new model version Every time the assistant searches
What it covers Knowledge up to the cutoff Current pages it can find and read
Crawler examples GPTBot, ClaudeBot, Google-Extended OAI-SearchBot, Claude-SearchBot, PerplexityBot
Shows citations No Usually yes
What you can influence Your long-term public footprint Rankings, page content, third-party coverage
Time to effect Months, unpredictable As soon as pages are recrawled

Training and search crawlers are separate switches

Because vendors split crawlers by purpose, your robots.txt choices affect the channels differently. OpenAI’s documentation states that “each setting is independent of the others”, so a site can allow OAI-SearchBot while disallowing GPTBot. Google says Google-Extended “does not impact a site’s inclusion in Google Search” and is “not used as a ranking signal in Google Search.”

So blocking training crawlers doesn’t hide you from AI search, while blocking search crawlers, sometimes by accident through a broad rule, can remove your pages from retrieval. Our guide on AI crawlers and robots.txt lists the user agents, and the AI robots.txt generator helps you write rules per crawler.

An LLM SEO plan for both channels

  1. Diagnose which channel each answer uses. Run your buyer prompts and note which answers cite sources and which don’t. Uncited answers point to the model’s own knowledge.
  2. Fix retrieval first. It’s faster. Check crawler access, make your pages state plain facts, and target the searches assistants actually run.
  3. Make your brand description consistent everywhere. Use the same category, positioning, pricing and product names on your site, profiles, directories and partner pages. See entity SEO for AI.
  4. Earn independent coverage. Reviews, comparisons, media and community discussions feed both channels: they’re retrieved today and may be part of tomorrow’s training data. Our article on backlinks vs brand mentions explains why mentions matter.
  5. Re-measure after model updates. A new model version can change the uncited answers overnight. Re-run your prompts when a vendor ships one.

Track both channels with Citegram

Citegram is a free Chrome extension, with a dashboard, that shows what each assistant says and what it relied on:

  1. Add your buyer prompts and the competitors to compare against.
  2. Run them in the real apps. Citegram asks each prompt in ChatGPT, Gemini, Perplexity, Claude, DeepSeek, Grok and Google AI Overviews, in background tabs, with the accounts already signed in to your browser. No API key is needed.
  3. Separate cited from uncited answers. Each answer is saved with mentions, your position and the sources and links it cited. Answers without sources point you to the training-data side.
  4. Read ChatGPT’s searches. For ChatGPT, Citegram records the web searches it ran, so you see the queries the model wrote for itself.
  5. Watch the trend. Share of voice per brand and per assistant, and the most cited domains, over repeated runs, so a model update shows up as a change in the curve.

FAQ

What is LLM SEO?

LLM SEO, or LLMO, is the practice of influencing how large language models mention and describe your brand. It covers both what models learned during training and what they find through web search when answering. It overlaps heavily with GEO and AEO, with more attention on the model’s own knowledge.

How do large language models decide which brands to mention?

From two sources: associations learned from training data before the model’s cutoff, and pages retrieved from the web at answer time. Vendor docs say models tend to search when a question involves current or changing information, such as products and prices. When they don’t search, they answer from training alone.

Can I get my brand into ChatGPT’s training data?

Not directly, there’s no submission process. Public web content may be collected by training crawlers such as GPTBot, but what ends up in a model isn’t disclosed. The practical route is a consistent, widely published brand description, plus strong retrieval visibility meanwhile.

Does blocking GPTBot hurt my visibility in ChatGPT?

According to OpenAI’s crawler documentation, GPTBot is for training and OAI-SearchBot is for search, and each setting is independent. Blocking GPTBot signals your content shouldn’t be used for training. It doesn’t block ChatGPT’s search crawler, as long as you allow OAI-SearchBot.

See what models say about you

LLM SEO starts with reading the answers, cited or not. The free plan tracks 1 prompt and 1 competitor across all major assistants, forever. Pro is $10 per seat per month with unlimited prompts and competitors. See pricing.

Free Chrome extension

See if ChatGPT recommends your brand

Check one question for free in ChatGPT, Gemini, Perplexity, Claude and 3 more assistants.

Check my brand for free