How to Optimize Your Website for AI Search

How ChatGPT, Claude, Perplexity, and Google's AI Overviews each find a page, which AI crawlers to allow, how to write a page they quote, and how to check.

Author
Mike Murphy, Owner and Founder
Published
Reading time
10 min read
A flat illustration of a signal mast on a rocky point beside a small stone hut, flying five plain colored flags, with three boats at different distances on the water.

You optimize a website for AI search by making each page reachable by the engines' crawlers, easy to quote, and consistent with what other sites say about you. Those three jobs cover Google's AI Overviews and AI Mode, ChatGPT search, Claude, and Perplexity. The advice that ranks for this question tends to list the same dozen tips without saying how any of it works. This guide goes engine by engine instead. If the terms are new, start with what AEO is and come back.

  1. Reach. Each engine gets to your page by a different route. If the route is closed, nothing else on this list matters.
  2. Quotability. An engine lifts a passage, not a page. The passage must make sense by itself.
  3. Corroboration. The engine reads other sites about you, like directories, reviews, and profiles. When they disagree with your site, the engine hedges or picks another business.

What is not on the list: special AI markup, a new file format, or anything that guarantees a place in an answer. No such thing exists, and the engines say so.

How does each AI search engine find your pages?

Google's AI answers draw on Google's own search index, while ChatGPT, Claude, and Perplexity each fetch pages with a search crawler of their own. That difference decides where you look when a page is missing.

  1. Google AI Overviews: Googlebot, through Google's search index. You control indexing and snippets.
  2. ChatGPT search: OAI-SearchBot, fetching the page directly. You control robots.txt.
  3. Claude: Claude-SearchBot, fetching the page directly. You control robots.txt.
  4. Perplexity: PerplexityBot, fetching the page directly. You control robots.txt.
Figure 1. Four routes from one website to four AI answers. Google's runs through its search index; ChatGPT, Claude, and Perplexity send their own crawlers.
  • Google AI Overviews and AI Mode. Google says a page must be "indexed and eligible to be shown in Google Search with a snippet" to appear as a supporting link, and that there are no additional requirements (Google Search Central, 2025). Both features may run several related searches for one question and cite pages from any of them. The route is Googlebot and the search index you already know.
  • ChatGPT search. OpenAI's documentation names OAI-SearchBot as the crawler "used to surface websites in search results in ChatGPT's search features" (OpenAI, crawler overview). A site that blocks it will not be shown in ChatGPT's search answers.
  • Claude. Anthropic names Claude-SearchBot, which "navigates the web to improve search result quality for users," and says that disabling it "may reduce your site's visibility and accuracy in user search results" (Anthropic, crawler documentation).
  • Perplexity. Perplexity names PerplexityBot as the crawler "designed to surface and link websites in search results on Perplexity," and recommends allowing it in robots.txt (Perplexity, crawler documentation).
EngineHow it reaches your pageWhat you control
Google AI Overviews, AI ModeGooglebot and Google's indexIndexing, and snippet controls such as nosnippet
ChatGPT searchOAI-SearchBotIts line in robots.txt
ClaudeClaude-SearchBotIts line in robots.txt
PerplexityPerplexityBotIts line in robots.txt

For Google, the test is Search Console. Inspect the URL and confirm it is indexed. For the other three, open your robots.txt (it lives at /robots.txt on any domain) and read what it says about each crawler by name, and about all crawlers under User-agent: *.

Which AI crawlers should you allow, and which can you block?

Allow the search crawlers if you want to be cited, and decide about the training crawlers separately, because the vendors document them as different bots with different jobs.

  • OpenAI runs GPTBot to collect content that may be used to train its models, and OAI-SearchBot for search. OpenAI states that "each setting is independent of the others" and that a robots.txt change can take about 24 hours to take effect (OpenAI, crawler overview). Blocking GPTBot does not take you out of ChatGPT search.
  • Anthropic runs three agents. ClaudeBot collects web content that could contribute to model training, and Anthropic says restricting it "signals that the site's future materials should be excluded from our AI model training datasets." Claude-SearchBot indexes pages for Claude's search results. Claude-User fetches a page when a person asks Claude a question, and Anthropic says disabling it "may reduce your site's visibility for user-directed web search" (Anthropic, crawler documentation). Each has its own line in robots.txt, so a site can block ClaudeBot and still allow the two that serve Claude's answers.
  • Perplexity says PerplexityBot "is not used to crawl content for AI foundation models." Its second agent, Perplexity-User, fetches a page when a person asks a question, and Perplexity says it "generally ignores robots.txt rules" because a user requested the fetch (Perplexity, crawler documentation).
  • Google offers Google-Extended to control whether content Google crawls may be used to train future Gemini models. Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" (Google Search Central, 2026). To keep content out of AI Overviews, Google points to the same controls it uses for snippets: nosnippet, data-nosnippet, max-snippet, and noindex.

The common mistake goes the other way. A site blocks every AI user agent to stay out of training data, and quietly disappears from the search answers it wanted to be in. Read the file line by line before assuming it is fine. On rivetbay.com both kinds are allowed, and our AEO vs GEO vs SEO guide shows the rest of that setup.

How do you write a page an AI engine will quote?

Put the question in the heading and the answer in the first sentence under it, so the sentence still makes sense when it is lifted out alone. That habit does more than any technical change.

Compare two openings under the heading "How long does it take to build a custom website?"

  • Weaker: "Every project is different, and a lot of factors go into how long things take."
  • Stronger: "A custom marketing site typically takes four to six weeks, and a larger project eight to twelve."

The first says nothing an engine can use. The second answers the question, gives the range, and holds up out of context. It is also the answer we give about our own projects.

A few more habits follow from the same idea:

  • One idea to a section. A section that answers two questions gets quoted for neither.
  • Facts stated plainly, with dates where they change. Hours, service areas, and prices change over time. An engine that finds two versions trusts neither.
  • The copy in the HTML. If the text only appears after a script runs, some crawlers see an empty page. Check a page's source code and find its first sentence.
  • A named author. A person with a bio page gives an engine someone to attribute the answer to.

What AEO is, and what it looks like on a real site walks through each of these on rivetbay.com.

Neither is required, and Google says plainly that neither helps in Google's AI features. Its guidance states that "structured data isn't required for generative AI search, and there's no special schema.org markup you need to add," and that an llms.txt file "will neither harm nor help your site's visibility or rankings in Google Search" (Google Search Central, 2026).

We use both on our own site anyway, for a narrower reason. Structured data states once, in a form a machine cannot misread, who the business is, who wrote each page, and how the two relate. An llms.txt is a plain summary that some tools outside Google read. Neither is a lever. Treat them as good housekeeping, and be wary of anyone who sells either as the fix.

Why do AI engines cite directories instead of your website?

AI engines cite directories and review sites because those pages already compare many businesses in one place, which is the shape of answer a "best" or "who should I hire" question needs. We saw this in our own check in September 2026. We asked ChatGPT and Perplexity seven questions a buyer of web design might ask. For the questions about which agency to hire, both engines leaned on directories and lists, among them Clutch, DesignRush, Built In Boston, and Semrush's agency directory, and named the firms those lists carried.

Two things follow for any business, not only agencies:

  • Your listings are part of your AI search presence. A directory profile, a review site, and your Google Business Profile are pages an engine reads about you. If these pages are missing, thin, or out of date, the engine will describe you using what it finds, or not at all.
  • They have to match the site. The same name, phone number, address, and list of services everywhere. Our Google Business Profile guide covers the profile most engines see first.

The same check turned up a second pattern. For several of the plainer questions, such as what AEO is, ChatGPT answered from its model with no sources at all. Those answers cite nobody, so no website change earns a place in them. The questions where sources do appear are the ones worth working toward.

How do you check whether ChatGPT, Claude, Perplexity, or Google cite you?

Ask them. Take ten questions your customers really ask, put each to ChatGPT, Claude, Perplexity, and Google, and write down who is cited. Do it again next month with the same questions. Your list of who is cited is your baseline and competitor list; it is more useful than any tool score.

  1. Collect the questions. Ask whoever answers the phone or the inbox for the ten they hear most.
  2. Ask each engine the same way. Use a private window, and turn on web search in ChatGPT and Claude so they look at the web.
  3. Log the sources, not the prose. For each answer, note which sites are cited and which businesses are named.
  4. Check your analytics. Google says that traffic from AI Overviews and AI Mode appears in Search Console's Performance report under the "Web" search type, which is part of your ordinary search traffic (Google Search Central, 2025). Visits from the assistants that carry a referrer show up in your analytics under their domains, such as chatgpt.com, claude.ai, and perplexity.ai.

The AEO website checklist turns the whole routine into checks and a 30-day plan.

What should you not pay for?

Do not pay for guaranteed placement in AI answers, because nobody can sell it. The engines choose sources per question, from their own index or crawl, and none of them offers a paid route into the cited links.

  • A promised ranking in ChatGPT or Claude. There is no ranking to promise. There is a set of sources for each question, and it changes.
  • An llms.txt as a product. It is a short text file. Publishing one takes an hour, and Google says it does not affect Google Search.
  • "AI optimization" that skips the basics. If the pages are not indexed, the answers are not in the first sentence, and the listings disagree with the site, no add-on fixes that.

An opinion, with reasons: most of what is sold as AI search optimization to small businesses is SEO with a new label. The honest part of the new label is real, and it is narrow. Know which crawlers serve which engine, write passages that stand alone, keep your facts consistent across the web, and check what the engines say.

Where should you start?

Start with reach, because a page an engine cannot fetch cannot be quoted. Read your robots.txt this week and confirm your key pages are indexed in Search Console. Then rewrite the headings and first sentences on the five pages customers visit most. Run the ten-question check and correct the listings it finds.

If you would rather have a site built this way from the start, our AEO services cover the build and the content work after it. Start a project and we will begin with the questions your customers ask.

Frequently asked questions

How can I rank my website in AI search engines?

Nobody ranks in an AI answer the way a page ranks in a list of results. The engine picks a few sources to cite for each question. To be one of them, let each engine's crawler reach the page, answer the question in the first sentence under a heading that asks it, and make sure the directories and profiles the engine also reads describe your business the same way your site does.

How can I optimize content to get cited by AI search engines?

Write each section so it can be lifted out whole. Put the question in the heading and the answer in the first sentence, keep one idea to a section, and state facts plainly, with dates where they change. Then make sure the page is indexed by Google and open to the search crawlers from OpenAI, Anthropic, and Perplexity. A page an engine cannot fetch cannot be quoted, however well it is written.

Can I optimize my website for AI search myself?

Much of it, yes. Checking robots.txt, confirming the page is indexed in Search Console, rewriting headings as questions, and asking the engines your customers' questions each month are all work an owner can do. The harder parts are a site that loads its copy in the HTML rather than after a script runs, and a structure that holds up as the site grows. Those are build decisions, and they are easier to make once than to retrofit.

Is SEO still worth it in 2026?

Yes, and AI search makes it more so. Google says its AI Overviews and AI Mode are rooted in its core Search ranking and quality systems, and that a page must be indexed and eligible for a snippet to be shown in them. A page that does not do well in ordinary search has little chance of being cited in Google's AI answers. The fundamentals are the entry ticket for both.

Does blocking GPTBot remove my site from ChatGPT?

No. OpenAI documents GPTBot and OAI-SearchBot as separate crawlers with independent settings. GPTBot collects content that may be used to train its models. OAI-SearchBot is the one that surfaces sites in ChatGPT's search answers. You can block training and still allow search, and a site that blocks OAI-SearchBot will not be shown in ChatGPT's search answers.