Most of what gets a page quoted by an AI assistant is plain old SEO done properly. The rest is making sure every fact about your product says the same thing everywhere a machine might read it. That second part is where I lost the most time.

Over the last months I worked through search and AI discovery on several product sites: a regulated-industry AI tool, a compliance SaaS, a local services marketplace, and this site. Different stacks (static React with a prerender step, Next.js, a tiny Node build), same problems. This is the runbook I wish I had at the start. It is summarised on purpose. The order matters more than any single trick.

The order that worked
01CrawlableHTML, canonicals, sitemap
02ConsistentOne source for every fact
03QuotablePassages that stand alone
04AnnouncedTell engines what changed
05MeasuredFixed prompts, real baselines

Skipping a step does not save time. It moves the bug to a place you cannot see.

AI answer engines do not rank pages. They lift passages. Write and structure every page so that any one paragraph still makes sense when it is the only thing quoted.

Get the boring foundations right.

Google says it plainly: AI Overviews and AI Mode use the same indexing and snippet eligibility as normal search. No special file, no magic schema. The same is roughly true for the others, because ChatGPT search, Copilot, Perplexity and Claude all lean on a search index before they read your page. If the page is not crawlable and indexable, nothing downstream matters.

01Render the content as HTML

If your marketing pages are a client-side app, prerender them. Crawlers that run JavaScript are the exception, and most AI fetchers are not in that club. Important text must be in the initial HTML, not behind a tab, an accordion that loads lazily, or an image.

02One page, one intent, one canonical

03Keep the sitemap honest

Only canonical, 200-status, indexable URLs. Change lastmod only when the visible content really changed. Bumping every date on every deploy teaches engines that your dates mean nothing, and then they ignore them when it counts.

Small pitfall that cost me an afternoon: meta description tags split across several lines. Browsers and Google handle it fine. Several audit tools and line-based scrapers reported every page as having no description at all. Write head tags on one line and save yourself the confusion.

One source of truth for every fact.

This is the big one. Every product site I checked had the same disease: facts drifted apart. The FAQ structured data on one homepage contained four questions. Two of them were no longer on the page. Pricing in the JSON-LD said less than the pricing page a human could read. The plain-text summary for AI crawlers left out the costs that actually mattered to a buyer.

Nobody did anything wrong. Someone edited the page, and the three other copies of the same fact did not move with it. That is the normal outcome when the same claim lives in four places.

Drifts Page copy, JSON-LD, llms.txt and the FAQ each hold their own hand-written version of the price.
Holds One plans module. The pricing grid, the Offer nodes and the text summary all render from it.

The fix is architectural, not editorial. Put prices, plans, FAQ answers, feature limits and coverage numbers in a data module. Render the visible page, the structured data and any machine-readable copies from that module. Then add one test that fails when they disagree:

test("FAQ schema matches the visible FAQ", () => {
  const html = renderPage("/faq");
  for (const { question } of faq) {
    expect(html).toContain(question);
  }
  expect(faqSchema().mainEntity).toHaveLength(faq.length);
});

Boring. Also the single most valuable change in this whole list, because an AI answer that states the wrong price is worse than no answer.

Structured data that helps, not hurts.

Schema.org markup matters more for answer engines than for classic rich results, which Google shows inconsistently anyway. Assistants read it as a compact statement of what the entity is. So it has to be correct, and it must never claim more than the page does.

04Give entities stable identities

Use one Organization and one SoftwareApplication with fixed @id values and reference them from every page. I had a product page that quietly built a second application node and lost the shared offers and features in the merge. From the outside it looked like two products with half the facts each.

05Validate the fiddly fields

{
  "@type": "Offer",
  "name": "Team",
  "price": "49",
  "priceCurrency": "EUR"
}

06Treat sameAs as a pointer, not a wish

sameAs links your entity to profiles elsewhere. Add a URL only once that profile exists, is filled out and links back. Pointing at empty or missing directory pages lowers confidence instead of raising it. On one site I also found the organization marked as "the same as" a developer's personal GitHub profile. It is not, and it muddied who the company actually is.

knowsAbout and areaServed on the Organization are underrated. Before any third party writes about you, they are one of the few direct signals tying a brand name to a topic and a market.

07FAQPage only where the FAQ is visible

Mark up questions and answers only on the route where they appear, and only if every marked-up pair is on that page. Copying the homepage FAQ schema onto every route looks generous and reads as spam.

Write passages worth quoting.

This is where GEO differs from SEO in practice. A search result sends a person to your page. An assistant takes one paragraph and puts it in its own answer, surrounded by other sources. Your paragraph loses its heading, its neighbours and your logo.

08Answer first, explain second

09Be specific, and be careful

Vague claims do not get quoted. Numbers, dates, document names, limits and procedures do. But specificity cuts both ways, because now the claim can be checked. Things I had to walk back after a review:

If an assistant repeats your page, it repeats your exaggerations too, under your name. A mention with a wrong capability is not a win.

10Build a real answer page

The most useful single page was a plain product FAQ covering what buyers actually ask: who it is for, what it costs, which integrations exist today, what is not supported, and how data is handled. Server-rendered, stable fragment links per answer, linked from navigation and the product page. It turned out to be the page assistants had been missing.

llms.txt and friends, in proportion.

I shipped llms.txt, llms-full.txt and Markdown copies of long articles on several sites. Cheap to add, easy for an agent to read. But let us be honest: no major engine has said it ranks on them. They are a convenience layer, not a strategy.

Crawlers, caching, and deploys.

11Know which bot does what

Search crawlers and training crawlers are different user agents. OpenAI separates OAI-SearchBot from GPTBot. Anthropic has Claude-SearchBot plus a user-triggered fetcher. Perplexity has PerplexityBot and publishes IP ranges. You can allow search and still decide about training separately.

User-agent: *
Allow: /
Disallow: /api/
Disallow: /account/

Sitemap: https://example.com/sitemap.xml

Pitfall: adding a friendly User-agent: GPTBot group with Allow: /. A bot that matches its own group ignores the * group, so you just dropped your exclusions for that bot. If the wildcard already allows public pages, leave it alone. And never treat robots.txt as access control. Private data needs real authorization.

A curl -A "PerplexityBot" returning 200 is a smoke test, not proof. Real crawlers come from their own IPs and can still hit a WAF or geo rule. Check your CDN logs before concluding anything.

12Do not cache crawler files like assets

This one bit me twice, on two different sites. The deploy synced everything with the year-long immutable cache meant for hashed JS and CSS. Sitemap, robots.txt, llms.txt, RSS feeds and Markdown copies included. So crawlers were served a sitemap that predated the new pages. For weeks.

# hashed assets: cache forever
aws s3 sync dist/ s3://bucket --cache-control "public,max-age=31536000,immutable" \
  --exclude "*.xml" --exclude "*.txt" --exclude "*.md" --exclude "*.html"
# crawler files and pages: five minutes
aws s3 sync dist/ s3://bucket --cache-control "public,max-age=300" \
  --exclude "*" --include "*.xml" --include "*.txt" --include "*.md" --include "*.html"

Then add a post-deploy check that fails if the sitemap comes back with immutable. Also check that an unknown slug is a real 404, not a 200 with a "not found" message. Soft 404s and stale CDN route allowlists are quieter versions of the same bug.

Announce changes and build clusters.

13Push changes with IndexNow

Bing's index feeds ChatGPT search and Copilot, and Bing supports IndexNow. After the production smoke tests pass, the pipeline sends the changed URLs. It is a notification, not a guarantee, but it shortens the wait from weeks to days. I saw one engine still quoting homepage copy three weeks old before this was in place.

curl -X POST https://api.indexnow.org/indexnow \
  -H "Content-Type: application/json" \
  -d '{"host":"example.com","key":"KEY","urlList":["https://example.com/pricing"]}'

14Cluster content, do not multiply it

Measure without fooling yourself.

This is where the discipline slips, because it is tempting to ship the changes and declare victory. There is no clean AI analytics yet. What worked well enough:

  1. Fix a prompt set. About ten questions per product, a mix of branded ("What is X and who is it for?") and unbranded ("Which tools help with Y?").
  2. Run each prompt three times per engine. Fresh sessions, web search on. Answers vary more than you would like.
  3. Record everything. Date, engine, mode, locale, full answer, cited URLs, factual errors.
  4. Track three things separately. Mentioned or not, correct or not, clicked through or not.
  5. Use the tools that exist. Bing Webmaster Tools has an AI performance report with citations and grounding queries. Search Console mixes AI traffic into normal web data, so do not label it all AI.
  6. Baseline before you change anything. Otherwise any number you quote later is a guess.

The honest finding across all of this: technical work makes a site eligible to be quoted. Independent references make it likely. Case studies, reproducible walkthroughs, third-party mentions, visible maintenance. No schema field substitutes for someone else saying you are worth reading.

Pitfalls, ranked by the time they cost me.

  1. Fact drift. Structured data, FAQ, text summaries and pages disagreeing. Fix with one data source and a parity test.
  2. Immutable cache on crawler files. Engines read old sitemaps for weeks.
  3. Overclaiming. Assistants quote it back, confidently, under your brand.
  4. Broken entity identity. Duplicate application nodes, wrong sameAs, invented prices.
  5. Middleware in the way. Locale redirects and route allowlists hiding .txt files and new pages.
  6. Mistaking smoke tests for proof. A user-agent curl is not a crawl. A branded search is not a recommendation.

A pre-release checklist.

Crawl and index

  • Content in the initial HTML?
  • One h1, self canonical, unique title and description?
  • Sitemap only lists live canonical URLs with honest dates?
  • Private routes noindex and out of the sitemap?

Facts and schema

  • Prices, plans and FAQ rendered from one source?
  • Parity test for page, JSON-LD and text copies?
  • Stable @id, valid Offer fields, no invented numbers?
  • sameAs points only to real profiles that link back?

Answers

  • First paragraph answers the page on its own?
  • Product named inside the answers?
  • Limitations stated where the feature is described?
  • Visible review date?

Deploy and measure

  • Short cache on sitemap, robots, llms and feeds?
  • Post-deploy route and 404 checks passing?
  • IndexNow sent for changed URLs?
  • Prompt baseline recorded before the change?

If you only do one thing from this list, make every fact about your product come from one place. The engines will get the rest wrong in more interesting ways, but at least not because you contradicted yourself.