Generative engine optimization for your help center

Generative engine optimization for help centers: write and measure articles so ChatGPT, Perplexity and Google AI features quote them.
TL;DR: Generative engine optimization for a help center means writing each article so an AI system can lift one passage and quote it correctly. Google says AI Overviews and AI Mode need no special markup or AI text files, only indexable pages with snippets. OpenAI and Perplexity document separate search crawlers you should allow. FAQ rich results are gone from Google Search since May 2026, so write visible question-and-answer blocks for readers and retrieval, not for a badge. Measure with Bing's AI Performance report, referral data, and a weekly prompt audit.
I come at this as a help-center editor, the person who has to defend every sentence in a support article because someone will act on it. For years the reader of that article was a customer with a search box. Now a second reader sits between the customer and the page: a system that splits the question into several searches, pulls passages, and writes an answer with a citation.
The customer may never see your page. They see a sentence from it, in someone else's voice, with your policy inside.
That changes the editing job. A paragraph that reads fine inside a long page can become wrong when it is lifted out alone. A date left off a policy can turn a 2023 rule into a 2026 answer.
This guide applies what Google, OpenAI, Perplexity, and Microsoft document about their systems, plus one peer-reviewed study, to the specific shape of a support knowledge base. Where a source is silent, I say so rather than fill the gap with folklore.
It builds on our earlier guides to structuring a knowledge base for AI and to self-service rate. Those cover your own agent. This one covers the public assistants that answer questions about you before the customer ever opens your widget.
What generative engine optimization is, and what the research shows
The term comes from a 2023 paper by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, accepted to KDD 2024. Their arXiv abstract describes generative engine optimization as a paradigm to help content creators improve their visibility in generative engine responses. The authors built GEO-bench, a benchmark of diverse user queries across domains, and report that the methods can boost visibility by up to 40% in generative engine responses, with effectiveness varying by domain.
Two limits matter for a help center. The study measured visibility in generated answers on a benchmark, not support outcomes, and the abstract notes that results differ by domain. Nothing in it tests support documentation specifically.
Treat "up to 40%" as evidence that how content is written affects whether it is cited, not as a forecast for your site.
The practical reading is modest and useful. Wording and structure influence citation. That is good news for an editor, because wording and structure are the parts of a help center you fully control.
You cannot control a model's ranking, but you can control whether your answer is complete, specific, and easy to lift.
Shopify's merchant guidance frames it similarly. In its GEO playbook, Kyle Risley notes that AI tools run query fan-out, splitting one prompt into several search queries, and that strong search rankings improve the odds of being cited. He also suggests asking an LLM what it says about your brand and then correcting or supplementing the pages it cites.
That second tip is the seed of the audit routine later in this guide.
What the platforms actually document
Much GEO advice online is a mix of documented behavior and guesses. The safest approach is to sort claims by who said them. The table below lists only statements from the platform owners' own documentation.
| Platform | What it documents | What it does not say |
|---|---|---|
| Google AI Overviews and AI Mode | Pages must be indexed and snippet-eligible; no additional requirements; no special schema.org markup or AI text files needed | ✗ Any guarantee of inclusion |
| OpenAI ChatGPT search | Allow OAI-SearchBot in robots.txt to appear in search answers; changes take about 24 hours | ✗ How passages are ranked |
| OpenAI GPTBot | Used for foundation model training; disallowing it signals content should not be used for training | ✗ Any effect on search appearance |
| Perplexity | PerplexityBot surfaces and links sites in results; Perplexity-User fetches pages for user questions and generally ignores robots.txt | ✗ Ranking factors |
| Bing and Copilot | AI Performance report shows citations, average cited pages, and grounding queries | ✓ Citation counts, ✗ clicks and rankings |
Google's AI features documentation is the clearest. It states there are no additional requirements to appear in AI Overviews or AI Mode, no other special optimizations necessary, and no need to create new machine readable files, AI text files, or markup. It adds that there is no special schema.org structured data you need.
It also says both features may use "query fan-out," issuing multiple related searches across subtopics and data sources.
Read together, those lines tell an editor where to spend effort. Do not buy a schema plugin for AI search. Do not build a parallel set of AI-only files for Google.
Make the normal page good, make sure it can be indexed with a snippet, and cover the subtopics a fan-out would search for.
Crawler access: the part that is binary
Access is the one area where a mistake removes you entirely. OpenAI's bot documentation says: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." It recommends allowing OAI-SearchBot in robots.txt and notes it can take about 24 hours from a robots.txt update for systems to adjust.
The same page separates three bots by purpose. GPTBot is for training generative models. ChatGPT-User handles actions a person initiates inside ChatGPT, and because a user starts them, robots.txt rules may not apply.
OpenAI says ChatGPT-User is not used to decide whether content appears in Search. So a company that blocks GPTBot to stay out of training can still allow OAI-SearchBot to stay visible in answers.
Perplexity draws a similar line in its bot guide. PerplexityBot surfaces and links websites in search results and is not used to crawl content for foundation models. Perplexity-User supports user actions, and since a user requested the fetch, it generally ignores robots.txt rules.
For a help center, the decision is usually simple. Support content exists to be found and used. Unless your legal team has a specific reason, allow the search bots, and decide separately about training bots.
Then verify, because many help centers inherit a restrictive robots.txt from a platform default or a security vendor's rule.
- Open yourdomain.com/robots.txt and read it line by line for blanket Disallow rules that cover /help or /support.
- Check whether a CDN or bot-protection service blocks unknown user agents before the request reaches your server.
- Confirm the help center subdomain has its own robots.txt, since each host serves its own file.
- Publish an XML sitemap with accurate last-modified dates for every article.
- Test that the article text appears in the raw HTML response, not only after scripts run.
I could not find OpenAI, Perplexity, or Google documentation that states precisely which of their fetchers execute JavaScript. In the absence of a guarantee, server-rendered text is the conservative choice. Many hosted help-center products render article bodies on the server already, but check yours with a plain curl request.
Query fan-out and the shape of a good help article
Fan-out changes what "covering a topic" means. When a customer asks an assistant, "Can I get a refund if I cancel mid-month?", the system may search separately for the refund policy, the cancellation steps, proration rules, and the billing cycle. Your answer might be spread across four pages, and the assistant will stitch together whichever passages it finds.
That is a problem when the four pages disagree or when one is missing. The assistant fills the gap from another source, which may be a forum post or an outdated review. Your best defense is a set of articles where each subtopic has one authoritative page that states the rule plainly.
Here is the editing pattern I use for a policy or how-to article. It is a house style, not a platform requirement, but each rule maps to a failure I can name.
- One intent per article. A page titled "Refunds" that also covers invoices and tax gets quoted wrongly.
- Put the answer in the first sentence under the heading, with the condition inside the sentence: "You can get a prorated refund within 14 days of the charge."
- Make every paragraph survive being read alone. Replace "this" and "it" with the noun.
- Include the numbers: days, dollar amounts, limits, plan names. Vague policies produce confident wrong paraphrases.
- State exceptions in the same section as the rule, not three scrolls down.
- Name the date of the last review and the plan or region the rule applies to.
These rules also help the customer who reads the page directly. A passage that works alone is a passage that scans well. The two audiences want the same thing, which is why GEO for a help center is mostly good editing.
Before and after: rewriting one article
Consider a common refund page opening. The before version reads: "We want you to be happy with our service. If you are not satisfied, please contact us and we will look into your situation.
Terms may vary." Lifted out, the passage contains no rule, so an assistant will invent one or cite someone else's.
The after version reads: "Annual plans can be refunded in full within 14 days of purchase. After 14 days, we do not refund annual plans, but you can cancel to stop the next renewal. Monthly plans are not refunded for partial months.
Last reviewed October 2026." Every sentence is a fact an assistant can quote and a customer can check.
Notice the work the second version does. It states the condition, the time limit, the exception, and the date. It avoids phrases like "terms may vary," which tell a model nothing and tell a customer to open a ticket.
Writing a version like this takes a few minutes, but only if the policy itself is decided. Often the real finding is that nobody has written the rule down.
That finding connects to a theme in our support knowledge gap analysis guide: the questions your team answers repeatedly in tickets are the questions your articles do not answer clearly. Start your rewrite queue from the top of that list.
FAQ blocks and schema in 2026
FAQ sections are still worth writing, but the reason changed. Google's structured data documentation records that the FAQ rich result stopped appearing in Google Search as of May 7, 2026, with related Search Console tooling removed afterward. Before that, since 2023, it had been limited to well-known government and health sites.
See Google's FAQ structured data page and its changelog for the authoritative dates.
So FAQPage markup no longer earns a visible badge, and Google's AI features documentation says no special schema is needed to appear there. Does that make the markup harmful? I found no source saying so, and Google has said unused structured data does not cause problems for Search.
It is simply not a lever.
What remains valuable is the visible content: a question phrased the way customers phrase it, followed by a short, complete answer. That format matches how assistants retrieve and quote. If you keep FAQPage markup because your platform generates it, make sure it matches the visible text exactly, since mismatches between markup and page content have always been a quality risk.
| Element | Worth doing in 2026 | Reason |
|---|---|---|
| Visible question-and-answer blocks | ✓ | Matches how customers ask and how passages are lifted |
| FAQPage JSON-LD for a rich result badge | ✗ | FAQ rich results no longer appear in Google Search |
| Special AI markup or AI-only text files for Google | ✗ | Google says they are unnecessary |
| Accurate Article markup with dates | ✓ | Helps freshness signals and clarity, low cost |
| Clear headings in sentence form | ✓ | Gives each passage a label that survives extraction |
Freshness and consistency across pages
An assistant that quotes your help center quotes whatever version it last saw. If the pricing page says one thing and an old help article says another, the answer depends on which passage the system retrieves. You will not know which until a customer repeats it to your support team.
The fix is dull and effective. Give every article an owner and a review date, show the date on the page, and update the date only when someone has actually read the article against the product. A "last reviewed" stamp that changes automatically on every save teaches nobody anything and can mislead a model about how current the content is.
- Keep one source of truth per fact. Prices, limits, and trial lengths live in one place, and every other page links to it instead of restating it.
- Search your help center for each plan name, price, and limit once a quarter and confirm every mention matches.
- When a product change ships, list the articles it touches in the release checklist, not after customers complain.
- Write dates into time-bound statements: "As of October 2026, the export limit is 10,000 rows."
- Redirect merged or renamed articles with a 301 so old links and old citations still land on the current answer.
This is also where an internal agent benefits. If you train a support agent on the same help center, it inherits the same inconsistencies. Our guide to training AI on your help center covers cleaning the source before you index it, and the RAG for customer support explainer shows why retrieval returns conflicting passages when sources disagree.
Measurement without fantasy numbers
There is no Search Console report that says "you were quoted by ChatGPT 412 times." Be suspicious of any dashboard that claims precision here. What exists is a small set of real signals, each partial, and a manual routine to fill the gaps.
Microsoft ships the only first-party report I know of. Bing Webmaster Tools added an AI Performance report in public preview in February 2026. It shows how often your site is cited in Copilot and partner AI answers, the average number of cited pages, and the grounding queries that led to citations.
It does not report clicks or rankings, and the preview status means the feature may change.
| Signal | Where to get it | What it tells you | Limit |
|---|---|---|---|
| AI citations and grounding queries | Bing Webmaster Tools AI Performance | ✓ Which pages and queries earn citations | ✗ Covers Microsoft surfaces only; no click data |
| AI referral sessions | Web analytics, filtered by referrer domain | ✓ Visits from chat.openai.com, perplexity.ai, and similar | ✗ Misses answers read without a click |
| Prompt audit | Manual weekly run of 20 to 30 real customer questions | ✓ What assistants say about your policies | ✗ Answers vary by run, account, and location |
| Self-service rate | Your help-center and chat analytics | ✓ Whether customers resolve issues without a ticket | ✗ Does not isolate AI-search effects |
| Ticket topics | Helpdesk tags | ✓ New wrong-policy questions that quote an assistant | ✗ Depends on agents tagging consistently |
The prompt audit is the most useful and the cheapest. Take your 25 most common customer questions, from ticket tags rather than imagination. Ask each one in ChatGPT, Perplexity, and Google, in a fresh session.
Record whether the answer is correct, whether your site is cited, and whether any detail is wrong. Run it weekly for a month and the pattern will show which articles need work.
Treat the results as a sample. Assistants return different answers across runs, and I would not draw conclusions from one week. What matters is repeated errors on the same topic, because that points to a missing or unclear article rather than model noise.
For the human-facing side, track self-service rate and your analytics for questions the widget could not answer. Those are the same gaps the public assistants will fill badly.
Retire, merge, or rewrite stale articles
Most help centers carry dead weight. Articles for features that no longer exist, three versions of the same setup guide, and screenshots of an interface from two redesigns ago. A human can usually tell.
A retrieval system sees three passages that all look relevant and picks one.
Run a simple triage. List every article with its last review date, its traffic, and the ticket volume on its topic. Then sort each into one of four bins.
- Keep and refresh. The topic is current and customers need it. Update the facts and the review date.
- Merge. Two articles answer the same question. Combine into the clearer one and redirect the other.
- Split. One article answers four questions. Break it into four pages with one intent each.
- Retire. The feature is gone or the policy changed. Redirect to the closest current page, or return a 410 if nothing replaces it.
Retiring is the step teams skip, and it often has the largest effect on wrong answers. An outdated page is not neutral. It is a confident source of misinformation with your domain name on it.
Where Communicate fits
I will keep this part short and honest. Public AI search decides what a stranger hears about you. Your own support agent decides what a customer hears once they are on your site.
Both draw on the same material, so the work you do on articles pays twice.
Communicate lets you connect a help center, a website, or documents as data sources so your AI agent answers from the articles you maintain. When the agent cannot answer, the conversation lands in a shared inbox for your team, and the unanswered question tells you which article to write next. We do not control how ChatGPT or Google cite you, and we make no claim that our product changes your ranking in them.
If you want agents outside your site to read your knowledge base directly, two related pieces are worth reading. One covers help center MCP servers. The other covers making a product readable to agents with llms.txt.
Note that Google states llms.txt is not needed for its AI features, so treat it as optional elsewhere.
A 30-day plan for a small team
You do not need a program. Here is a sequence one editor can run alongside normal work, with each week producing something checkable.
- Week 1, access. Audit robots.txt, CDN rules, and sitemap. Allow OAI-SearchBot and PerplexityBot unless there is a documented reason not to. Confirm article text appears in raw HTML. Verify Bing Webmaster Tools and request access to the AI Performance preview.
- Week 2, baseline. Pull your 25 top ticket topics. Run the prompt audit in three assistants and record every wrong or missing answer. Capture self-service rate and AI referral sessions as starting numbers.
- Week 3, rewrite. Take the five topics with the worst audit results. Rewrite each to one intent, answer first, conditions inline, dated. Merge duplicates and redirect them.
- Week 4, repeat and decide. Rerun the audit. Compare against week 2 with the caveat that answers vary. Keep what visibly improved, and schedule the next ten articles.
Expect slow movement. OpenAI notes a roughly 24 hour delay for robots.txt changes, and no platform publishes how quickly new or updated pages enter their retrieval. A month is enough to see whether corrected articles are being quoted, not enough to claim a trend.
Writing patterns that survive extraction
The rewrite example earlier showed one policy. The same habits apply to every article type in a help center, and each type fails in its own way when a passage is lifted out. Here is how I edit the four types that cover most of a knowledge base.
Policy articles fail through omission. A refund, cancellation, or data-retention page that states a rule without its conditions produces a confident but wrong paraphrase. Write the rule, the window, the exceptions, and the effective date in adjacent sentences.
If the rule differs by plan or region, name each case in its own sentence rather than a vague "may vary."
How-to articles fail through missing context. Step 4 of a setup guide means nothing alone. Open with one sentence that names the goal and the prerequisite, number the steps, and put the expected result after the last step.
When an assistant quotes steps 3 to 5, the customer at least sees the goal in the surrounding answer.
Troubleshooting articles fail through ambiguity about the symptom. Title the page with the exact error text or the visible symptom, since that is how customers phrase it. Then list causes from most to least common, each with its fix in the same bullet.
Avoid the pattern where a fix sits three screens below the cause it addresses.
Comparison and plan articles fail through stale numbers. Pricing, limits, and feature availability belong on one canonical page. Elsewhere, link to it with a short summary and a date.
Duplicated numbers drift, and drift is how an assistant ends up quoting last year's limit with this year's confidence.
A short list of structural habits covers most of the gain. Use a descriptive heading for each section in sentence case, so the heading itself says what the passage contains. Keep tables for genuine comparisons and give them a text sentence that states the takeaway, because a table may be flattened when extracted.
Spell out acronyms on first use. Write alt text that describes what the screenshot shows For the wider structure of a knowledge base, our piece on knowledge base structure for AI goes deeper on hierarchy, naming, and article size.
Reviewing with real tickets
A rewrite is only an improvement if it answers the question customers actually ask. The cheapest test is a ticket replay. Pull ten recent tickets on one topic, strip names, and read each customer message to the article as if you were the retrieval system.
Does a single passage answer it? If you need two paragraphs from different sections, add a section that does.
Ticket replay also exposes vocabulary gaps. Customers say "cancel my plan" while the article says "terminate subscription." Fan-out searches use the customer's phrasing and its variants, so match it. Add the customer's words as headings or as the first sentence, and keep your internal terms in the body.
If you track unanswered chat questions, you already have this list. A support knowledge gap analysis turns those misses into a ranked writing queue. The analytics view in Communicate surfaces questions the agent could not resolve, which is a practical place to start choosing which article to fix first.
Finish with a quick fidelity check before publishing. Ask a colleague who did not write the article to answer three customer questions using only the edited page. If they hesitate or contradict each other, an assistant will too.
This costs fifteen minutes and catches the most embarrassing errors before a model repeats them.
Teams that want a fuller quality loop can pair this with review of AI answers themselves. Our guide on reducing AI hallucinations in support describes how grounding and source hygiene limit invented answers, and it applies equally to how you edit the source material.
Common mistakes to avoid
- Buying a GEO tool before fixing the articles. A dashboard cannot repair a vague refund policy.
- Blocking all AI bots with one rule, then wondering why you are missing from answers.
- Stuffing keywords or adding promotional language to chase citations. Support content gets lifted into answers, and puffery reads badly there.
- Hiding answers in images, PDFs, or tabbed widgets that do not render as plain text.
- Writing for the assistant and forgetting the customer who lands on the page. If a passage is confusing out of context, it is confusing in context too.
Conclusion
Help-center GEO is mostly the editing discipline good support teams already respect, applied with the knowledge that every paragraph may travel alone. Keep crawlers allowed, make each article answer one question in its first sentence, put numbers and dates in the text, retire what is stale, and measure with the modest tools that exist. If you want your own site to give those same corrected answers instantly, see how Communicate pricing works and connect your help center as a data source.
Frequently asked questions
Is GEO different from SEO for a help center?
They overlap heavily. Google says AI Overviews and AI Mode use the same indexing and snippet eligibility as Search, so classic technical SEO still gates visibility. GEO adds attention to passage-level clarity, since an assistant may quote a single paragraph without the rest of the page.
Do I need special schema markup to appear in AI Overviews?
No. Google's AI features documentation says no special schema.org structured data is required and no new machine readable files or AI text files are needed. Standard Article markup with accurate dates is still reasonable hygiene but is not a requirement.
Should I add FAQPage structured data?
It will not earn a rich result in Google Search, because that result stopped appearing as of May 7, 2026. If your platform already outputs it and it matches the visible text, there is no known harm. Spend your effort on the visible question-and-answer content instead.
Which crawlers should I allow?
For visibility in answers, allow OAI-SearchBot and PerplexityBot, and keep normal Googlebot access. Training bots such as GPTBot are a separate decision. Check your legal and content policies, then confirm your CDN is not blocking these user agents upstream of robots.txt.
What is the difference between GPTBot and OAI-SearchBot?
OpenAI documents GPTBot as the crawler for training foundation models and OAI-SearchBot as the one that powers ChatGPT search results. Opting out of OAI-SearchBot removes you from ChatGPT search answers. Opting out of GPTBot signals you do not want content used for training.
Does blocking Perplexity-User or ChatGPT-User work?
Not reliably. Both vendors describe these as user-initiated fetches, and Perplexity says Perplexity-User generally ignores robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User.
If content must stay private, use authentication rather than robots.txt.
How long until my changes show up?
OpenAI says robots.txt updates can take about 24 hours to take effect for its search system. No platform publishes how fast new content enters retrieval. Plan on weeks for editorial changes and judge results from repeated audits, not single checks.
What is query fan-out and why does it matter to me?
Google describes it as issuing multiple related searches across subtopics for a single question. For a help center it means one customer question can pull passages from several articles. Each subtopic therefore needs one clear authoritative page, and those pages must agree.
How long should a help article be for AI retrieval?
No source gives a magic length. Aim for one intent per page, an answer in the first sentence under each heading, and paragraphs that stand alone. Length follows from the topic.
Padding with generic introductions only dilutes the passage a system might lift.
Does the Aggarwal GEO paper prove these tactics work for support content?
No. The paper reports visibility gains of up to 40% on its GEO-bench queries and notes results vary by domain. It did not test support documentation or measure customer outcomes.
Use it as evidence that wording affects citation, then test on your own topics.
Can I measure how often ChatGPT cites my help center?
Not directly with a first-party report. You can track referral sessions from AI domains in analytics and run a regular prompt audit. Bing Webmaster Tools offers an AI Performance report in public preview for Microsoft surfaces, which shows citations and grounding queries but not clicks.
What is a prompt audit?
A manual routine. Take your most common real customer questions from tickets, ask each in the assistants customers use, and record accuracy and whether you are cited. Repeat weekly.
Answers vary by run, so look for repeated errors on the same topic rather than one-off misses.
Should I publish a separate page just for AI?
No. Google says special AI files and parallel markup are not needed, and a hidden or duplicate page invites inconsistency. Make the one public article clear.
If an assistant can quote it accurately, a customer reading it will also understand it.
How do dates help?
A date lets a reader and a retrieval system judge whether a rule is current. Put time-bound statements in the text, such as "As of October 2026," and show a genuine last-reviewed date. Do not auto-update the stamp on every save, since that hides real staleness.
What should I do with outdated articles?
Sort each into keep, merge, split, or retire. Redirect merged and retired pages with a 301 to the closest current answer, or return 410 when nothing replaces them. Leaving stale pages live gives retrieval systems conflicting passages to choose from.
Is JavaScript-rendered help content a problem?
I found no vendor statement that guarantees their fetchers run JavaScript, so server-rendered text is the conservative choice. Fetch a page with curl and check that the article body is present in the raw HTML. If it is missing, ask your help-center vendor about server rendering.
Does llms.txt help?
Google says no AI text files are needed for its AI features, and other vendors have not committed to using it. It is optional, cheap, and harmless. Do not treat it as a substitute for clear articles or open crawler access.
How does this relate to my own AI support agent?
They share a source. A support agent that answers from your help center repeats whatever your articles say, including contradictions. Fixing the articles improves both public AI answers and your own agent.
Gaps your agent cannot answer show you what to write next.
Who should own GEO work on the team?
The help-center editor or support content owner, with a named reviewer from product for factual accuracy. It is an editing and maintenance job, not a one-time project. Give it a weekly slot for the prompt audit and a quarterly slot for the retire-merge triage.
What is the fastest first step?
Open your robots.txt and confirm you are not blocking OAI-SearchBot or PerplexityBot by accident. Then rewrite your single most-asked policy article so the first sentence states the rule with its conditions and a date. Both take under an hour.