How to Rank on AI Search Engines in 2026: ChatGPT, Perplexity and Gemini

Dark futuristic illustration of an AI search engine core surrounded by glowing data source panels, navy background with electric blue and purple neon accents

Introduction

Every major guide to AI search repeats the same nine or ten tactics: write clear content, add statistics, use schema, stay fresh, get mentioned elsewhere. Almost none of them say which of those tactics is actually backed by a controlled experiment and which one just keeps getting repeated because everyone else repeats it. This guide separates the two, using the real research behind AI search visibility the Princeton and Georgia Tech study that coined the term Generative Engine Optimization, and newer data from Ahrefs’ Brand Radar covering millions of citations across ChatGPT, Perplexity, Gemini and Google’s AI Overviews. Where the evidence is thin, or flatly contradicts popular advice, schema markup is the clearest example that gets flagged directly instead of glossed over, so your effort goes where it’s actually rewarded.

Quick Answer

To rank on AI search engines in 2026, start with content that already ranks well in Google. About three-quarters of AI Overview citations pull from pages in the organic top ten. Beyond that: answer questions directly, cite real sources, and keep pages genuinely updated. Schema markup and llms.txt files, despite the hype, show little proven effect.

Key Takeaways

About three in four AI Overview citations come from pages that already rank in Google’s organic top ten, based on Ahrefs’ analysis of 1.9 million citations so traditional SEO fundamentals still decide most of the outcome.

Adding schema markup to a page AI already cites produced no measurable citation gain on Google AI Mode or ChatGPT, and a small decline on Google AI Overviews, in a controlled Ahrefs test of 1,885 pages because none of the major AI systems read JSON-LD during live retrieval.

The widely quoted “40% visibility boost” from the Princeton GEO study is a ceiling reached by specific technique combinations on a research benchmark, not a guaranteed lift from adding one statistic or citation to an article.

ChatGPT’s citations skew toward content published roughly 458 days more recently than what shows up in Google’s organic results, while Google’s own AI Overviews barely favor fresh content at all so “stay fresh” means different things on different platforms.

Google has stated on the record that it does not crawl or use llms.txt files, and that ordinary SEO practices are what determine AI Overview visibility making llms.txt one of the least evidence-backed items on most 2026 AI SEO checklists.

Whether your content is actually being cited is directly checkable, through Google Search Console’s AI-related reports and third-party visibility trackers, rather than something you have to guess at from ranking position alone.

What “Ranking” Means in AI Search (Not Google’s Ten Blue Links)

AI search engines don’t return a page of ten ranked links. ChatGPT, Perplexity, Gemini and Google’s AI Overviews each generate a written answer and cite a small set of sources alongside it usually three to five for an AI Overview, sometimes more inside a longer ChatGPT response. Being visible here means being chosen as one of those sources for a given prompt, not holding a numbered position.

That distinction changes what “optimizing” even means. A page can rank nowhere near position one in Google and still get cited, because the system is retrieving and scoring individual passages, not whole pages, against a specific question. It can also rank first in Google and never get cited, if a competitor’s page answers the exact sub-question the AI generated internally more directly. The rest of this guide treats “ranking” as shorthand for citation likelihood, since that’s the metric that actually exists in AI search.

How ChatGPT, Perplexity and Gemini Actually Choose What to Cite

Most AI search systems break a single prompt into several narrower “fan-out” queries, retrieve top-ranking pages for each one from an underlying search index, then have a language model synthesize an answer and cite whichever passages answered each sub-query best. This is retrieval-augmented generation, or RAG: the model isn’t answering from memory, it’s answering from what it just fetched.

For a prompt like “best running shoes,” the system might separately search for a specific shoe model’s most recent review, a durability comparison and a price check, then combine the strongest passage from each result into one answer. This is why a single article rarely gets cited for an entire broad topic; it gets cited for the specific sub-question it happens to answer best.

The platforms differ in how they surface this. Perplexity displays numbered citations inline next to each claim, which rewards content with a clear, checkable fact per sentence. Gemini leans on Google’s own search grounding for factual queries. ChatGPT’s search feature runs its own retrieval and browsing layer on top of the base model. Each platform runs its own pipeline on its own schedule, so a change that helps citation in one can do nothing in another treat this as several audiences, not one “AI search” target.

Where Traditional SEO Still Decides Whether You Rank on AI Search Engines

Google’s AI Overviews pull the large majority of their citations from pages that already rank in ordinary Google search 76.1% of cited pages sit in the organic top ten, and the top-cited source in a given AI Overview has a median Google ranking of position 2, according to Ahrefs’ study of 1.9 million citations. Ranking well by classic SEO measures remains the single strongest predictor of AI citation.

The pattern holds as citation position moves down the AI Overview response:

AI Overview citation positionTypical Google ranking (median)
12
24
35

About 14% of AI Overview citations come from pages that don’t rank in Google’s top 100 at all, which is real but the exception rather than the rule. It’s tempting to assume these lower-ranked citations come from content specifically targeting AI’s internal fan-out queries longer, more specific search terms. Ahrefs tested this and found the opposite: pages cited despite ranking outside the top ten actually appeared for fewer keywords and shorter search terms than top-ranking pages, not more. The honest conclusion is that AI citation beyond position ten is driven by a mix of factors nobody has fully isolated yet, including freshness and topical fit not a specific long-tail-content trick you can reliably aim for.

What the Research Actually Shows Moves the Needle

The most cited academic source in this field is the 2024 “GEO: Generative Engine Optimization” paper, from researchers at Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, which ran roughly 10,000 queries through a benchmark of generative search engines and tested nine content-modification techniques. Three findings are worth acting on:

  1. Adding cited sources and specific statistics to a passage improved its visibility more consistently than any other single change the researchers tested though the paper’s own headline figure, a visibility boost of “up to 40%,” describes the strongest technique combinations on its benchmark, not a typical result from any one edit.
  2. Stuffing a passage with keywords measurably hurt visibility in the same tests, a rare case where a classic SEO instinct points the wrong way for AI citation.
  3. Results varied significantly by topic domain, so a technique that helps an explainer-style “how does X work” article may do little for a product comparison or a how-to guide.

Treat the 40% figure as a ceiling reached under favorable conditions in a research setting, not a return you can bank on from any single edit to your own article.

The paper’s own authors are more cautious than the marketing built on top of it. Their limitations section states plainly that they did not test whether these techniques affect a page’s actual search ranking, and they expect that as future language models handle longer context windows, search ranking will matter less overall, since the systems will simply be able to read more sources per query. Even the researchers who coined the term GEO don’t claim it replaces ranking well in the first place the sourcing and structure work below still carries the load.

Content Structure That Gets Extracted

AI systems favor passages that make sense on their own a direct answer in the first sentence of a section, without needing the paragraph before or after it for context. Writing for extraction means naming the subject explicitly instead of leaning on “it” or “this,” and giving every heading a complete, self-contained answer underneath it.

A page explaining how retrieval-augmented generation works, for example, gets lifted more easily if each subsection opens with a plain definition “RAG is a technique where a model retrieves external documents before generating an answer” than one that opens with “Now let’s look at how this works,” which means nothing without the sentence before it.

Four habits do most of the work here:

  1. Phrase headings as the real questions people ask, then answer them in the first two sentences underneath not the fifth.
  2. Name the entity every time, especially in a section’s opening sentence: “ChatGPT’s search feature” beats “it” or “this tool.”
  3. Use a table or numbered list whenever you’re comparing three or more things on the same criteria extraction systems handle structured data more reliably than a dense paragraph.
  4. Put your single strongest, most quotable sentence directly under a heading rather than three paragraphs in if it’s buried, most systems pull a weaker sentence around it instead.

Being right isn’t enough in AI search you also have to be the easiest right answer to lift out of the page. If you’re still fuzzy on how large language models generate an answer in the first place, our explainer on generative AI covers the mechanics this section assumes.

The Schema Markup and llms.txt Myth

Adding schema markup to a page AI systems already cite does not reliably increase citations, and Google has said outright that it doesn’t use llms.txt at all making these two of the most commonly recommended 2026 AI SEO tactics also two of the least supported by evidence.

In a controlled experiment, Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matching each one against thousands of control pages that were already receiving similar levels of AI citation but never added schema. Across four separate statistical tests, the result held: adding schema produced no meaningful citation gain.

PlatformEffect on citations after adding schema
Google AI Mode+2.4% statistically indistinguishable from noise
ChatGPT+2.2% statistically indistinguishable from noise
Google AI Overviews−4.6% small but statistically significant decline

One limitation matters: every page in that study already had over 100 AI Overview citations before schema was added, meaning it only tested pages AI systems had already found. For a page that isn’t being crawled or surfaced by AI at all, schema might still help it get parsed correctly in the first place the study can’t speak to that scenario. A separate test reported in the same Ahrefs report, run by an independent researcher across five major AI systems including ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, found that during live retrieval, every one of them read only the visible HTML on a page and ignored JSON-LD, hidden Microdata and hidden RDFa entirely.

Common mistake: treating schema as an AI-citation strategy on its own. If you’re already doing the rest of the SEO work, strong content, real authority, technical cleanliness, schema markup isn’t the shortcut the AEO industry has been selling. 

The llms.txt file, a plain-text index of a site’s key pages, proposed as an AI-era sitemap has even less backing. At a Google Search Central event in July 2025, Google’s Gary Illyes stated directly that Google does not crawl or use llms.txt, and that ordinary SEO practices are what get a page into AI Overviews. Other AI companies’ crawlers behave inconsistently, some fetch the file occasionally but no major platform has documented using it as a ranking or citation signal.

What does help, on the access side, is making sure robots.txt isn’t accidentally blocking the crawlers that power AI answers. OpenAI runs GPTBot, used for model training, and OAI-SearchBot, used for live search retrieval, as separate, independently controllable user agents you can allow OAI-SearchBot so your pages can appear in ChatGPT’s search results while still disallowing GPTBot if you don’t want your content used for training. Anthropic and Google run comparable crawler pairs. Check these settings before assuming a visibility problem is a content problem.

How Much Content Freshness Really Matters

Freshness helps with AI citation, but by very different amounts depending on the platform and updating a publish date without changing anything else won’t fool any of them. Ahrefs’ analysis of 17 million citations across seven AI search platforms found that AI-cited content is, on average, 25.7% “fresher” than content in organic Google results.

That average hides a wide range. Google’s own AI Overviews show almost no freshness preference; the content they cite is actually 16 days older, on average, than what shows up in ordinary Google results, since they largely inherit Google’s existing ranking signals. ChatGPT sits at the opposite end: its cited sources run 458 days newer than organic search results on average, and its in-text references appear to be deliberately ordered from newest to oldest. Perplexity and Gemini fall in between, each preferring content updated within the last two to three years without ChatGPT’s strong recency bias.

Common mistake: bumping a page’s published or updated date without making a real content change, hoping to game the freshness signal. Google’s own search advocate John Mueller has warned against this specifically a changed date with no substantive edit underneath it is a signal search engines increasingly discount, and it risks looking manipulative if a reader compares versions.

The practical takeaway isn’t to update everything constantly, it’s to prioritize genuine updates on pages targeting prompts likely to be asked through ChatGPT or Perplexity, where freshness carries real weight, over pages primarily competing in Google’s AI Overviews, where it barely moves the needle.

How to Check Your LLM Search Visibility

You can confirm whether AI systems are actually citing your content instead of guessing from ranking position. Google Search Console now reports AI Overview and AI Mode impressions directly, and several third-party tools track citations across ChatGPT, Perplexity and Gemini by running the same prompts repeatedly and logging which sources get mentioned.

Third-party visibility trackers work differently: they run a fixed set of prompts against each AI platform on a schedule and record every domain that gets cited or recommended the only way to see your standing on platforms that don’t run their own version of Search Console.

Before investing heavily in any tactic from this guide, run a baseline check first. Note which of your pages get cited today and for which prompts, make one change, and recheck after a few weeks the same logic Ahrefs used in its schema experiment, just at a smaller, one-site scale.

Common Mistakes That Undermine AI SEO

Beyond the schema and freshness traps already covered, a few other habits quietly work against AI visibility:

  1. Treating AI search as one target instead of several, a tactic that helps ChatGPT citation can do nothing for Google’s AI Overviews, since each platform runs its own retrieval and freshness logic.
  2. Copying a statistic from a competitor’s article instead of tracing it to its original source, which risks repeating an outdated or wrong number and gives AI systems a weaker, secondhand citation to work with.
  3. Optimizing only your own site’s content and ignoring earned coverage elsewhere a September 2025 study from researchers including Georgia Tech and the University of Toronto found AI search systems are systematically biased toward third-party, earned media over brand-owned content, more so than traditional Google results.
  4. Publishing content with no named alternatives or competitors, which reads as promotional and gives AI systems less reason to treat the page as a balanced, citable source rather than marketing copy.
  5. Chasing every new “must-do” AI SEO tactic llms.txt included without checking whether any platform has confirmed using it, instead of spending that time on content depth and genuine authority.

Most of these come down to the same root cause: treating AI visibility as a checklist to complete rather than a byproduct of being the clearest, most trustworthy source on a topic.

Conclusion

The pages that consistently rank on AI search engines in 2026 are, overwhelmingly, the pages that already do classic SEO well clear structure, real sourcing, genuine authority with a handful of AI-specific habits layered on top: answer-first writing, self-contained sections, and honest freshness where the platform actually rewards it. Skip the AI SEO tactics without evidence behind them, especially schema-as-a-silver-bullet and llms.txt, and put that time into the sourcing and structure work the research actually supports.

Good AI SEO right now looks less like a new discipline and more like SEO done properly, for an audience that reads faster and cites more selectively than a human ever did. If you want to go deeper on the broader content and structured-data landscape this article builds on, our 2026 digital marketing trends guide is the natural next read.

FAQs

1. Does ChatGPT use Google’s search index?

ChatGPT’s search feature runs its own retrieval process rather than simply mirroring Google’s rankings, though the two overlap significantly in practice. A majority of AI Overview citations also rank in Google’s organic top ten, and similar overlap patterns show up across other AI platforms, but each system applies its own additional scoring on top of whatever it retrieves.

2. Is GEO (Generative Engine Optimization) different from SEO?

GEO describes adjusting content to be cited by generative AI systems, and it shares most of its foundation with traditional SEO crawlability, authority and clear structure. The differences are narrower than the term suggests: better sourcing and self-contained sections matter more, while some classic tactics like schema markup and keyword-focused writing show weaker or even negative results for AI citation specifically.

3. How long does it take to start showing up in AI search citations?

There’s no fixed timeline, since it depends on how often each platform recrawls a page and whether it already ranks in classic search. Pages that already rank well in Google tend to appear in AI citations faster, sometimes within days of a content update, while new domains with no existing search visibility can take considerably longer to be picked up at all.

4. Do backlinks still matter for ranking on AI search engines?

Yes, backlinks remain one of the signals that help a page rank in classic Google search, and since a large majority of AI citations pull from pages already ranking well in Google, link-building indirectly supports AI visibility too. No verified study yet isolates how much backlinks alone matter to AI citation independent of ranking position.

5. Can a small or new website get cited by AI search?

Yes, more often than in classic Google rankings, because AI systems weigh passage-level relevance and clarity alongside domain authority rather than relying on it as heavily. A smaller site with a clear, well-sourced answer to a specific question can outcite a larger competitor that only covers the topic broadly.

logo-white.png

Subscribe to Our Newsletter