What Is Generative Engine Optimization (GEO)? How to Become Visible in AI Answers

Illustration: a man stands in front of a generated AI answer; below it sit three source cards, one of them highlighted in blue and connected to the answer.
Key Takeaways:

Generative Engine Optimization (GEO) is work on content aimed at getting cited inside the answers of generative AI systems. Those systems are ChatGPT, Perplexity, Google Gemini and the AI Overviews in Google Search.

  • Where the term comes from: from the research paper “GEO: Generative Engine Optimization” by Aggarwal et al. (arXiv:2311.09735), submitted in November 2023 and accepted to KDD 2024. It comes neither from Google nor from the consulting business.
  • What was measured: the paper reports visibility gains of up to 40 % inside generated answers. That was measured on the authors’ own benchmark GEO-bench with 10,000 queries, in a self-built setup using gpt3.5-turbo and additionally on Perplexity.ai. Google’s AI Overviews were not a test environment.
  • What worked best: verbatim quotations, citations and statistics. Keyword stuffing delivered “little to no improvement” according to the paper and landed 10 % below the untouched version on Perplexity.ai.
  • What Google says about it: the Search Central documentation names GEO and AEO explicitly and files them under SEO: “optimizing for generative AI search is optimizing for the search experience, and thus still SEO”.
  • My reading of both sources: they describe two different steps. Retrieval decides which pages become candidates for an answer and wording decides who gets quoted from them. Even the paper’s own setup pulled its sources from Google Search.
  • What you actually do: source it, quote it, quantify it, write it readably. Those are the methods with the best numbers in the test. No dedicated tool needed and no conflict with Google’s documentation.

Generative Engine Optimization (GEO) describes work on content with the aim of appearing inside the answers of generative AI systems. That means systems such as ChatGPT, Perplexity, Google Gemini or the AI Overviews in Google Search. They answer a question with a finished text instead of a result list and a few sources get quoted inside it while many others never show up at all.

So much for the common definition. The more interesting question is what it rests on.

I looked at the German search results for “generative engine optimization” on 16 August 2026. Among the first fifteen results sit eight guides from agencies and tool vendors. Between them sit the two sources the whole topic feeds on: in position 6 the research paper that invented the term, and in position 9 Google’s own documentation, which contradicts it rather directly.

Both are freely available and neither one says very much without the other. That is why they stand side by side in this article.

The route there: first the definition and the origin, then the evidence together with its limits, then the implementation. Anyone who only needs the recommendations can jump straight there through the table of contents.

What is generative engine optimization?

Key Takeaway: GEO is an optimisation framework from a research paper by six authors, submitted in November 2023 and accepted to KDD 2024. It originates neither with Google nor with the consulting industry.

The source is a paper with the preprint number arXiv:2311.09735 titled “GEO: Generative Engine Optimization”. The authors are Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande. Version 1 was submitted on 16 November 2023 and version 3 revised on 28 June 2024. The note on the arXiv page reads “Accepted to KDD 2024”, meaning the 30th ACM SIGKDD conference held in Barcelona from 25 to 29 August 2024.

The paper introduces an umbrella term first: systems such as BingChat, Google’s SGE at the time and Perplexity.ai are called “generative engines” because they search for information and generate an answer from it. The step after that is the actual contribution: a framework website owners can use to influence their visibility inside such answers.

The reasoning in the abstract is remarkably non-technical. It is about the third party in the game: “it poses a huge challenge for the third stakeholder”, and the original names websites and content creators explicitly. Anyone producing content loses control over whether and how they appear in a generated answer. The authors propose a way out of that, which they call “the first general creator-centric framework”.

Where AIO, AEO and LLMO fit into that picture. GEO targets generative answers and AEO targets direct answer engines. LLMO refers to how language models understand content. AIO is the umbrella above all of them. The boundaries are soft and interchangeable across many texts. I have compared AIO, GEO, AEO and LLMO in a separate overview. Only one point matters here: out of the four abbreviations I could find a named primary source with disclosed methodology for exactly one.

That is the origin of GEO: no product and no Google announcement, but an academic proposal with a benchmark next to it.

Why the term is suddenly everywhere

Key Takeaway: The boom around the term has two verifiable triggers. An answer surface now absorbs clicks and in May 2026 Google itself documented that it treats this surface as part of Search.

The paper appeared in November 2023. That GEO is now sold as a discipline of its own therefore has less to do with the paper than with the surface search results land on these days.

A generated answer settles the question on the results page already and turns the click on the source into an optional extra. This pattern is older than AI search and goes by the name zero-click search. Generated answers give it a new order of magnitude: whoever supplies the answer and still goes unnamed loses the visibility entirely rather than just a rank.

The second trigger has a date. On 15 May 2026 Search Engine Land reported on a new documentation page from Google about optimising for generative AI features. That made the surface official. A documented selection process creates a market for advice about it.

Part of the same development is the report Google now offers for it and which the documentation explicitly points to for measurement: the Generative AI performance report in Search Console. A platform does not build a report for a surface it considers unimportant.

What those two triggers do not supply is a documented statement about effect: the importance of a topic says nothing about the effect of any single measure. This is exactly the point where it gets tempting to jump straight from market observation into a list of measures. The following sections take the detour through the evidence instead.

GEO and SEO: what differs and what does not

Key Takeaway: The two terms differ in what they measure rather than in what they demand. SEO measures positions in a result list and GEO measures shares of a generated text.

The common comparison runs: SEO brings rankings and GEO brings mentions. That is not wrong, it merely blurs the actual difference and the comparison only gets useful once you lay the two evidence bases side by side.

 Classic SEOGEO according to the paper
Target surfaceposition in the result listshare and position inside a generated answer
What gets optimisedpage, structure, linking, relevance to the querywording of a page that has already been retrieved
Metricrank, clicks, impressionsPosition-Adjusted Word Count, Subjective Impression
Who defines the rulesthe search engine, documentednobody bindingly so far
Evidence basedecades of practice, vendor documentationone study with disclosed methodology, two engines tested

Table: own comparison based on the paper “GEO: Generative Engine Optimization” (Aggarwal et al., v3 of 28 June 2024) and the Google Search Central documentation, status 10 July 2026.

The second row is the one that counts. GEO in the sense of the paper starts at a page that has already been retrieved. That is the test setup itself rather than a side condition, as the section after next shows in detail: for a page outside the candidate set even the best wording changes nothing.

This is why the formula “GEO replaces SEO” does not hold. It assumes that generative systems find their sources differently from search engines. At Google that is explicitly not the case according to the AI Optimization Guide. For the other vendors comparable documentation is missing and the claim there stays neither proven nor disproven.

What the GEO paper actually tested

Key Takeaway: Measurement happened in a self-built setup: five sources per query taken from Google Search, answer generated by gpt3.5-turbo. Perplexity.ai was added as a second and commercially running engine. Google AI Overviews and Bing Chat were not part of the measurement.

The 40 % are quickly quoted. The setup behind them needs a few more paragraphs. So let me take this in order.

The benchmark. Lacking existing data, the authors build their own test set: “GEO-bench, a benchmark consisting of 10K queries from multiple sources”. Those 10,000 queries come from nine different sources. According to the paper they cover 25 domains from arts through health to games. Different difficulty levels and search intents are layered on top.

The engine. This is where it gets important for any interpretation, because per query the setup fetches the five best sources: “only the top 5 sources are fetched from the Google search engine for every query”. A language model then produces the answer: “The answer is then generated by the gpt3.5-turbo model”. To dampen random effects the authors sample five answers per query at a temperature of 0.7.

The second engine. The same test additionally runs against a real product: “we evaluate the same Generative Engine Optimization methods on Perplexity.ai, which is a commercially deployed generative engine”. The limitations section has the decisive number: “we rigorously test our proposed methods on two generative engines, including a publicly available one”. Two engines, then. Google’s AI Overviews appear in the paper as an example of generative systems rather than as a test environment.

The metrics. Two figures score the result: Position-Adjusted Word Count counts the words of an answer that stem from a particular source. Weighting follows position, so whatever appears at the top counts more. Subjective Impression bundles several subjective factors into one value and is assigned by a language model. Those factors include relevance and influence of the cited passage as well as its uniqueness. Both measure visibility inside a generated answer rather than rankings or clicks.

The paper draws that line cleanly itself: “owing to the black-box nature of search engine algorithms, we didn’t evaluate how GEO methods affect search rankings”.

Careful: The 40 % are the authors’ own measurement, on their own benchmark, with their own metric, in their own setup plus Perplexity.ai. It is not a measurement against Google’s AI Overviews, Gemini or ChatGPT. Passing the number on without that boundary passes on something other than what was measured.

Then there is the age of the experiment, which ran on gpt3.5-turbo and therefore on the state of 2023. Today’s models are different ones and so are the retrieval layers above them. That does not devalue the results. It limits them to a careful measurement of a system from back then.

The nine methods and what they delivered

Key Takeaway: Quotations, citations and statistics performed best. Keyword stuffing dropped below the untouched original. That ranking matches what solid editorial work demands anyway.

Nine variants were tested: “We propose 9 different Generative Engine Optimization methods to optimize website content for generative engines.” Each one alters the same source text in a different way. The results sit in table 1 of the paper, in absolute values. I have reproduced them here because the absolute values are the most interesting part of the whole paper and get lost as soon as everything is boiled down to a single percentage.

MethodWhat was changed in the textPosition-Adjusted Word CountSubjective Impression
No optimisation (baseline)original text, untouched19.319.3
Keyword Stuffingadd more keywords from the query17.720.2
Unique Wordsuse rarer, more unusual words20.520.4
Authoritativemake the tone more persuasive and authoritative21.322.9
Easy-to-Understandsimplify the language22.020.5
Technical Termsadd technical terminology22.721.4
Cite Sourcesadd citations24.621.9
Fluency Optimizationimprove flow and phrasing24.721.9
Statistics Additionreplace qualitative statements with numbers25.223.7
Quotation Additionadd verbatim quotations from relevant sources27.224.7

Table: absolute values from table 1 of the paper “GEO: Generative Engine Optimization” (Aggarwal et al., GEO-bench test split, averaged over five runs, version v3 of 28 June 2024). Higher is better. Observed in the authors’ test setup, not general proof.

The paper summarises the leading group like this: “our top-performing methods, Cite Sources, Quotation Addition, and Statistics Addition, achieved a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric”. The best value sits in the table caption: “The best methods improve upon baseline by 41% and 28%”. That is the Quotation Addition row.

I find the bottom end more interesting. Keyword stuffing is arguably the best known tactic from the early days of search engine optimisation. In this visibility metric it drops below the untouched original text. The paper phrases it carefully: “we find such methods offer little to no improvement on generative engine’s responses”. On Perplexity.ai it gets clearer, where it “performs 10% worse than the baseline”.

The second engine confirms the pattern in smaller numbers. On Perplexity.ai, Quotation Addition leads “with a 22% improvement over the baseline”. Cite Sources and Statistics Addition reach “improvements of up to 9% and 37% on the two metrics” according to the paper.

One detail regularly gets lost in the shortened versions: the effect depends on the topic. The abstract says “the efficacy of these strategies varies across domains”. The results section makes that concrete: domains such as “Law & Government” and question types such as “Opinion” benefit particularly from numbers in the text according to the paper. The authors conclude from this that domain-specific optimisation is needed.

What Google itself says about GEO

Key Takeaway: Google’s documentation names GEO and AEO explicitly and files them under SEO. The reason given is that the AI features sit on the same ranking systems as ordinary Search.

The AI Optimization Guide in the Search Central documentation now has the status of 10 July 2026. When it appeared I covered in detail why Google files GEO and AEO under SEO. The core sentence is short:

“From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”

The reasoning sits one paragraph above. According to Google the AI features run on top of the existing systems: “our generative AI features on Google Search are rooted in our core Search ranking and quality systems”. Two techniques are named explicitly. Retrieval-augmented generation is also called grounding there and pulls current pages from the search index to support the answer. How that anchoring works I took apart in a separate piece on grounding. Query fan-out generates several related queries for one request at the same time to gather more material. The example in the documentation: “how to fix a lawn that’s full of weeds” turns into “best herbicides for lawns” and “remove weeds without chemicals” among others. What that means for the way AI Overviews work is already covered on this blog too.

The documentation gets more explicit in a section headed “Mythbusting generative AI search: what you don’t need to do”. It lists what you can ignore for Google Search:

  • llms.txt and similar files: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search”. Google Search simply does not use them. What the format is about and who else reads it sits in my llms.txt guide.
  • Content chunking: “There’s no requirement to break your content into tiny pieces for AI to better understand it.”
  • Rewriting for AI: “You don’t need to write in a specific way just for generative AI search.”
  • Bought mentions: “seeking inauthentic ‘mentions’ across the web isn’t as helpful as it might seem”.
  • Structured data as an AI lever: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.” The same passage keeps it useful for rich results, which I work through in my piece on structured data and AI Overviews.

On the terms themselves the documentation gets unusually direct. It writes that many of the circulating tricks “aren’t effective or supported by how Google Search actually works”. Then it adds a pointer: “If you’re considering third-party ‘AEO’ or ‘GEO’ advice or services, review our guidance on evaluating third-party SEO advice.”

That is a statement about Google Search and says nothing about how ChatGPT, Perplexity or Claude pick their sources. The distinction matters because both camps like to skip it.

Retrieval and generation: why both sources are right

Key Takeaway: The paper measures the generation step, Google’s documentation describes the retrieval step. Whatever is not retrieved cannot be quoted. This is my reading of both sources rather than a statement from either.

A generative answer comes about in two steps: first the system looks for sources and afterwards a language model formulates an answer from them, deciding which passages it takes over and which source it names.

Google’s documentation talks about the first step. Retrieval-augmented generation and query fan-out describe how pages enter the candidate set. The paper talks about the second. All nine methods alter the text of a page that has already been retrieved.

The evidence for that split sits inside the paper’s own test setup. Its sources come from Google Search and specifically five per query. Whatever is missing from those five takes no part in the measurement at all. Put differently: the experiment presupposes classic ranking instead of replacing it.

That dissolves the two positions. “Up to 40 % more visibility” does not mean “40 % more visibility than without SEO”. It means that among the pages retrieved anyway, the citation shifts. And Google’s “it’s still SEO” does not mean wording is irrelevant. It means Google recognises no second rulebook for it.

An order of work follows from that in practice: first get into the candidate set, then work on the place inside the answer. Anyone tackling the second stage without holding the first is optimising a text that no system retrieves. I took that apart elsewhere: how source selection differs between AI Overviews and AI Mode.

Putting GEO to work: the four steps with evidence

Key Takeaway: What worked in the paper is editorial work: source it, quote it, quantify it, write it readably. None of that needs a dedicated tool and none of it contradicts Google’s documentation.

Look at the ranking in the table above once more. At the top sit verbatim quotations, numbers and citations, below them flow and readability and at the very bottom the tactic that stuffs the text with keywords.

That is not an AI discipline. That is the description of a well sourced specialist article.

Checklist: four things that measurably worked in the paper and that you can implement without a new tool.
  • Quantify instead of asserting. Wherever a qualitative statement sits (“considerably faster”), a number with a source belongs there.
  • Quote verbatim. The strongest single method in the test was adding real quotations from relevant sources.
  • Make sources visible. Do not just name them, link them exactly where the statement sits.
  • Stay readable. Fluency and comprehensibility landed clearly above the baseline without any change to the substance.

The order in which to tackle this. The following sequence is my rule of thumb rather than a requirement from either source. It follows the split from the previous section: get retrieved first, get quoted second.

  1. Check indexability and ranking. Pages that do not sit in the top results for the question appear in no candidate set at all. That is classic work and it comes first.
  2. Actually answer the question in the text. An answer that sits in the text can be taken over. One that is merely implied cannot.
  3. Build in the evidence. Numbers, quotations and linked sources right where the statement falls. Those are the three methods with the best values in the test.
  4. Bring the readability along. Fluency and comprehensibility landed above the baseline in the test without any change to the substance. That is the cheapest step of them all.

What I derive from that for my own work is a rule of thumb too. I treat those four points as quality rules for any specialist article rather than as an AI measure. The reason is simple. They sit on my editorial checklist anyway and cost nothing extra. They also still work when the next engine picks differently.

The reverse applies as well: everything Google’s documentation explicitly calls unnecessary costs time without a documented return. For Google Search at least. If you maintain an llms.txt because another system reads it, that is a separate decision with its own reasoning. Whether your text is more than a prettier version of the usual is what my commodity self-test settles.

How to measure whether you appear in AI answers

Key Takeaway: For Google Search there is an official data source, the Generative AI performance report in Search Console. For the other engines none comparable is known to me.

The measurement question is more awkward with GEO than with classic SEO. A position in a result list you can look up and whether your text appears in a generated answer depends on the individual query and on the moment and on the model.

Google’s documentation names exactly one source: “To measure how your content is performing in generative AI features on Google Search and Discover, use the Generative AI performance report in Search Console.” How that report is built and what it shows sits in my analysis of the generative AI performance reports in Search Console.

Three things help when reading those numbers:

  • Impressions and clicks drift apart. A mention inside an answer produces an impression. The click is then decided by a reader who already holds the answer.
  • Visibility spreads per topic rather than per domain. The same website can get quoted regularly on one topic and never on the next. Why that is I described under AI visibility per topic rather than per domain.
  • Single spot checks do not count as measurement. The same query returns different sources depending on the run and the paper’s authors sampled five answers per query for exactly that reason.

For ChatGPT, Perplexity, Claude and Gemini an official reporting surface for website owners is missing and what remains in the end is your own spot checks and the numbers of third-party vendors. Both are estimates and both should be called that.

What GEO tools can do and what they cannot

Key Takeaway: A tool can observe whether a brand shows up in answers. It cannot read out a ranking factor, because no vendor has access to those systems.

By now there is a range of tools reporting an AI visibility. Between what such a tool can do and what it claims runs a fairly clear dividing line.

Observing is something a tool can do. It puts questions to the engines and logs the answers and counts the mentions of a domain in them. That is a real measurement with the limits from the previous section: a spot check at one moment and with a question set somebody picked.

Reading out is something it cannot do. Google’s documentation is unmistakable here: “No third-party tool has access to our internal ranking or AI systems.” A reported “GEO score” is therefore a model built by the vendor rather than information out of the system.

A usable purchasing check follows from that in two questions. First: which questions does the tool ask and how often and against which engines? Second: how does the reported score come about? On the first question there is a verifiable answer and on the second you will hear either a disclosed formula or an explanation of why there is none.

Frequently asked questions (FAQ)

What is generative engine optimization in one sentence?

Generative engine optimization is the work on content with the aim of appearing inside the answers of generative AI systems such as ChatGPT, Perplexity or the AI Overviews. The term comes from the paper “GEO: Generative Engine Optimization” by Aggarwal et al., which appeared as a preprint in 2023 and was accepted to KDD 2024.

Is GEO the same thing as SEO?

From Google Search’s point of view, yes. The documentation writes: “optimizing for generative AI search is optimizing for the search experience, and thus still SEO”. The paper behind the term describes something narrower. It is about text changes that shift how an already retrieved source gets quoted in a generated answer. Both statements are compatible once you read retrieval and generation as two steps.

Who invented the term generative engine optimization?

A six-person author team around Pranjal Aggarwal in the paper “GEO: Generative Engine Optimization”, available as a preprint under arXiv:2311.09735. Version 1 dates from 16 November 2023 and version 3 from 28 June 2024, and it was accepted to KDD 2024.

Do the tested methods help with Google AI Overviews as well?

The paper does not answer that question because it did not test AI Overviews. Measurement happened in a self-built setup with gpt3.5-turbo and additionally on Perplexity.ai. For Google Search, Google’s documentation points to the usual SEO fundamentals rather than to AI-specific measures. What remains of the methods are general quality rules: source it and quantify it.

How long does GEO take to show an effect?

No solid source on that is available to me. The paper measures the effect of a text change in a laboratory setup and makes no statement about timeframes in real systems. What can be said is this: a text has to be retrieved and processed again before a change can become visible in an answer. Any concrete number of weeks you come across is an estimate drawn from experience.

Do I need an llms.txt for AI search?

Not for Google Search. The documentation is unambiguous there: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search”. By its own account Google Search does not use those files, although they do no harm either. For other systems the file can make sense, and that is then a separate decision.

Why does keyword stuffing fail in generative answers?

The paper does not explain the cause, it only measures the outcome: “we find such methods offer little to no improvement on generative engine’s responses”. On Perplexity.ai the variant landed 10 % below the untouched version. The authors conclude that techniques from classic search engine optimisation do not automatically transfer to generative systems.

How do I measure whether I appear in AI answers?

For Google Search there is the Generative AI performance report in Search Console, which Google’s documentation points to explicitly. A comparable official data source from the other engines is not known to me. Third-party tools estimate. Google’s documentation notes on that: “No third-party tool has access to our internal ranking or AI systems.”

Conclusion: GEO is the last metre, not the road

Key Takeaway: The term has a solid source with a clear methodology and a clear boundary. What was measured there presupposes classic ranking. Google’s documentation says the same thing from the other direction.

Two sources, one term, two perspectives. The paper shows that the wording of a text shifts how a generative answer quotes it. Google’s documentation shows that the selection of those sources still happens where it always did. Taken together the two findings paint a fairly unexcited picture of an abbreviation that is now being sold as its own discipline.

One thing stands out: none of the nine tested methods is a trick. Quoting, sourcing, quantifying and writing readably. That is the work a specialist article demands anyway, except that in a generated answer it becomes visible faster than in a classic result list.

Tip: if somebody sells you a GEO package, ask for the source of the promised effect. The one measurement with disclosed methodology that I know of ran on two engines, neither of which was Google. Everything beyond that is experience or assumption. Both may be called by their name.

Status: August 2026. All information without warranty: despite careful research, no warranty is given as to topicality or completeness. The study results reproduced here are based on the information published by the authors of the paper “GEO: Generative Engine Optimization”, for whose accuracy no warranty is assumed. Details about Google Search reflect the documented state and may change. The assessment of the search results page of 16 August 2026 is an own pull through the SERP interface of DataForSEO (Google, Germany, German) and therefore a snapshot. All brands and product names mentioned are the property of their respective owners. This article does not replace individual consulting.

Christian Ott - Gründer von www.seo-kreativ.de

Christian Ott – Creative SEO Thinking & Knowledge Sharing

As the founder of SEO-Kreativ, I live out my passion for SEO, which I discovered in 2014. My journey from hobby blogger to SEO expert and product developer has shaped my approach: I share knowledge in a clear, practical way-without jargon.