Key takeaways
- GEO means improving how often AI answer engines cite or mention your content. The term comes from a paper first submitted in November 2023 and presented at KDD 2024.
- In that paper's tests, adding quotations, statistics and cited sources helped most and keyword stuffing did little. The main engine was simulated, so treat the results as a hypothesis.
- Google says optimising for generative AI search is still SEO, while Microsoft recommends schema and modular structure. The engines don't agree on every tactic.
- Nobody outside the AI companies knows how sources are weighted, so treat any guaranteed GEO tactic with suspicion.
Generative engine optimisation (GEO) is the practice of improving how often AI answer engines such as ChatGPT, Google’s AI Overviews and Perplexity cite, link to or mention your content. The term has a real academic origin, but much of what’s sold under it is still unproven.
If you work in SEO, GEO feels familiar and slightly irritating at the same time. The goal hasn’t changed, which is being found by people who are looking, but the results page has: instead of ten blue links, the reader gets a written answer with a few sources attached.
This guide covers where the term came from, what the original research found, where Google and Microsoft disagree, and what’s worth doing first. Every figure comes from a page we’ve read in full, and we say plainly where nobody knows.
For the service side, see our AI search page.
Where Did the Term GEO Come From?
It was introduced in a research paper, “GEO: Generative Engine Optimization”, first submitted to arXiv on 16 November 2023 and presented at the KDD 2024 conference. Six authors from IIT Delhi, Princeton University and two independent researchers wrote it.
Who wrote the paper, and why?
The lead author of the GEO paper is Pranjal Aggarwal, and Princeton’s Karthik Narasimhan is among the co-authors. Their starting point was that content creators have little control over when and how generative engines show their material.
They framed GEO as black-box optimisation: you can change your content, but you can’t see inside the engine. That framing still holds, and it explains why so much GEO advice is inference.
How were the tests run?
The authors built GEO-bench, a set of 10,000 queries across 25 domains, each paired with the text of the top five Google results. The headline results use its 1,000-query test split.
The main experiments used a simulated engine. It fetched the top five Google results for a query and had OpenAI’s gpt-3.5-turbo write an answer, sampled five times at a temperature of 0.7.
The team then rewrote one source at a time using nine methods and measured how much of the answer came from it. They repeated the work on Perplexity.ai, a live engine, to check the findings travelled.
What Did the GEO Research Find?
Adding quotations, statistics and cited sources gave the biggest gains in the paper’s tests, and keyword stuffing gave little or none. The best method improved visibility by 41% on the paper’s main measure, which the abstract rounds to “up to 40%”.
Which methods helped, and by how much?
The paper scores each method two ways. Position-adjusted word count weighs how much of the answer comes from your source and how early it appears, and a subjective impression score covers factors such as relevance and likelihood of a click.
| Method | Position-adjusted word count | Subjective impression (average) |
|---|---|---|
| No optimisation (baseline) | 19.3 | 19.3 |
| Keyword stuffing | 17.7 | 20.2 |
| Cite sources | 24.6 | 21.9 |
| Statistics addition | 25.2 | 23.7 |
| Quotation addition | 27.2 | 24.7 |
The scores come from the paper’s Table 1. The authors say their top three methods, cite sources, quotation addition and statistics addition, improved the first measure by 30 to 40% and the second by 15 to 30%.
What happened to keyword stuffing?
The paper found that keyword stuffing offers little to no improvement in generative engine responses. On Perplexity.ai it performed 10% worse than the baseline.
That’s a useful warning for anyone tempted to carry old SEO habits across. It’s also a result on one benchmark, so read it as “this didn’t help in these tests”, not as a law.
Who gained most?
When the authors optimised every source at once, lower-ranked pages gained more. Cite sources lifted visibility by 115.1% for pages ranked fifth in the search results, while the top-ranked page lost 30.3% on average.
That is a simulation with five sources and one model, so don’t expect those numbers on a live site. The direction is still interesting: if every page cites sources, the advantage moves to whoever has something the others lack.
Where does the paper stop?
The authors list their own limits. Methods may need to adapt as engines evolve, the queries will date, and they didn’t test how GEO methods affect search rankings.
Add the age of the work: the main engine was gpt-3.5-turbo, and the paper was first submitted in 2023. It’s good evidence that wording and sourcing can move visibility, and weak evidence for any specific lift on your site today.
What Are the Other Names for GEO?
You’ll also see AEO (answer engine optimisation), LLM optimisation and AI search optimisation, and none of them has an agreed technical definition. Google’s guide says AEO and GEO are both terms for work focused on visibility in AI search experiences.
What does Google call it?
Google’s guide, last updated on 10 July 2026, says that from Google Search’s perspective optimising for generative AI search is optimising for the search experience, and so still SEO. It also tells site owners to review Google’s guidance on evaluating third-party SEO advice before buying “AEO” or “GEO” services.
What does Microsoft call it?
Microsoft uses GEO openly. Its February 2026 announcement of AI Performance in Bing Webmaster Tools describes the release as an early step toward GEO tooling.
A Microsoft Advertising post from October 2025 goes further. It says that whether you call it GEO, AIO or SEO, visibility is what matters.
Do the labels matter?
Not much. Some agency guides split the terms into layers, with AEO for answering, GEO for trusting and LLMO for recommending, but none of the engine documentation we read makes that split.
Judge a provider by the method they describe and the evidence behind it. The acronym tells you nothing.
How Is GEO Different From SEO?
GEO shares most of its foundations with SEO but changes what you aim for and how you measure it. Search ranks pages; an answer engine builds a response and decides which pages to credit.
| SEO | GEO | |
|---|---|---|
| Output | Ranked list of links | Generated answer with citations |
| Goal | Position and clicks | Citations and brand mentions |
| Reporting | Search Console, Bing Webmaster Tools | Google and Bing reports; none from most other engines |
| Stability | Positions move gradually | Answers vary from run to run |
What stays the same?
Google says its generative AI features are rooted in its core Search ranking and quality systems, and that a page must be indexed and eligible to show in Search with a snippet to be a supporting link. Microsoft’s October 2025 post says crawlability, metadata, internal linking and backlinks remain the starting point.
So the groundwork is shared. A site that can’t be crawled or indexed gets nowhere in either world.
What changes?
The unit of competition shifts from the ranked page to the cited source. The GEO paper makes the same point: average ranking position works for a list of links but not for a generated answer with interleaved citations.
You also lose the neat feedback loop of rank tracking. Reporting exists for Google’s AI features and for Microsoft’s, but not for most other engines.
Why is measurement so much harder?
SparkToro’s Rand Fishkin and Patrick O’Donnell of Gumshoe.ai published a study on 27 January 2026. They asked 600 volunteers to run 12 prompts through ChatGPT, Claude and Google’s AI features, which produced 2,961 runs in November and December 2025.
Fishkin reported a less than 1 in 100 chance that ChatGPT or Google’s AI would give the same list of brands in any two responses, and less than 1 in 1,000 for the same order. One caveat: O’Donnell works at Gumshoe, an AI tracking vendor, which Fishkin discloses, and the work isn’t peer reviewed.
Related: How Do You Measure AI Search Visibility?
Want a team that tests changes one variable at a time before rolling them out? Talk to us about AI search for your business.
Do the Engines Agree on What Works?
No. Google says you don’t need special markup, content chunking or llms.txt, while Microsoft’s guidance recommends schema, modular headings and snippable answers.
Both are documented positions from different products, so neither cancels the other.
| Topic | Google (guide updated 10 July 2026) | Microsoft (Advertising blog, 8 October 2025) |
|---|---|---|
| Structured data | Not required; no special schema.org markup | Lists schema among its core practices |
| Content structure | No requirement to chunk; no ideal page length | Headings, Q&A blocks, lists and tables so pages split into reusable pieces |
| llms.txt | Google Search ignores it; neither helps nor harms | Not mentioned in the post |
| What to prioritise | Unique, non-commodity content from first-hand experience | Fresh, authoritative, structured and semantically clear content |
Where do they differ?
Structured data is the sharpest split. Google says it isn’t required for generative AI search but is still worth using for rich results, so schema markup earns its place anyway.
On the proposed llms.txt file, Google says Search ignores it. We haven’t found it recommended in the Microsoft post, and the OpenAI pages we read don’t advise it for ChatGPT search.
Where do they agree?
Both say the SEO fundamentals still matter, and both say there’s no secret strategy. Google puts the weight on content people find unique and useful, while Microsoft talks about clear, current content that its systems can surface with confidence.
What should you do about the gaps?
Where the cost is low and the benefit doesn’t depend on AI, such as clear headings, tables and sensible schema, just do it. Where a tactic exists only because someone says an engine likes it, test it on your own pages before you commit budget.
What Do We Know About How Engines Pick Sources?
Less than the GEO industry implies: each engine has published some detail, and none has published a formula. What follows is what each company has said.
Google says its generative AI features rely on retrieval-augmented generation, which uses its core ranking systems to find relevant pages, and on query fan-out, where the model issues several related searches at once. Pages must be indexed and eligible to show in Search with a snippet.
Related: How Do Google AI Overviews Choose Their Sources?
OpenAI
OpenAI’s help centre says ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information, and that placement isn’t guaranteed. The pages we read don’t list the factors.
Related: How Do You Get Your Brand Cited by ChatGPT?
Microsoft Copilot
Microsoft says its assistants break content into smaller pieces, a step it calls parsing, and evaluate those pieces for authority and relevance before assembling answers. Its Bing Webmaster Tools report shows which pages were cited and a sample of the grounding queries behind them.
The honest gap
What’s unpublished includes how heavily each factor counts, how much brand mentions matter and how often retrieval behaviour changes. Anyone claiming to know is guessing, or has run their own tests and is showing you only the wins.
What Should You Do First?
Start with the basics every engine documents: let the right crawlers in, publish content worth citing and set up a way to measure. Most of it is good SEO anyway.
1. Check access
Make sure robots.txt, your CDN and your firewall aren’t blocking the search crawlers you want. OpenAI says sites that opt out of OAI-SearchBot won’t be shown in ChatGPT search answers, though they can still appear as navigational links.
Whether to block training crawlers is a separate decision, and we cover it in should you block AI crawlers.
2. Publish what others can’t
Google’s guide says unique, non-commodity content will likely influence your presence in generative AI search more than any other tip in it. Its example of commodity content is a generic “7 Tips for First-Time Homebuyers” listicle.
Original data, first-hand testing and named methods all qualify, and that’s the job of content marketing done properly. The same guide warns that creating separate pages for every fan-out query, primarily to manipulate responses, violates its scaled content abuse policy.
3. Back claims with evidence
The paper’s best performers were statistics, quotations and cited sources, and Microsoft’s February 2026 post also tells publishers that examples, data and cited sources build trust when content is reused. That fits ordinary editorial practice and the trust signals in E-E-A-T.
It’s cheap to try and unlikely to hurt, but test it on your own pages before assuming the paper’s lift transfers.
4. Skip what Google calls myths
Google’s guide says you can ignore chunking, rewriting content just for AI systems and over-focusing on structured data. It adds that seeking inauthentic mentions across the web isn’t as helpful as it might seem.
Genuine coverage in publications is a different matter, and that’s what digital PR is for.
5. Measure with a repeatable prompt set
Track a fixed list of prompts over time, run each several times and record mentions and citations. If you’d rather have a baseline before you start, an AI visibility audit records what the engines say about you today.
How Do You Spot Poor GEO Advice?
Look for guarantees, single-run screenshots and claims with no named source. Good GEO advice states its method, its sample size and its limits.
| What you’ll hear | What the sources say |
|---|---|
| “Quotations raise AI visibility by 40%” | The abstract says up to 40% for the best methods on its benchmark, using a simulated engine plus Perplexity.ai |
| “You need an llms.txt file” | Google says its Search ignores them; the OpenAI pages we read don’t advise them for ChatGPT search |
| “We track your ranking position in AI answers” | SparkToro’s study found positions too random to track, and Fishkin called any tool that offers one “full of baloney” |
| “We guarantee citations” | OpenAI says placement isn’t guaranteed; Google says indexing and serving aren’t guaranteed |
| “Build a page for every fan-out query” | Google says doing so primarily to manipulate responses violates its scaled content abuse policy |
Which questions should you ask a provider?
Ask how many prompts they track, how many times they run each one and what account settings they use. Ask whether they report ranges or a single number, and how they’d tell a real change from noise.
A provider who can’t answer those hasn’t done the maths. We work through the arithmetic in our guide to measurement.
What does a fair test look like?
Change one thing on one set of pages, leave comparable pages alone and re-run the same prompts on a schedule. At Utterly Digital we test ranking factors in single-variable experiments before applying them to client sites, and the same discipline suits GEO.
Need an outside view of how your site shows up in AI answers? Explore Utterly Digital’s AI search service.
| SEO | GEO | |
|---|---|---|
| Where you appear | A ranked list of links | Inside a generated answer |
| Win condition | Position and click | Citation or brand mention |
| Official reporting | Search Console, Bing | Google and Bing only |
| Result stability | Fairly stable | Changes run to run |
FAQs
Is GEO worth it for a small business?
Start with the basics before paying for anything specialist. Google's guide says plenty of content does well in its generative AI features without any overt SEO, and for local firms it points to keeping Business Profile details up to date.
Do I need to rewrite my old content for AI?
Google says you don't need to write in a specific way just for generative AI search, because its systems understand synonyms and meaning. Fix pages that are thin, out of date or unsourced, but don't rewrite good pages to chase a format.
Is there an ideal page length for AI answers?
Google says there's no ideal page length and that shorter or longer pages can both work, depending on the audience and subject. Write as much as the topic needs and no more.
Can I use AI to write GEO content?
Google says you can use generative AI tools to help create content, as long as the result meets its Search Essentials and spam policies. The unique-viewpoint test still applies, and a page that only restates what's already online is the weak spot.
Do clicks from AI Overviews convert better?
Google says it has seen that clicks from results pages with AI Overviews are higher quality, meaning people spend more time on the site. That's Google's own observation and it hasn't published the data, so check your analytics before assuming it holds for you.
How much does it cost to find out where we stand?
An AI visibility audit from Utterly Digital starts at £500. It records how the main AI engines describe your brand today, which gives you a baseline before you change anything.
Sources
- GEO: Generative Engine Optimization (abstract page), arXiv (Aggarwal et al., KDD 2024)
- GEO: Generative Engine Optimization (full text, version 3), arXiv (Aggarwal et al., KDD 2024)
- Optimizing your website for generative AI features on Google Search, Google Search Central
- AI features and your website, Google Search Central
- Optimizing Your Content for Inclusion in AI Search Answers, Microsoft Advertising blog
- Introducing AI Performance in Bing Webmaster Tools Public Preview, Microsoft Bing Webmaster Blog
- Searching the web with ChatGPT, OpenAI Help Center
- NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility, SparkToro
Related services