Guide 1 of 9 · Free, in full, no email required

How AI Actually Finds You

This guide is free because it has to be. It sets up the framework the rest of this series leans on, and a series about getting AI systems to cite you would be a bad advert for itself if its own foundational thinking were locked behind a paywall an AI system can't read. What follows is the complete guide, not a preview.

Crawling and citation are not the same problem

For twenty years, "being found online" meant one thing: get indexed by a search engine, rank for the right queries, and show up as a blue link. That's still true, and it still matters. But it answers a narrower question than most businesses now need answered, because a growing share of the people who might buy from you never see a list of links at all. They ask a question inside an AI system and get a written answer, sometimes with sources named, sometimes not.

Being indexed by Google tells you almost nothing about whether that answer engine can use your page as a source. These are different systems doing different jobs, and treating them as the same problem is the single most common mistake businesses make when they first try to get "AI visible."

What a search engine actually does

A traditional search engine's core job is retrieval by keyword match and relevance ranking: it crawls pages, builds an index of what words appear where, and at query time it ranks candidate pages against hundreds of signals (topical relevance, links, engagement, technical health) to produce an ordered list. Your page either makes that list or it doesn't. Nobody reads it for you first; the searcher does that themselves after clicking through.

What an AI answer system actually does

An AI search or answer system does something structurally different. Instead of handing the user a list of pages to go read, it reads a set of candidate pages itself, in the moment, and writes an answer using what it found - sometimes citing sources, sometimes just absorbing the information into a general answer with no attribution at all. Broadly, that happens in three stages, and each one is a separate place your business can fail to show up, regardless of how well you're doing at the others.

Stage 1: Discovery

Before anything can be retrieved, it has to be discoverable at all. Some AI systems crawl the web with their own bots, similar in principle to a search engine crawler, and some lean on an underlying search index (their own or a licensed one) to find candidate pages for a given query. Either way, if your content is blocked by robots.txt for the relevant agents, hidden behind client-side JavaScript with nothing in the initial HTML response, or simply never linked to from anywhere crawlable, it is invisible at this stage and nothing downstream matters. This is a purely technical gate, and it's the subject of Guide 2.

Stage 2: Retrieval

Once a page is discoverable, the system still has to decide it's worth pulling into context for a specific question. Most modern AI answer systems use some form of retrieval: given a query, find the passages of text (not necessarily whole pages - often small chunks) that are most relevant to answer it, and feed those into the model as source material. This is where writing that's easy for a human to skim but hard to extract a clean, self-contained answer from starts to lose out to writing that states a fact or definition plainly, in a paragraph that makes sense on its own. A page can be perfectly discoverable and still be a poor retrieval candidate if every sentence depends on three sentences before it for context.

Stage 3: Citation

Being retrieved into context still isn't the same as being cited. A model can read your page, quietly use the fact it contains, and never name you - or name a competitor's page that said the same thing more clearly, more concisely, or with a credential the model weighted more heavily. What tips a system toward actually naming a source rather than just absorbing it varies by system and changes over time, but two things show up consistently: the source answers the specific question cleanly, and the source looks like a credible, real, checkable entity rather than an anonymous or thin page. The rest of this series is mostly about that last stage - proving you're a real, citable source, not just a discoverable one.

Why this distinction is the whole framework

Every guide after this one assumes you already understand that discovery, retrieval, and citation are three separate hurdles, because the fixes for each are different and businesses routinely pour effort into the wrong one. A site with excellent technical crawlability (stage 1 solved) but content written entirely as persuasive marketing copy rather than citable fact (stage 2 weak) will still be invisible in AI answers, even though it would rank perfectly well in traditional search. A site with genuinely citable content but no verifiable identity behind it (stage 3 weak) may get read and quietly used, but never named. Knowing which stage you're actually failing at is most of the work; guessing wastes effort on the wrong fix.

What's next

Guide 2 is the technical checklist for stage 1. Guides 3 and 4 cover the site structure and page types that make stage 2 work. Guides 5 and 6 are about stage 3 - proving you're real, and making sure your identity is consistent everywhere an AI system might cross-check it. Guide 7 is how to actually measure whether any of this is working. Guides 8 and 9 are about compounding an advantage once the basics are in place.

PDF of this guide£4.99
Complete series (all 9)£69.99

See all nine guides