Skip to main content
Module 1: The New Search Landscape

How AI search picks its sources

Retrieval and grounding — and being honest about what's unknown.

To influence whether AI cites you, it helps to understand — at a high level — how these systems choose sources. The general architecture is known; the exact selection logic is not. Being clear about both keeps you effective and skeptical of anyone claiming certainty.

The general architecture: retrieval + grounding. Most AI answer engines work by retrieval-augmented generation (RAG): given your question, the system searches for relevant documents, ranks and filters them, and the AI writes an answer grounded in those retrieved sources, usually with citations. The key implication: to be cited, your content generally has to be retrievable (findable and indexed) and relevant to the question. Grounding in real sources is also why these answers cite anything at all — and why being a trusted, findable source matters.

How the big systems differ (roughly). Perplexity is retrieval-first — it searches and assembles sources on essentially every query. ChatGPT's search is more selective — it decides whether to search before answering, so citations are less consistent. Google's AI Overviews, per Google, run on its 'core Search ranking and quality systems' — the same index and ranking pipeline as normal search, not a separate secret system. So classic search visibility feeds AI visibility.

Be honest about what's unknown. Here's the crucial part most 'GEO gurus' skip: the exact logic for which sources get cited is proprietary and undisclosed. Detailed vendor descriptions of Google's pipeline ('narrows 200 candidates to 5–15 via semantic retrieval, then an E-E-A-T filter, then re-ranking…') are reverse-engineered guesses, not confirmed facts. Specific numbers ('96% of citations have strong E-E-A-T,' 'exactly 5–15 sources') are illustrative vendor claims. Treat them as informed speculation, not gospel — and be wary of anyone selling certainty about a black box.

One useful, if uncertain, observation. Studies suggest AI-answer citations don't map neatly onto the classic top-10 rankings — pages ranking lower, or off page one, still get cited sometimes. This is evolving third-party data, but it hints that AI citation is related to, but not identical to, traditional ranking. So strong SEO helps, but it isn't the whole story.

The takeaway: AI answer engines mostly work by retrieval-augmented generation — they search, rank, and filter sources, then write an answer grounded in them with citations — so to be cited, your content must be retrievable and relevant. Systems differ (Perplexity searches every query; ChatGPT decides whether to; Google's AI Overviews run on its core Search systems, so classic visibility feeds AI visibility). But be honest: the exact citation-selection logic is proprietary and undisclosed, so detailed vendor 'pipeline' descriptions and precise numbers are reverse-engineered guesses, not facts. Strong SEO clearly helps, AI citation relates to but isn't identical to ranking, and you should distrust anyone selling certainty about a black box.

Try it

Pick a question in your field and compare how Google's AI Overview, ChatGPT, and Perplexity each answer it — do they cite the same sources? Different ones? Any that rank #1 in normal search, and any that don't? You're observing directly that AI citation relates to, but differs from, classic ranking — and that no one fully controls the black box.

Stay in the loop

Enjoying the free lessons? Get an email when we publish new courses and updates — no spam, unsubscribe anytime.

Discussion (0)

Ask a question or share what worked for you. Comments are reviewed before they appear.

Log in to join the discussion and ask questions about this lesson.

No comments yet. Be the first to start the discussion!