The two-stage reality: retrieve, then quote
Engines with browsing (Perplexity, ChatGPT search, Gemini) work in two stages. First a retrieval step finds candidate pages — this looks a lot like classic search, which is why SEO authority still matters. Then a synthesis step decides which passages actually make it into the answer. You can win retrieval and still lose synthesis: your page was read, but nothing in it was worth quoting.
Most GEO advice fails by only addressing stage one. The passage-level game — being the sentence an engine wants to lift — is where sites that consistently get cited differ from sites that merely rank.
What the evidence says engines prefer
- Answer-shaped passages: a heading that matches the question with a direct 2–3 sentence answer under it.
- Provenance signals: named sources, specific numbers, and recent years IN the text — engines prefer quoting claims that carry their own evidence.
- Freshness: visibly dated, recently updated pages beat undated ones for time-sensitive questions.
- Machine-readable identity: structured data (Organization, FAQ) doesn't guarantee citations but removes ambiguity about who's making the claim.
- Domain trust: for "best X" questions, engines disproportionately cite aggregators, editorial comparisons, and community threads — consensus sources — rather than the vendors being compared.
What this means for your strategy
- Audit passages, not just pages: for each question you want to win, does a single liftable paragraph exist anywhere on your site? If not, write it.
- Give your claims receipts: numbers with sources and dates. "Rated 4.8/5 across 214 reviews (G2, 2026)" is quotable; "customers love us" is not.
- Work the consensus layer: identify the exact domains engines cite for your category and earn honest presence there. This is slower than on-site fixes, and it is where most of the cited evidence sits.
- Sample repeatedly: answers vary run to run, so judge trends across scheduled checks, not single spot-checks.
Promvia's citation tracking records which domains each engine cites for your tracked questions — turning "where do engines look?" from guesswork into a ranked list you can act on.
Frequently asked questions
Do engines cite the same sources every time?
No — answers and citations vary between runs, engines, and phrasings. That variance is why single checks mislead and scheduled multi-engine sampling is the honest measurement.
Does being in the training data matter more than being crawlable now?
For engines with live search, current crawlability and current sources dominate. Training-data presence matters most for offline answers, and you influence it the same slow way: by being widely and credibly mentioned.
Can I pay to be cited?
Not in organic answers today. Sponsored placements are appearing in some engines, but they're labeled — the synthesis itself draws on organic sources.