Cited sources and the pages ChatGPT read
When ChatGPT searches the web before answering, two lists exist. The first is the pages it cites: the links attached to sentences in the answer. The second, longer one is every page it consulted while searching. A brand can be in the second list and absent from the first. For a buyer question about your category, the first list is the one your buyers see, and the one this guide is about.
An answer written without a web search has no sources to show. If an answer about your category comes back with no citations, ask again and make sure the search ran, or the row in your notes records the model's memory rather than the web it draws on.
The manual method: ChatGPT's own interface
- Write ten buyer questions the way customers type them: "best X for Y", "X alternatives", "X vs Y", "how to choose X". Leave your brand name out; a question that names you mostly measures whether ChatGPT can repeat your name.
- Open a fresh chat for each question and, where your account allows it, switch memory off, so earlier conversations do not shape the answer.
- Ask the question and check that the answer carries citations. If it does not, ask it to search the web and ask again.
- Open Sources under the answer and copy every link, in the order shown, with the sentence each inline citation is attached to.
- Note whether your brand is named, whether your own domain is among the sources, and which competitors appear.
- Do the same next week, with the same wording. The comparison between runs is where the information is.
The limit is labour. Ten questions once a week is ten answers to read and perhaps a hundred links to copy. That is manageable for a month; it does not stretch to several engines or a few dozen questions, which is the point where a tracking tool starts to pay for itself.
The API method: OpenAI's Responses API with web search
The same check can be scripted. You send the question to the Responses API with the web_search tool switched on, and the response comes back with the searches the model ran and the pages it cited, in fields a script can read.
- Request: a model, tools set to a single web_search tool, and the question as input. Add include with web_search_call.action.sources if you want the consulted list as well as the citations.
- Force the search. With the tool only offered, the model may answer from memory: in our own collector on 20 August 2026, an unforced call about "Otterly alternatives" came back about a different product with no citations. Setting tool_choice to the web_search tool makes it search every time.
- Read the output: for each output item of type message, walk its content and keep every annotation of type url_citation. Keep the url and the title; deduplicate by URL within one answer.
- Store the web_search_call queries too. They show what ChatGPT searched for on the way to the answer, which is often not the question you typed.
One caveat applies to every API measurement, ours included: an API answer has no account memory or custom instructions, so it can differ from what a signed-in person sees in the ChatGPT app. Treat API results as a consistent sample of ChatGPT's search behaviour, not as a screenshot of one buyer's screen. Our methodology documents how we collect each engine.
What to record for each answer
- Date, question wording and collection route (app or API, with the model if you know it).
- Whether your brand is named, and whether it is in a list of picks or mentioned in passing.
- Every cited URL, its domain, and its order in the answer.
- Whether any cited URL is on your own domain, and which page.
- The competitors named and the competitor pages cited.
- For API runs, the searches the model reported running.
Domains are the unit that holds still long enough to act on; individual URLs move more. Group by domain first, then look at which pages on the domains that keep appearing are doing the work: a comparison article, a review site's category page, a Reddit thread, a vendor's own pricing page.
One run is noise: repeat before you conclude
ChatGPT's sources change from run to run even when the question does not. Our persistence study asked the same 40 buyer questions of ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews and Google AI Mode every week for six weeks and counted how many cited domains came back:
- ChatGPT: 611 of 1,300 domains cited for a question were cited again for the same question one run later, 47.0%.
- All engines pooled: 3,335 of 6,947, 48.0%.
- A domain cited for the first time came back the following week 533 of 1,950 times, 27.3%.
So a single run tells you which domains ChatGPT drew on that day, not which ones it draws on. Run the same questions weekly for at least four weeks and count, for each domain, in how many runs it appears. A domain present in four runs out of four is a fixture of the answer; a domain present once is a draw. The full study has the per-engine detail, and the sample-size guide covers how many questions a rate needs.
Which sources ChatGPT cites for buyer questions, in public data
Before measuring your own questions, look at the category-wide picture. Our AI Source Index page for ChatGPT lists, every week, the domains ChatGPT cited for 40 fixed English buyer questions, collected through OpenAI's API with web search on, with the rows published under an open licence. The AI Source Index has the same view for the other engines on the panel.
How Promvia does it — full disclosure: this is our product
Promvia runs the API method on a schedule for the questions you track. For every answer it stores every URL the engine cited, the unique domains among them, the cited URLs on your own domain, the searches the engine reported running, whether your brand was named, and which competitors were. ChatGPT is asked through OpenAI's Responses API with the web search tool forced on every call, twice per weekly or on-demand check by default.
- "Where AI gets its answers" ranks the domains cited across your questions by how many questions cite each, top 20, with your own domain and competitors marked.
- A run-to-run comparison lists which domains entered the cited set and which dropped out between the last two runs.
- Each question's page lists the competitor domains cited instead of you.
- On Pro and Agency, the API and MCP server return the full per-question lists (every cited URL and domain, plus the searches the engines ran) for your own scripts or an AI assistant.
What it does not do, stated plainly: ChatGPT is not on the Free plan, which runs Perplexity only, and paid checkout had not opened on 7 October 2026 (see pricing). ChatGPT is checked weekly on every paid plan, not daily. And it measures API answers, with the caveat above about the app.
What to do with the list
- If your own pages never appear, check that OpenAI's search crawler can fetch them: verified AI crawler user agents lists OAI-SearchBot and how to confirm it in your logs, and robots.txt for AI crawlers covers the rules.
- If the same third-party domains appear run after run, those are the pages to study: what they say about your category, whether you are in them, and whether they are out of date. Off-site authority covers that work.
- If a competitor's own page is cited, read it as a reference for what that answer quotes: the format, the numbers, the date.
- Then measure again. How to get cited by ChatGPT collects the on-page and off-page work; no list of fixes can promise the next answer, and the next run is what tells you whether anything changed.
Frequently asked questions
Can I see which sources ChatGPT used without a paid tool?
Yes. In the ChatGPT app, answers that searched the web carry inline citations and, when available, a Sources button listing the cited pages and other relevant links. Recording them by hand for ten questions a week is manageable. The OpenAI API returns the same kind of data in a form a script can store.
Does ChatGPT cite the same sources every time?
No. In our six-week panel of 40 buyer questions, 47.0% of the domains ChatGPT cited for a question were cited again for the same question one run later. Treat one run as a sample and look for the domains that recur across several runs.
Is the API answer the same as what users see in ChatGPT?
Not exactly. An API answer has no account memory or custom instructions, so it can differ from what a signed-in user sees. It is a consistent, repeatable sample of ChatGPT's search behaviour, which is what tracking needs, but it is not one particular buyer's screen.
Does every ChatGPT answer have sources?
Only answers that used web search. An answer written from the model's own memory cites nothing. Over the API you can force the search on every call; in the app, ask for current information or ask it to search the web.