ChatGPT searching a specific subreddit happens directly whenever a community is judged relevant to the user's question. This finding is interesting because it shows ChatGPT's retrieval doesn't always stop at the domain level like reddit.com; in a specific case, the system can target a particular community such as r/whatnotapp before it ever fetches a page.
That finding is covered in a Search Engine Journal article on ChatGPT searching subreddits by name. The article summarizes an experiment by Suganthan Mohanadasan that captured ChatGPT Search's internal queries and found that one search was explicitly scoped to a specific subreddit.
What makes this finding important isn't just the fact that Reddit can show up as a source. The case shows that source selection in AI search is highly query-dependent. In one question, Reddit dominated retrieval and citation. In a different commercial question tested just a few days earlier, Reddit was still fetched heavily but got zero citations at all.
Table of Contents
- What Did the ChatGPT and Reddit Experiment Actually Find?
- Why Does the Fact That ChatGPT Targets a Specific Subreddit Matter?
- Does Reddit Always Get Cited by ChatGPT?
- What's the Difference Between Retrieval and Citation in AI Search?
- Why Does This Change How We Should Read AI Citation Data?
- Why Did the "Whatnot Seller Tips" Query Favor Reddit More?
- Does ChatGPT Only Use Bing for Web Search?
- What's the Connection to the OpenAI–Reddit Partnership?
- Does Reddit's Robots.txt Block ChatGPT?
- Why Do OAI-SearchBot and GPTBot Need to Be Distinguished?
- What Does This Mean for a Reddit Strategy in the Generative AI Era?
- Does a Brand Need to Create Its Own Subreddit?
- How Can a Brand Monitor Its Reddit Visibility for AI Search?
- What's the Risk of Misreading a Drop in Reddit Citations?
- What Can't Be Concluded From This Experiment Yet?
- What's the Implication for Generative AI Search Optimization?
- FAQ About ChatGPT, Reddit, and Subreddits
- Can ChatGPT search a specific subreddit?
- How many Reddit pages were fetched in that experiment?
- Will every fetched Reddit page get cited?
- Does ChatGPT only use Bing for web search?
- Does Reddit block OpenAI's crawlers?
- Should a brand market aggressively on subreddits?
- Conclusion
What Did the ChatGPT and Reddit Experiment Actually Find?
The main experiment used a query about tips for selling on Whatnot. ChatGPT generated a search that was explicitly scoped to one subreddit:
site:reddit.com/r/whatnotapp seller tips Whatnot live sellingThat search used a time window of 3,650 days, or roughly ten years. Out of 71 pages that entered the retrieval pool, 48 came from Reddit. That means roughly 68% of the pages fetched in that conversation came from the same platform.
Even more notably, Reddit wasn't just a retrieval source. Six of the eight citations that finally appeared came from threads on r/whatnotapp.
Metric | "whatnot seller tips" query |
|---|---|
Total pages fetched | 71 |
Reddit pages | 48 |
Reddit's share of the retrieval pool | About 68% |
Total citations | 8 |
Citations from Reddit | 6 |
But that result comes from a single experiment with a single query and shouldn't be generalized into a rule that ChatGPT always prioritizes Reddit.
Why Does the Fact That ChatGPT Targets a Specific Subreddit Matter?
Because it shows retrieval can operate at the community level, not just the domain level. ChatGPT isn't simply requesting results from Reddit in general; the system selects the community it judges most relevant before ever fetching a page.
In the context of information retrieval, this step shows a fairly specific query-expansion or source-targeting process. The system can narrow its search space based on the context of the question.
For marketers and generative AI teams, the implication is significant. Visibility on Reddit doesn't always mean "the brand has to be popular on Reddit overall." What's more relevant may be a presence in the specific community that holds the strongest knowledge on a given topic.
For example, for a question about the experience of selling on Whatnot, the r/whatnotapp subreddit carries very specific practical knowledge. For a question about live-chat software, vendor sources, documentation, and comparison sites can compete far better.
Does Reddit Always Get Cited by ChatGPT?
No. This is one of the most important findings from the experiment.
Four days before the Whatnot test, Mohanadasan tested a different commercial query: "best ai live chat support software." For that query, ChatGPT fetched 221 pages, including 84 Reddit threads.
But the final result was completely different.
Query | Pages Fetched | From Reddit | Total Citations | Reddit Citations |
|---|---|---|---|---|
whatnot seller tips | 71 | 48 | 8 | 6 |
best ai live chat support software | 221 | 84 | 11 | 0 |
So Reddit can contribute a large share of the retrieval pool while still getting zero credit in the final citations.
This difference shows that retrieval and citation are two stages that need to be analyzed separately.
What's the Difference Between Retrieval and Citation in AI Search?
Retrieval is the stage where a system searches for and fetches candidate sources that might be useful for building an answer. Citation is the stage where some of those sources get selected and shown to the user as a reference.
In simple terms:
User Prompt
↓
Query generation
↓
Source retrieval
↓
Candidate pages
↓
Answer generation
↓
Citation selection
↓
Final responseA page can enter the retrieval stage without ever appearing in the citations.
In the live-chat software case, 84 Reddit threads made it into the retrieval pool, but not one became a citation. That means the problem wasn't that ChatGPT couldn't access Reddit for that query.
The Reddit pages were available to the system, but simply weren't chosen as a visible source.
Why Does This Change How We Should Read AI Citation Data?
Because a drop in citations for a domain doesn't automatically mean that domain is no longer being used as input.
If a team only looks at citations, they might conclude Reddit has "disappeared" from ChatGPT. This experiment shows that conclusion is too broad.
In one query, Reddit can be read in large volume yet get no credit. In another query, Reddit is not only read but dominates the citations.
That means aggregate citation share and retrieval relevance can move in completely different directions.
Why Did the "Whatnot Seller Tips" Query Favor Reddit More?
The source author offers a hypothesis: the most useful knowledge for that query genuinely lives inside a niche community.
Whatnot is a marketplace with very specific seller behavior, live-selling strategy, and user experience. A lot of practical insight — like when to start an auction, how to build entertainment value into a livestream, or how to schedule a show — comes from community members' own experience.
For a question like that, a forum can be the most natural source available.
By contrast, a category like live-chat software has plenty of alternative sources:
- a vendor's pricing page;
- documentation;
- a comparison article;
- a review platform;
- a product landing page.
But it's worth noting that this explanation is the experiment author's own interpretation, not a mechanism OpenAI has confirmed.
Does ChatGPT Only Use Bing for Web Search?
The experiment challenges the simple assumption that ChatGPT's web results are just a direct reflection of Bing's results.
The author checked six queries on Bing, including "whatnot seller tips," "best crm for small business," and a few other product queries. In the rendered Bing SERPs checked, Reddit didn't appear for any of those queries.
But on ChatGPT, the Whatnot query produced a retrieval pool that was 68% Reddit.
That shows, at least in this experiment, that the sources available to ChatGPT aren't identical to the search-result page a user would actually see on Bing.
The author raises a few possibilities: OpenAI may use a different index, or a licensed data feed from its partnership with Reddit. But the experiment's own sources can't determine which path is actually being used.
So the safe conclusion is: ChatGPT's retrieval can't be assumed to be identical to an ordinary Bing SERP.
What's the Connection to the OpenAI–Reddit Partnership?
OpenAI and Reddit have a data partnership that lets OpenAI access Reddit content in a structured way. That makes crawler-based analysis more complicated.
Robots.txt only governs crawling. If data is obtained through a licensed feed or an API, the mechanism is completely different from a crawler opening pages one at a time.
Mohanadasan's experiment can't determine whether the Reddit threads in the retrieval pool came from:
- OpenAI's own index;
- a licensed Reddit data feed;
- a direct crawl of
www.reddit.com; - a combination of several mechanisms.
So the data shows Reddit content is available to the system, but it doesn't prove the exact technical path it took.
Does Reddit's Robots.txt Block ChatGPT?
It's not as simple as reading a public robots.txt file.
The experiment summarized in the article found that Reddit responds differently depending on user-agent, and possibly on network-identity verification too.
In tests run from the same connection:
User-Agent | Observed Response |
|---|---|
Chrome | 200 |
OAI-SearchBot | 200 |
GPTBot | 403 |
ChatGPT-User | 403 |
Googlebot | 403 |
The author interprets this pattern as a sign that Reddit isn't just trusting the user-agent string — it may also be verifying the crawler's IP or identity.
But a spoofed-user-agent test from an ordinary connection can't show what an officially verified crawler IP address actually receives.
So a single curl test from a regular computer can't be used to conclude how Reddit treats OpenAI's official crawlers.
Why Do OAI-SearchBot and GPTBot Need to Be Distinguished?
Because crawlers from the same company can serve different functions.
In that test, OAI-SearchBot got a 200 response while GPTBot got a 403 from the same testing address. That reinforces why every OpenAI user-agent shouldn't be treated as a single, interchangeable thing.
In practice, a web team needs to know a crawler's function before writing a blocking rule.
A bot used for search discovery can carry different implications than a bot used for model training or user-requested fetching.
If your team has previously covered blocking AI crawlers, the internal article AI Search Citation Sources Differ by Industry: What Should You Optimize? can serve as an internal link once the previous article's URL is verified.
What Does This Mean for a Reddit Strategy in the Generative AI Era?
This finding doesn't mean a brand should spam subreddits to earn citations. That approach is actually likely to fail, since Reddit communities are usually quick to spot inauthentic self-promotion.
The more useful insight is this: a niche community can become a high-value knowledge source when a topic doesn't have a better formal source available.
For a brand, the sensible strategy is making sure organic discussion about the product, category, and user problems carries accurate information.
A few areas worth monitoring:
- industry-category subreddits;
- product-specific subreddits;
- comparison threads;
- troubleshooting threads;
- buyer questions;
- genuine usage experience;
- reviews and recommendations.
The goal isn't chasing citations manipulatively — it's understanding where users share experience a retrieval system might judge as relevant.
Does a Brand Need to Create Its Own Subreddit?
Not automatically. Creating a subreddit purely for "AI SEO" purposes likely won't help if there's no real community actually living in it.
A subreddit has value when it has:
- genuine discussion;
- helpful answers;
- user experience;
- good moderation;
- an archive of questions and solutions;
- ongoing activity.
In the Whatnot experiment, the subreddit's value came from specific community knowledge that was genuinely relevant to a seller's question.
So community quality matters far more than simply owning the URL /r/brandname.
How Can a Brand Monitor Its Reddit Visibility for AI Search?
Start from queries that resemble real user questions, not just brand keywords.
- Map out buyer questions. Gather discovery, comparison, troubleshooting, and recommendation questions.
- Find the relevant subreddits. Identify the communities that discuss the category most often.
- Audit the threads that rank or show up often. Look at what's being said about the brand, the product, and competitors.
- Test prompts on AI search. Note whether Reddit shows up as a citation.
- Separate retrieval from citation where the data allows it. Don't conclude a source isn't being used just because it's not visible.
- Monitor for change regularly. Citations can shift based on query and time.
What's the Risk of Misreading a Drop in Reddit Citations?
The biggest risk is making a strategic decision based on an aggregate metric without looking at query-level behavior.
For example, a dashboard shows Reddit's overall citation share dropping. A team then concludes forums no longer matter for generative AI visibility.
This experiment shows two things can be true at the same time:
- Reddit's aggregate citation share is declining;
- Reddit still dominates citations on a specific query when that community happens to be the most relevant source.
So the strategy should focus on query-category fit, not just a domain-level trend.
What Can't Be Concluded From This Experiment Yet?
The original source is very explicit about its limitations. The testing was done on only four queries, one account, and within a single week.
Two of the four queries didn't pull in Reddit at all, while one query didn't trigger a search at all.
The researcher also couldn't determine:
- exactly when ChatGPT chooses to target Reddit;
- whether the retrieval came from a crawl, an OpenAI index, or a licensed feed;
- whether a Reddit page that wasn't cited still influenced the answer;
- how the citation-selection mechanism actually works;
- whether the same pattern holds globally.
So the headline that "ChatGPT searches subreddits by name" is supported by the observation. But a claim like "ChatGPT always prioritizes niche subreddits" isn't supported by the data yet.
What's the Implication for Generative AI Search Optimization?
The biggest lesson is that AI search optimization needs to move from domain-level thinking to source-context thinking.
In traditional SEO, we're used to asking:
Which domain is ranking?In generative AI search, the question needs to expand:
What sources get fetched?
Which community gets selected?
Which pages enter the retrieval pool?
Which sources ultimately get cited?
What changes based on the query?This shift in perspective matters, because a single domain can have a completely different fate on two very similar queries.
FAQ About ChatGPT, Reddit, and Subreddits
Can ChatGPT search a specific subreddit?
Yes, at least in the experiment documented by Suganthan Mohanadasan. ChatGPT generated a search query explicitly targeting reddit.com/r/whatnotapp.
How many Reddit pages were fetched in that experiment?
For the "whatnot seller tips" query, 48 of the 71 pages that entered the retrieval pool came from Reddit — about 68%.
Will every fetched Reddit page get cited?
No. On the live-chat software query, ChatGPT fetched 84 Reddit threads out of 221 total pages but gave Reddit zero citations.
Does ChatGPT only use Bing for web search?
The experiment shows ChatGPT's retrieval pool can contain many Reddit pages even when the rendered Bing SERP for the same query shows no Reddit at all. That means ChatGPT's results can't be assumed identical to an ordinary Bing SERP.
Does Reddit block OpenAI's crawlers?
The situation is more complex than that. Reddit can treat different user-agents differently, and it also has a data partnership with OpenAI. A test from a regular user can't confirm what an official crawler actually receives.
Should a brand market aggressively on subreddits?
No. This finding shows the value of a genuinely relevant community, not a validation of spam or manipulation. Participation should follow community rules and provide real value.
Conclusion
ChatGPT searching a specific subreddit directly happened in this recent experiment — it even set a ten-year search window and pulled most of its retrieval pool from that one community. On the Whatnot query, Reddit ended up winning six of eight citations.
But four days earlier, a different query pulled in 84 Reddit threads and gave Reddit zero citations. This contrast shows that ChatGPT's use of Reddit is highly query-dependent, and that retrieval isn't the same thing as citation.
For generative AI, SEO, and brand teams, the implication is that monitoring needs to happen at the prompt and source level, not just the domain level. A niche subreddit can become an extremely strong knowledge source when community experience happens to be the best answer to a user's question.
If your business wants to build an AI visibility monitoring system, citation analysis, community intelligence, or a content distribution strategy that's better prepared for generative search, you can discuss your business's technology needs with our technical team.




Comments
Got a question or feedback? Leave a comment!