Skip to main content

Can You Allow OAI-SearchBot While You Block GPTBot in Robots.txt?

admin9 min read
Illustration showing how to allow OAI-SearchBot while you block GPTBot in a robots.txt policy

Yes, you can allow OAI-SearchBot while you block GPTBot in your robots.txt file, and doing so lets you separate ChatGPT search participation from AI training data collection. OpenAI documents these as two distinct crawlers with different jobs: OAI-SearchBot fetches pages that may be considered for ChatGPT search results, while GPTBot gathers content that may train OpenAI’s generative models, according to OpenAI’s crawler documentation. Setting up this allow-OAI-SearchBot, block-GPTBot robots.txt policy tells cooperating crawlers your preference, but it does not obligate OpenAI to show your pages in any answer, and it cannot retroactively remove anything already used for training. This guide covers the two crawlers’ purposes, a copyable robots.txt policy, how to verify it actually reached your live server, and where a policy like this commonly goes wrong.

OAI-SearchBot vs. GPTBot: Why You Can Allow One and Block the Other in Robots.txt

OpenAI operates several distinct crawlers, and two of them matter most for this decision. OAI-SearchBot is the crawler behind ChatGPT’s search features; when it fetches a page, that page becomes eligible for consideration when ChatGPT search answers a query, according to OpenAI’s crawler documentation. GPTBot is a separate crawler that gathers content which may be used to train OpenAI’s generative AI foundation models. These are not the same job, and OpenAI documents separate robots.txt preferences for each one, which is why you can address them independently instead of writing one blanket rule for OpenAI as a whole.

Because the two crawlers serve different purposes, a publisher who wants search visibility without contributing to model training has a legitimate, documented way to express that preference. The rest of this guide shows the policy itself, how to verify it took effect, and what it does not promise.

A Sample Robots.txt Policy to Allow OAI-SearchBot and Block GPTBot

The example below expresses a search-allowed, training-crawl-disallowed policy for a site that wants its public pages available to OAI-SearchBot but keeps an /account/ path out of crawling entirely:

User-agent: OAI-SearchBot
Allow: /
Disallow: /account/

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /account/

Replace /account/ with your site’s actual private paths, and remember that robots.txt only asks cooperating crawlers to stay away; it is not an access control mechanism. Use authentication, not robots.txt, to actually protect private content. Publish this as a UTF-8, plain-text file at the root /robots.txt of each host that needs the policy.

Notice that the /account/ disallow appears twice. RFC 9309, the Robots Exclusion Protocol standard published in September 2022, specifies that a crawler matches against the group whose user-agent token equals its own name, and a named group does not inherit rules from the wildcard User-agent: * group. Because OAI-SearchBot has its own dedicated group, it only reads what is inside that group, so the account path has to be repeated there if you want it respected. Within a single group, RFC 9309 and Google’s robots.txt documentation both describe most-specific-path matching, with Allow preferred only when an Allow and a Disallow rule match a path equally; here the longer /account/ pattern is more specific than Allow: /, so it wins even inside the OAI-SearchBot group.

Draft and format a policy like this with the Robots.txt Generator, then treat the generated draft as a starting point, not a finished deployment, until you confirm it on the live server.

Test the Live File Before You Trust the Draft

A draft is not evidence of what your server actually serves. After publishing, fetch the live /robots.txt directly and read it exactly as a crawler would, then check specific URLs against it.

  1. Open the Robots.txt Tester and check whether OAI-SearchBot and GPTBot are available as selectable or custom user agents; if they are, test your homepage, a representative public article, and the excluded path against each one separately.
  2. Use the Robots.txt Rule Tester to confirm individual allow and disallow lines behave the way you expect, particularly the interaction between the wildcard group and the OAI-SearchBot group.
  3. Confirm the expected outcome for the sample policy above: public pages allowed for OAI-SearchBot, every path disallowed for GPTBot, and /account/ disallowed for OAI-SearchBot even though the rest of the site is open to it.
  4. Repeat the check for every relevant scheme, host, and port, including both www and non-www versions and any subdomain that serves its own content, since Google’s robots.txt documentation confirms that a robots.txt file’s scope is limited to the exact protocol, host, and port it was fetched from.

If either tool does not currently accept OAI-SearchBot or GPTBot as a named agent, fall back to testing against a generic or wildcard agent to confirm the wildcard group’s behavior, then rely on careful manual reading of the named groups for the OpenAI-specific rules.

What Allowing OAI-SearchBot Does Not Guarantee

Allowing OAI-SearchBot changes eligibility, not outcome. OpenAI’s crawler documentation states that a robots.txt change can take approximately 24 hours to affect its search systems, so do not expect an instant result after editing the file.

Being crawlable also does not mean a page will appear in a ChatGPT search answer or receive any referral traffic; OpenAI decides what to surface for a given query, and a passing crawl is only one input into that decision. OpenAI’s publisher FAQ separately describes that a page opted out of OAI-SearchBot will not be shown in ChatGPT search answers but can still appear as a navigational link, and it also describes limited link-and-title surfacing in ChatGPT Atlas when a disallowed URL is learned through another source; that Atlas behavior is specific to Atlas and should not be assumed to apply to every ChatGPT search feature.

Disallowing GPTBot going forward also does not retroactively remove content that a training crawl already collected before the rule existed, and it has no effect on unrelated crawlers that ignore robots.txt entirely. Treat this policy as a forward-looking preference, not a removal request.

Troubleshooting Sitewide Blocks and Server Access

Most unintended blocks trace back to a handful of predictable mistakes. Work through this list before assuming your policy is broken:

  • Search the file for a stray Disallow: / sitting inside the wrong group. A sitewide User-agent: * / Disallow: / does not by itself block OAI-SearchBot if a separate OAI-SearchBot group already contains Allow: /, but an explicit group that only contains Allow: / can unintentionally remove narrower restrictions that used to apply through the wildcard group, which is why the account path is repeated in the sample policy above.
  • Check for duplicate named groups for the same user agent; conflicting duplicate groups create ambiguity that different parsers may resolve differently.
  • Confirm the file is actually reachable. RFC 9309 draws a distinction between a robots.txt file that is unavailable, such as one returning a 4xx response, which may be treated as permitting crawling, and one that is unreachable due to a server error such as a 5xx response, which calls for assuming a complete disallow. To check the basics, run your site through the Website Down Checker, which tests whether a website is up and reports availability from multiple locations. It is an availability check only: it does not tell you which HTTP status code your robots.txt returned, so it cannot separate the 4xx and 5xx cases above. To see the exact response for /robots.txt, request that URL directly in a browser or with your server logs and look at the status code yourself.
  • Rule out CDN or firewall interference. A challenge page, a login wall, or a web application firewall rule can block a crawler even when robots.txt allows it, and robots.txt compliance is voluntary for the crawlers that choose to honor it in the first place, per RFC 9309’s explicit statement that the protocol is not a form of access authorization.
  • If you maintain an IP allowlist for OAI-SearchBot, verify requests against OpenAI’s published IP ranges rather than trusting the user-agent string alone, since that string can be spoofed by unrelated bots.
  • Use the Spider Simulator to see what a crawler-style fetch actually returns for a given page, including whether a robots meta directive in the HTML contradicts your robots.txt policy. Neither this nor any other generic tool can confirm that OpenAI itself fetched, rendered, or selected a page; it only shows you what is being served to a crawler-style request.

Frequently asked questions

What is the difference between OAI-SearchBot and GPTBot?

OAI-SearchBot crawls pages that may be considered for ChatGPT search features, while GPTBot crawls content that may be used to train OpenAI’s generative AI foundation models, according to OpenAI’s crawler documentation. They are separate crawlers with separate robots.txt groups.

Does allowing OAI-SearchBot guarantee my pages will appear in ChatGPT search results?

No. Allowing OAI-SearchBot only makes a page eligible for consideration. Whether it appears in a specific answer, and whether it drives any traffic, depends on factors OpenAI does not publish, and a passing crawl is not a placement guarantee.

How long does a robots.txt change take to affect OAI-SearchBot?

OpenAI’s crawler documentation states that a robots.txt change can take approximately 24 hours to affect its search systems, so wait at least that long before assuming a fresh policy did not work.

Is ChatGPT-User the same crawler as OAI-SearchBot?

No. ChatGPT-User is a separate, user-initiated fetch agent that retrieves a page when someone asks ChatGPT to visit a specific URL, rather than the crawler that determines ChatGPT search opt-outs. OpenAI notes that robots.txt rules may not apply the same way to that kind of user-initiated action.

Can robots.txt alone keep GPTBot or OAI-SearchBot out of private pages?

No. RFC 9309 explicitly describes robots.txt as a voluntary exclusion mechanism, not an access authorization system. Use authentication, permissions, or a paywall for anything that must actually stay private, and treat robots.txt as a crawling preference for pages that are already public.

Will disallowing GPTBot remove content that was already used for training?

No. A robots.txt change is forward-looking. It has no ability to retroactively withdraw content that a training crawl already collected before the rule was published.

Does opting a page out of OAI-SearchBot remove it from every OpenAI product?

Not necessarily. OpenAI’s publisher FAQ describes a page opted out of OAI-SearchBot as excluded from ChatGPT search answers but still eligible to appear as a navigational link, and it separately describes limited link-and-title surfacing in ChatGPT Atlas when a disallowed URL is learned elsewhere. Treat that Atlas behavior as specific to Atlas rather than universal.

Next steps

You can allow OAI-SearchBot while you block GPTBot in robots.txt, and doing so is a documented, reasonable way to separate ChatGPT search participation from AI training data collection. Publish the policy at the root of every relevant host, verify it against the live file rather than the draft, and remember that eligibility for search consideration is not the same as guaranteed placement or traffic. Revisit the file whenever you add new sections to your site, and recheck it after any migration, since a host, CDN, or platform change can silently reintroduce a sitewide block.

Related tools and resources on AllEasySEO:

Sources and further reading

Try these free SEO tools

Free, browser-based and no signup needed.

See all free SEO tools

Comments

Leave a comment