Control AI crawler access

GUIDE / PROTOCOL

Control AI crawler access

A clear access policy starts with the intended use of each public page rather than a list of crawler names.

Direct answer

How should AI-related crawlers be controlled?

Decide which pages may be discovered, which may appear in search and which should be excluded from a specific use. Express those choices through robots.txt, page directives and server responses, then test what a crawler can actually read.

01 / Review method

Review method

01

Inventory the surfaces

Separate public pages, resources needed for rendering, private areas and variants without independent value.

02

Choose by use

Decide access for conventional search, AI search and training collection using the crawler identities documented by each operator.

03

Inspect the response

Test HTTP status, robots.txt, redirects, robots meta or X-Robots-Tag and served HTML, including language versions.

04

Observe carefully

Compare server logs, webmaster tools and referral traffic. A crawler visit proves neither indexing nor citation.

02 / Evidence to keep

Evidence to keep

Dated rule

Keep the version, date and reason for each robots directive.

Tested URL

Record the exact page, status, canonical and page-level rules.

Identified agent

Separate declared search, preview and training crawlers.

Outcome check

Document what remains accessible and what actually appears in the available tools.

03 / FAQ

Frequently asked questions

Does allowing OAI-SearchBot guarantee a citation?

No. It enables access needed for possible appearance in ChatGPT search but does not guarantee use of a page.

Can search and training be treated separately?

Yes when an operator documents separate agents. Verify names and behaviour in official documentation before changing the rules.

Describe a context