最終更新日:2026-09-10 12:05:16
Reading time: About 5 minutes
As AI technology advances rapidly, an increasing number of AI providers (such as OpenAI, Google, and Anthropic) use crawlers to harvest website content at scale for model training, AI search, and content generation. This AI crawler traffic has become the fastest-growing category of bot traffic, creating challenges for website operators, including rising bandwidth costs and unauthorized use of content.
AI Crawler Control is an integrated governance solution for publicly available, identifiable large language model (LLM) service traffic across the Internet. It combines visual analytics on AI crawler trends with fine-grained, multi-dimensional control capabilities to build a comprehensive AI crawler governance framework built on visibility, analysis, and control. It helps enterprises fully understand and regulate how AI providers crawl and access website content.
AI Crawler Control applies to the following typical scenarios:
AI Crawler Control provides two core capabilities:
AI Crawler Control is available in Basic and Advanced editions to meet different customer needs:
| Key Feature | Basic | Advanced |
|---|---|---|
| Basic AI crawler control: Log, Deny, Skip | ✔ | ✔ |
| AI crawler trends | ✘ | ✔ |
| AI crawler activity monitoring | ✘ | ✔ |
| AI crawler security event and attack log queries | ✔ | ✔ |
Activation requirements:
Note: If you need to upgrade to the Advanced edition, contact your account manager to request activation.
Currently, the AI Bots library includes more than 40 AI crawlers from major AI providers — including OpenAI, Anthropic, Google, Meta, and Amazon — spanning the following categories:
| Crawler Category | Description | Example |
|---|---|---|
| AI Search Crawlers | Crawl content in batches on a regular basis to build or update indexes for AI search products, enabling AI search engines to retrieve and rank crawled content. | AmazonBot, Perplexity Crawler |
| AI Training Crawlers | Crawl content continuously and at scale for training or fine-tuning large language models (LLMs). The collected data is typically retained long term to enhance AI capabilities. | GPTBot, ClaudeBot, Google-Extended Crawler |
| AI Agents | Perform specific tasks on behalf of users in real time and on demand, crawling or accessing specific web pages. The data is generally not retained or used for training; the purpose is to complete user-assigned tasks immediately, rather than to build indexes or train models. | ChatGPT-User Crawler, Claude-User Crawler |
If an included AI crawler no longer meets the conditions above, it is removed from the AI Bots library. Examples include:
The AI Bots library is updated periodically to include mainstream AI tools. If you have a special request, contact technical support to request the addition of a specific bot. When submitting a request, we recommend providing the following information: bot name and purpose, a link to official documentation, the User-Agent, the source IP/ASN range, and the desired handling policy.