Overview

最終更新日:2026-09-10 12:11:28

Reading time: About 5 minutes

What Is AI Crawler Control

As AI technology advances rapidly, an increasing number of AI providers (such as OpenAI, Google, and Anthropic) use crawlers to harvest website content at scale for model training, AI search, and content generation. This AI crawler traffic has become the fastest-growing category of bot traffic, creating challenges for website operators, including rising bandwidth costs and unauthorized use of content.

AI Crawler Control is an integrated governance solution for publicly available, identifiable large language model (LLM) service traffic across the Internet. It combines visual analytics on AI crawler trends with fine-grained, multi-dimensional control capabilities to build a comprehensive AI crawler governance framework built on visibility, analysis, and control. It helps enterprises fully understand and regulate how AI providers crawl and access website content.

Use Cases

AI Crawler Control applies to the following typical scenarios:

  • Copyright protection: Publishers and content creators can monitor and prevent AI crawlers from harvesting original content at scale without authorization.
  • Sensitive information access control: E-commerce and enterprise websites can identify AI crawler activity on product pages and company information, and control access to sensitive data such as pricing and inventory.
  • Traffic cost control: High-frequency AI crawler requests consume bandwidth resources; control effectively reduces unnecessary traffic costs.
  • Brand visibility: For businesses that want their content indexed and cited by LLMs (such as Perplexity and ChatGPT) to increase brand visibility in AI-generated answers.

Key Capabilities

AI Crawler Control provides two core capabilities:

  • AI Crawler Trends (Visibility & Analysis): Provides global visual analysis of AI crawler requests, including total request volume and trends, crawler category distribution, rankings of the most active crawlers, and the most frequently crawled hostnames and paths.
  • AI Crawler Control (Enforcement): Provides fine-grained, hostname-level control over AI crawlers, supporting actions (Not Used/Log/Skip/Deny) by crawler category, service provider, and other dimensions.

Editions

AI Crawler Control is available in Basic and Advanced editions to meet different customer needs:

Key Feature Basic Advanced
Basic AI crawler control: Log, Deny, Skip
AI crawler trends
AI crawler activity monitoring
AI crawler security event and attack log queries

Activation requirements:

  • Basic: If you enable the WAAP Bot Management service, basic AI crawler control is activated automatically.
  • Advanced: In addition to the WAAP Bot Management service, you must separately enable the AI Crawler Control value-added service.

Note: If you need to upgrade to the Advanced edition, contact your account manager to request activation.

Coverage

Currently, the AI Bots library includes more than 40 AI crawlers from major AI providers — including OpenAI, Anthropic, Google, Meta, and Amazon — spanning the following categories:

Crawler Category Description Example
AI Search Crawlers Crawl content in batches on a regular basis to build or update indexes for AI search products, enabling AI search engines to retrieve and rank crawled content. AmazonBot, Perplexity Crawler
AI Training Crawlers Crawl content continuously and at scale for training or fine-tuning large language models (LLMs). The collected data is typically retained long term to enhance AI capabilities. GPTBot, ClaudeBot, Google-Extended Crawler
AI Agents Perform specific tasks on behalf of users in real time and on demand, crawling or accessing specific web pages. The data is generally not retained or used for training; the purpose is to complete user-assigned tasks immediately, rather than to build indexes or train models. ChatGPT-User Crawler, Claude-User Crawler

Inclusion Criteria

An AI crawler must meet all of the following conditions to be included in the AI Bots library:

  • Traceable operator: The operator is a publicly operating LLM or AI search provider with verifiable corporate entity information.
  • Observable traffic: At least one of the following conditions is met.
    • The crawler generates stable traffic across the Internet (an average of more than 1,000 verified requests per day over the past month), to ensure the inclusion criteria apply broadly.
    • Or, even if it has not reached the traffic threshold, its operator is a leading AI provider (such as OpenAI, Anthropic, Google, or Meta AI) that warrants a high level of attention and control.

Removal Criteria:

If an included AI crawler no longer meets the conditions above, it is removed from the AI Bots library. Examples include:

  • Operator no longer traceable: The operator has discontinued the service, the corporate entity has been deregistered, or similar.
  • Prolonged inactivity: No valid traffic has been observed across the Internet for 6 consecutive months.

Special Requests

The AI Bots library is updated periodically to include mainstream AI tools. If you have a special request, contact technical support to request the addition of a specific bot. When submitting a request, we recommend providing the following information: bot name and purpose, a link to official documentation, the User-Agent, the source IP/ASN range, and the desired handling policy.