AI Bot Robots.txt Checker

Enter any website to see which of 2,000+ known AI agents, crawlers, scrapers, and other bots its robots.txt allows or blocks. Find gaps in its rules and get clear recommendations for what to fix.

AI Agent
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI Assistant
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding Agent
Fetches documentation and other resources to help build software
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AI Search Crawler
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Archiver
Archiver
Captures and stores historical website snapshots for long-term digital preservation
Automated Agent
Automated Agent
Automates browser interactions programmatically without direct human supervision
Developer Helper
Developer Helper
Assists with testing, debugging, and ensuring website functionality
Fetcher
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
Intelligence Gatherer
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
Scraper
Scraper
Extracts large amounts of web data, often without explicit website permission
Search Engine Crawler
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
Security Scanner
Security Scanner
Scans websites for security vulnerabilities, threats, and configuration weaknesses
SEO Crawler
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Uncategorized
Uncategorized
Not yet assigned a type
Undocumented AI Agent
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

Known Agents has been featured in

Frequently Asked Questions

Can't bots ignore robots.txt rules?

Technically yes, but a good robots.txt file will solve 90%+ of the problem. You should always use one as your first line of defense.

If a company is small enough to get away with ignoring robots.txt rules, it probably isn't much of a threat anyway, even if it's able to get hold of your data. The vast majority of large companies follow robots.txt rules, including OpenAI, Anthropic, Meta, Amazon, and Microsoft.

The hard part is knowing which of their bots to block. That's where Automatic Robots.txt can help.

Use Agent Analytics to monitor unidentified automated browsers and bots that ignore your rules. You can inspect attributes such as location and IP address, then block them with your firewall.


How does the checker work?

Enter a website and the checker fetches its public robots.txt file. It tests the file's user-agent rules against 2,000+ known agents and bots to determine which are fully blocked from the site, then highlights gaps and recommends rules to review.


Which AI bots and crawlers does it check?

This checker tests your robots.txt against 2,000+ crawlers, scrapers, and AI agents across 15+ categories in the Known Agents directory. This includes AI Data Scrapers, AI Search Crawlers, AI Agents, AI Assistants, Scrapers, SEO Crawlers, Search Engine Crawlers, and more.


Which AI bots and crawlers should I block?

Block bots based on what they do, not simply because they use AI. Consider blocking AI Data Scrapers and AI Data Providers if you don't want your content used for model training or resale. You may also want to block general Scrapers and Undocumented AI Agents when their purpose or benefit is unclear.

Usually allow AI Search Crawlers, AI Assistants, and user-directed AI Agents if you want visibility, citations, and potential customer traffic. Keep traditional Search Engine Crawlers allowed unless you want your pages removed from their search results.


Does blocking AI bots affect my search rankings?

Blocking AI Data Scrapers and other training crawlers does not directly affect your rankings in traditional search engines. Blocking Search Engine Crawlers such as Googlebot can prevent pages from being indexed and cause them to disappear from search results. Blocking AI Search Crawlers can reduce your visibility in AI search products, even if it does not change your traditional rankings.


Will a good robots.txt save me money?

Yes. When bots that follow robots.txt rules see a disallow rule, they stop visiting those pages altogether. Fewer requests mean lower bandwidth, compute, logging, and CDN costs.

A firewall acts only after a request is made. An edge firewall can protect your origin server, but it still has to process each attempt.