What is ApifyWebsiteContentCrawler?
ApifyWebsiteContentCrawler is a web crawler by Apify that extracts and downloads full website content for use in AI, data analysis, and automation workflows. Agent Analytics can track when it visits your website.
Overview
| Operated By | Apify |
| Source | Official Website |
| Expected To Follow Robots.txt | Yes |
| Insights Last Updated | August 11, 2026 |
Do you operate this agent? Contact us to suggest an update.
Category
Expected Behavior
ApifyWebsiteContentCrawler crawls systematically at high volume. Expect recurring passes that cover large portions of your site rather than isolated pages, with sudden bursts when demand increases on its end.
How To Block ApifyWebsiteContentCrawler With Robots.txt
Add this rule to your robots.txt file to block ApifyWebsiteContentCrawler from accessing your entire website, or use Automatic Robots.txt to block all AI data providers at once. You can customize which pages are blocked by swapping out / for a different path.
User-agent: ApifyWebsiteContentCrawler # https://knownagents.com/agents/apifywebsitecontentcrawler
Disallow: /
Global Statistics for ApifyWebsiteContentCrawler
As of August 11, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the top 1000 websites and their robots.txt files.
Robots.txt Blocked Percentage
Country of Origin
Robots.txt Blocking Trend
1% of top websites block ApifyWebsiteContentCrawler in their robots.txt files.
Overall AI Data Provider Traffic
0.3% of all web traffic came from AI data providers.
Frequently Asked Questions
Should I Block ApifyWebsiteContentCrawler?
It depends on how you feel about redistribution. A crawl from ApifyWebsiteContentCrawler can supply your content to multiple AI companies for training, search, or retrieval. Allowing it may expand your visibility across third-party AI products, while blocking it limits that reach without affecting traditional search rankings. For context, 1% of the top websites we track currently have robots.txt rules for ApifyWebsiteContentCrawler.
Does ApifyWebsiteContentCrawler Follow Robots.txt Rules?
Yes. ApifyWebsiteContentCrawler is expected to follow robots.txt directives, so a disallow rule is the appropriate first step. Automatic Robots.txt can add and maintain that rule, while Agent Analytics lets you verify whether ApifyWebsiteContentCrawler respects it.
Does ApifyWebsiteContentCrawler Access Private Content?
No special access. ApifyWebsiteContentCrawler can reach public pages, and some providers use proxy networks that can bypass rate limits or geographic restrictions. Protect sensitive content with authentication rather than relying on those softer barriers.
Why Is ApifyWebsiteContentCrawler Visiting My Website?
One of Apify's customers may have requested data from your site, or your pages may be part of ApifyWebsiteContentCrawler's standing index. A single fetch can support multiple downstream AI applications.
How Can I Tell if ApifyWebsiteContentCrawler Is Visiting My Website?
Agent Analytics tracks visits from ApifyWebsiteContentCrawler and every other known AI agent, crawler, and scraper, then authenticates each visit using ApifyWebsiteContentCrawler's published verification method. You can also check your server logs for requests whose user-agent string contains "ApifyWebsiteContentCrawler". Look for long sequences of requests spanning several unrelated sections of your site. A matching log entry alone is not proof because any bot can claim to be ApifyWebsiteContentCrawler.