What is Diffbot?

Diffbot is a web crawler that extracts and structures website content using AI-powered visual understanding, providing knowledge graph data for applications like market intelligence, e-commerce, and AI model training. Agent Analytics can track when it visits your website.

Overview

Operated By Diffbot
Source Official Website
Expected To Follow Robots.txt Yes
Insights Last Updated August 11, 2026

Do you operate this agent? Contact us to suggest an update.

Category

AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service

Expected Behavior

Diffbot crawls systematically at high volume. Expect recurring passes that cover large portions of your site rather than isolated pages, with sudden bursts when demand increases on its end.

Diffbot's User Agent

User Agent Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.2751.93 Safari/537.36 Diffbot-User/0.1 (+http://www.diffbot.com)

How To Block Diffbot With Robots.txt

Add this rule to your robots.txt file to block Diffbot from accessing your entire website, or use Automatic Robots.txt to block all AI data providers at once. You can customize which pages are blocked by swapping out / for a different path.

User-agent: Diffbot # https://knownagents.com/agents/diffbot
Disallow: /

Global Statistics for Diffbot

As of August 11, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the top 1000 websites and their robots.txt files.

Robots.txt Blocked Percentage

14%
14% of top websites are blocking Diffbot
Learn How →

Country of Origin

United States
Diffbot normally visits From the United States

Robots.txt Blocking Trend

14% of top websites block Diffbot in their robots.txt files.

Overall AI Data Provider Traffic

0.3% of all web traffic came from AI data providers.

Top Visited Website Categories

Health
Internet and Telecom
Computers and Electronics
Business and Industrial
Science

The types of websites most frequently visited by Diffbot.

Frequently Asked Questions

Should I Block Diffbot?

It depends on how you feel about redistribution. A crawl from Diffbot can supply your content to multiple AI companies for training, search, or retrieval. Allowing it may expand your visibility across third-party AI products, while blocking it limits that reach without affecting traditional search rankings. For context, 14% of the top websites we track currently have robots.txt rules for Diffbot.


Does Diffbot Follow Robots.txt Rules?

Yes. Diffbot is expected to follow robots.txt directives, so a disallow rule is the appropriate first step. Automatic Robots.txt can add and maintain that rule, while Agent Analytics lets you verify whether Diffbot respects it.


Does Diffbot Access Private Content?

No special access. Diffbot can reach public pages, and some providers use proxy networks that can bypass rate limits or geographic restrictions. Protect sensitive content with authentication rather than relying on those softer barriers.


Why Is Diffbot Visiting My Website?

One of Diffbot's customers may have requested data from your site, or your pages may be part of Diffbot's standing index. A single fetch can support multiple downstream AI applications.


How Can I Tell if Diffbot Is Visiting My Website?

Agent Analytics tracks Diffbot visits in real time alongside every other known AI agent, crawler, and scraper. You can also check your server logs for requests whose user-agent string contains "Diffbot". Look for long sequences of requests spanning several unrelated sections of your site. Because Diffbot does not publish a verification method, any client can claim its identity and a log match is only a clue.