Agent Directory
The definitive guide to every known AI agent, crawler, scraper, and other bot on the internet, updated daily.
AI Data Providers
DesearchBot
DesearchBot crawls and indexes public web content for Desearch's search engine and search APIs.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
Diffbot
Diffbot is a web crawler that extracts and structures website content using AI-powered visual understanding, providing knowledge graph data for applications like market i…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
ExaBot
ExaBot is a web crawler that indexes web content to power Exa's AI search engine and semantic search APIs for AI applications.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
FirecrawlAgent
FirecrawlAgent is a web crawler operated by Firecrawl that extracts web content and converts it into structured data for use in LLM and AI applications.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
HenkBot
Henkbot crawls the web on behalf of Valyu, an AI search infrastructure provider that indexes content for use in AI-powered retrieval pipelines.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
Keenable-User
Keenable-User accesses web content for Keenable's search and content-access infrastructure for AI agents and AI labs.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
KeenableBot
KeenableBot crawls web pages for Keenable's independent search index. Keenable provides this index through search and content-retrieval APIs used by AI agents, AI labs, a…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
Mozilla-Tabstack
Mozilla-Tabstack is an AI agent operated by Mozilla that performs programmatic, AI-driven interactions with web content through Tabstack.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
Querit-SearchBot
Querit-SearchBot is a web crawler operated by Querit that indexes web content for their search API service, which is designed to provide real-time search results for larg…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
QueritBot
QueritBot is a web crawler operated by Querit, a company providing a search API for large language model integration. This bot indexes web content to power the real-time …
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
ScryBot
ScryBot collects public web content for Scry, Unflatten's programmable search platform. Scry lets applications query and analyze an index spanning web pages, social conte…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
ShapBot
ShapBot is a web crawler by Parallel that collects and structures web content to power its search, extraction, and deep research APIs, providing AI agents with high-accur…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
SofyaBot
SofyaBot crawls public web pages to build the independent search index behind Sofya's API, which gives AI agents access to current web information.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
TalarionFrontierCrawler
TalarionFrontierCrawler is a web crawler operated by Talarion that collects public web content for its research corpus and AI knowledge services.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
TavilyBot
TavilyBot is a web crawler by Tavily that indexes and extracts content from billions of pages, providing real-time search, extraction, and research data to ground AI agen…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
Terra Cotta
Terra Cotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
TerraCotta
TerraCotta is Ceramic's web crawler that indexes public content to power their web-scale search API for AI and LLMs.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
YouBot
YouBot is a web crawler by You.com that indexes and extracts web content to power its real-time search, contents, and research APIs, delivering grounded web data to AI ag…
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
See More →
AI Data Scrapers
AI2Bot
AI2Bot is operated by Ai2, a non-profit AI research institute. It's used to download data to train open source AI models.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Ai2Bot-Dolma
Ai2Bot-Dolma is operated by Ai2, a non-profit AI research institute. It's used to download data to train open source AI models.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Amazonbot
Amazonbot is a web crawler operated by Amazon that builds an index of web content to improve Amazon products and services, including Alexa, Kindle, and Amazon Shopping. T…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Applebot-Extended
Apple-Extended is used to train Apple’s foundation LLM models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
BeansLLM-CorpusBot
BeansLLM-CorpusBot collects public web content to build corpora for language-model training.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
bedrockbot
bedrockbot is a web crawler operated by Amazon that crawls selected URLs for Amazon Bedrock knowledge bases and custom AI applications.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Bytespider
Bytespider is a web crawler operated by ByteDance, the Chinese owner of TikTok. It's allegedly used to download training data for its LLMs (Large Language Model) includin…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
CCBot
CCBot is Common Crawl's web crawler that creates an open repository of web data, making crawled content universally accessible for research, analysis, and AI training pur…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
ChatGLM-Spider
ChatGLM-Spider is a web crawler operated by Zhipu AI, the Chinese company behind ChatGLM. It is used for collecting data to train and evaluate the company's large languag…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
ClaudeBot
ClaudeBot is a web crawler operated by Anthropic to download training data for its LLMs (Large Language Models) that power AI products like Claude.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
CloudVertexBot
CloudVertexBot is a Google-operated crawler available to site owners to request targeted crawls of their own sites for AI training purposes on the Vertex AI platform.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
cohere-training-data-crawler
cohere-training-data-crawler is a web crawler operated by Cohere to download training data for its LLMs (Large Language Models) that power its enterprise AI products.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
CohereBot
CohereBot is a web crawler operated by Cohere, a company that provides large language models and enterprise solutions. This bot collects data from websites to support the…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Cotoyogi
Cotoyogi is a research crawler operated by Japan's Research Organization of Information and Systems that collects web content to build AI training datasets for research p…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Datenbank Crawler
Datenbank Crawler is a web crawler operated by German company netEstate used for collecting and selling international website data.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
DeepSeekBot
DeepSeekBot is a web crawler operated by DeepSeek that downloads website content for AI model training and other data collection use cases.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Doubaobot
Doubaobot is a web crawler operated by ByteDance that downloads website content to support Doubao training and generated answers.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
ERNIEBot
ERNIEBot is a web crawler operated by Baidu that downloads website content to support ERNIE training and generated answers across Baidu services.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
FacebookBot
FacebookBot is a web crawler used by Meta to download training data for its AI speech recognition technology.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Google-Extended
Google-Extended is a web crawler used by Google to download AI training content for its AI products like the Gemini assistant and its Vertex AI generative APIs.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
GoogleOther
GoogleOther is Google's generic crawler used by various product teams for fetching publicly accessible content, including one-off crawls for internal research and develop…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
GPTBot
GPTBot is OpenAI's web crawler that collects data from publicly accessible web pages to improve AI models like ChatGPT, while respecting robots.txt and opt-out preference…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Hunyuan
Hunyuan crawls public web content for Tencent's Hunyuan foundation models and AI applications.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
ICC-Crawler
ICC-Crawler is NICT's research crawler that automatically collects web pages from the Internet for academic research at Japan's National Institute of Information and Comm…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
imageSpider
imageSpider is a web crawler operated by ByteDance, the company behind TikTok, Douyin, and other content platforms. The bot collects images from websites across the inter…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
Kangaroo Bot
Kangaroo Bot is used by the company Kangaroo LLM to download data to train open source AI models tailored to Australian language and culture.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
KimiBot
KimiBot is a web crawler operated by Moonshot AI that collects web content for training Kimi's AI models. Sites can block this bot via robots.txt to prevent their content…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
KimiCrawler
KimiCrawler is a web crawler operated by Moonshot AI that downloads website content for Kimi and related Moonshot AI services.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
laion-huggingface-processor
LAION-huggingface-processor is a web crawler operated by LAION (Large-scale Artificial Intelligence Open Network), a non-profit organization that creates open datasets fo…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
LCC
LCC is a web crawler operated by the University of Leipzig that collects text data from websites to build large-scale linguistic corpora for research purposes. The bot ga…
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
meta-externalagent
meta-externalagent crawls web content for training AI models and improving Meta's products by indexing content directly across the internet.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →
micro-crawl
micro-crawl is a web crawler operated by Reflection AI that downloads website content for AI research and model development.
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
See More →