How 5 AI Platforms Actually Crawl Brand Data: What 111,286 Requests Taught Us
111,286
Total requests/week
19,454
AI crawler requests
17.5%
AI share of traffic
5,829
Brands in directory
OpenAI
10,816 requests (55.6%)| Crawler | Requests | What It Does |
|---|---|---|
| GPTBot/1.3 | 8,159 | Training/indexing crawler — builds ChatGPT's parametric 'memory' |
| OAI-SearchBot/1.3 | 1,250 | Real-time search crawler — fetches data when ChatGPT searches the web |
| ChatGPT-User/1.0 | 515 | Live browsing — a human asked ChatGPT to 'look up' a specific page |
| OAI-SearchBot/1.0 | 441 | Older search variant, still active |
| GPTBot/1.0 | 10 | Legacy version, phasing out |
OpenAI operates three distinct crawling modes, and your brand needs to serve accurate data to all three.
GPTBot builds parametric knowledge — the 'memory' that ChatGPT uses when it doesn't search the web. If your brand information is wrong on pages GPTBot crawls, ChatGPT will confidently state wrong information even when it's not searching.
OAI-SearchBot retrieves real-time data. When someone asks ChatGPT a question and it hits the search button, this crawler fetches the results. This is your opportunity for accuracy — real-time retrieval can override stale parametric knowledge.
ChatGPT-User fires when a human asks ChatGPT to 'look up' or 'visit' a specific URL. 515 times this week, someone asked ChatGPT to go check a brand profile on our directory. That's a direct signal of buyer behavior.
What this means for your brand: If you're optimizing for ChatGPT, focus on the pages GPTBot crawls most — your homepage, product pages, and structured data. These build the parametric knowledge that persists across conversations. Then ensure your real-time information (pricing, product updates, recent case studies) is accessible to OAI-SearchBot on pages that update frequently.
Anthropic (Claude)
4,669 requests (24.0%)| Crawler | Requests | What It Does |
|---|---|---|
| ClaudeBot/1.0 | 4,235 | Main indexing crawler |
| Claude-User | 303 | Live browsing — someone asked Claude to check a page |
| Claude-SearchBot/1.0 | 99 | Search-specific crawler |
| Claude-User/1.0 | 32 | Authenticated browsing variant |
Anthropic's crawling volume is about 43% of OpenAI's — which aligns with our AI representation data showing that Claude tends to have less comprehensive brand knowledge than ChatGPT.
The ClaudeBot crawler does the heavy lifting, but the Claude-User traffic (335 requests) is the most interesting signal: real people are asking Claude about brands, and Claude is visiting our directory to answer them.
What this means for your brand: Claude's lower crawl volume means your parametric representation may lag behind ChatGPT's by weeks. If you make a major positioning change, don't wait for ClaudeBot to discover it organically — ensure your Crunchbase, Wikipedia, and other high-authority sources are updated, because Claude's training pipeline weights these sources heavily.
Amazon
4,366 requests (22.4%)| Crawler | Requests | What It Does |
|---|---|---|
| Amzn-SearchBot/0.1 | 2,885 | Amazon's AI search crawler |
| Amazonbot/0.1 | 1,481 | General web crawler |
Amazon runs two separate bots with nearly equal presence. This matters because Amazon's AI assistants — Alexa and Rufus — use this data for product and brand recommendations.
If you sell to enterprises, you might think Amazon doesn't matter. But Rufus is increasingly used for B2B product research, and your brand data feeds into that recommendation engine.
What this means for your brand: If your product has any e-commerce, marketplace, or product comparison dimension, Amazon's crawl data feeds into Rufus and Alexa recommendations. Treat your Amazon-relevant pages (product pages, comparison content, pricing) as high-priority for structured data accuracy.
Perplexity
1,699 requests (8.7%)| Crawler | Requests | What It Does |
|---|---|---|
| PerplexityBot/1.0 | 1,699 | Search-and-answer crawler |
Perplexity is the purest signal in this data. Every single Perplexity crawl is in direct service of answering a real user query. When PerplexityBot visits a brand profile, it's because someone asked Perplexity a question about that brand or category.
1,699 times this week, someone asked Perplexity a question that led to our directory. These are the highest-intent requests because each one maps to a real human asking a real question.
What this means for your brand: Since every Perplexity crawl answers a real user query, you can think of PerplexityBot traffic as a proxy for 'how often real people are asking AI about brands in your space.' Our 1,699 requests/week translates to roughly 240 real-person brand queries per day being answered partly by our directory data.
ByteDance
15 requests (0.1%)| Crawler | Requests | What It Does |
|---|---|---|
| Bytespider | 15 | TikTok/Doubao crawler |
Minimal presence today, but worth watching. ByteDance's Bytespider is the crawler behind TikTok's AI features and their Doubao assistant.
15 requests is noise, but it's noise that suggests TikTok's AI is starting to index brand data. If your audience skews younger or if TikTok is a channel for your market, this will matter within 12 months.
What this means for your brand: 15 requests is a baseline, not a signal to ignore. TikTok is increasingly used for product research, especially in B2C and SMB-focused categories. If Bytespider's crawl volume increases over the next 3-6 months, it could signal TikTok's AI features gaining meaningful brand knowledge.
The Bigger Picture
Zooming out from AI crawlers, here's the full traffic landscape for a brand directory serving 5,829 profiles:
| Traffic Source | Requests/Week | Key Crawlers |
|---|---|---|
| AI model crawlers | 19,454 | GPTBot, ClaudeBot, PerplexityBot, Amazonbot |
| SEO tool crawlers | 10,936 | SemrushBot, MJ12bot, SERankingBot, AhrefsBot |
| Search engine crawlers | 6,730 | Googlebot (4,530), YandexBot (1,403), bingbot (768) |
| Human traffic (estimated) | ~13,000 | Direct, referral, organic |
The ratio that matters:
AI crawlers are hitting our directory at nearly 3× the rate of search engine crawlers. AI models are consuming brand data faster and more aggressively than Google. This is the clearest infrastructure-level signal that AI is becoming a primary consumer of brand information — not a secondary channel.
How Crawl Frequency Correlates with AI Accuracy
An obvious question: do brands that get crawled more frequently have higher AI representation scores?
From our directory data: yes, but with important caveats. Brands with strong structured data and consistent authoritative sources tend to get crawled more frequently — AI models learn that these pages are reliable and revisit them more often. It's a virtuous cycle: accuracy drives crawl frequency, which drives more up-to-date parametric knowledge, which drives accuracy.
The inverse is also true. Phantom brands (AI representation 0–19) tend to have the lowest crawl frequencies. AI models never formed a strong entity representation, so they don't prioritize re-crawling those pages.
The actionable insight: if you can get your page right once — accurate structured data, correct industry classification, consistent messaging — the AI crawlers will return often enough to keep it current. The hard part is the initial correction. The maintenance is largely self-sustaining.
What This Means for Your robots.txt
If you take one action from this report, audit your robots.txt for AI crawler permissions. Here's the minimum configuration we recommend:
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Amazonbot
Allow: /
We've seen brands unknowingly blocking AI crawlers because their IT team added broad bot-blocking rules. One brand in our directory went from Phantom (AI representation 3) to Challenger (AI representation 54) in 3 weeks — the only change was fixing their robots.txt. That's how impactful discoverability is. For the full implementation guide, see our AI-Friendly robots.txt Guide.
What's Coming Next
This is the first installment in what will be a monthly series. Future reports will include:
- • Month-over-month crawl volume trends per platform
- • New crawler variants detected (AI companies frequently deploy new bot versions)
- • Correlation analysis between crawl frequency and AI representation score changes
- • Category-level crawl patterns (which brand categories do AI crawlers visit most?)
Bookmark this page or read our full March 2026 report — the longitudinal data gets more valuable with every month.
Related
This is the first in our monthly "State of AI Brand Crawling" series. Read the full March 2026 report →
Data source: Cloudflare server logs from Optimly's AI Brand Index infrastructure, week of March 22–28, 2026. All request counts are approximate and based on user-agent classification. Some crawler variants may be underrepresented if they use non-standard user-agent strings.
Previously published under the name "AI Brand Directory" — renamed to "AI Brand Index" in 2026.
