We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic. By clicking “Accept,” you agree to our website's cookie use as described in our Cookie Policy. You can change your cookie settings at any time by clicking “Preferences.”
    AI is already influencing your customers. See what it recommends →
    Research

    How 5 AI Platforms Actually Crawl Brand Data: What 111,286 Requests Taught Us

    We operate one of the largest AI brand directories on the internet. This week, AI model crawlers accounted for 17.5% of all traffic. Here's exactly what each platform is doing.

    111,286

    Total requests/week

    19,454

    AI crawler requests

    17.5%

    AI share of traffic

    5,829

    Brands in directory

    OpenAI

    10,816 requests (55.6%)
    Crawler Requests What It Does
    GPTBot/1.3 8,159 Training/indexing crawler — builds ChatGPT's parametric 'memory'
    OAI-SearchBot/1.3 1,250 Real-time search crawler — fetches data when ChatGPT searches the web
    ChatGPT-User/1.0 515 Live browsing — a human asked ChatGPT to 'look up' a specific page
    OAI-SearchBot/1.0 441 Older search variant, still active
    GPTBot/1.0 10 Legacy version, phasing out

    OpenAI operates three distinct crawling modes, and your brand needs to serve accurate data to all three.

    GPTBot builds parametric knowledge — the 'memory' that ChatGPT uses when it doesn't search the web. If your brand information is wrong on pages GPTBot crawls, ChatGPT will confidently state wrong information even when it's not searching.

    OAI-SearchBot retrieves real-time data. When someone asks ChatGPT a question and it hits the search button, this crawler fetches the results. This is your opportunity for accuracy — real-time retrieval can override stale parametric knowledge.

    ChatGPT-User fires when a human asks ChatGPT to 'look up' or 'visit' a specific URL. 515 times this week, someone asked ChatGPT to go check a brand profile on our directory. That's a direct signal of buyer behavior.

    What this means for your brand: If you're optimizing for ChatGPT, focus on the pages GPTBot crawls most — your homepage, product pages, and structured data. These build the parametric knowledge that persists across conversations. Then ensure your real-time information (pricing, product updates, recent case studies) is accessible to OAI-SearchBot on pages that update frequently.

    Anthropic (Claude)

    4,669 requests (24.0%)
    Crawler Requests What It Does
    ClaudeBot/1.0 4,235 Main indexing crawler
    Claude-User 303 Live browsing — someone asked Claude to check a page
    Claude-SearchBot/1.0 99 Search-specific crawler
    Claude-User/1.0 32 Authenticated browsing variant

    Anthropic's crawling volume is about 43% of OpenAI's — which aligns with our AI representation data showing that Claude tends to have less comprehensive brand knowledge than ChatGPT.

    The ClaudeBot crawler does the heavy lifting, but the Claude-User traffic (335 requests) is the most interesting signal: real people are asking Claude about brands, and Claude is visiting our directory to answer them.

    What this means for your brand: Claude's lower crawl volume means your parametric representation may lag behind ChatGPT's by weeks. If you make a major positioning change, don't wait for ClaudeBot to discover it organically — ensure your Crunchbase, Wikipedia, and other high-authority sources are updated, because Claude's training pipeline weights these sources heavily.

    Amazon

    4,366 requests (22.4%)
    Crawler Requests What It Does
    Amzn-SearchBot/0.1 2,885 Amazon's AI search crawler
    Amazonbot/0.1 1,481 General web crawler

    Amazon runs two separate bots with nearly equal presence. This matters because Amazon's AI assistants — Alexa and Rufus — use this data for product and brand recommendations.

    If you sell to enterprises, you might think Amazon doesn't matter. But Rufus is increasingly used for B2B product research, and your brand data feeds into that recommendation engine.

    What this means for your brand: If your product has any e-commerce, marketplace, or product comparison dimension, Amazon's crawl data feeds into Rufus and Alexa recommendations. Treat your Amazon-relevant pages (product pages, comparison content, pricing) as high-priority for structured data accuracy.

    Perplexity

    1,699 requests (8.7%)
    Crawler Requests What It Does
    PerplexityBot/1.0 1,699 Search-and-answer crawler

    Perplexity is the purest signal in this data. Every single Perplexity crawl is in direct service of answering a real user query. When PerplexityBot visits a brand profile, it's because someone asked Perplexity a question about that brand or category.

    1,699 times this week, someone asked Perplexity a question that led to our directory. These are the highest-intent requests because each one maps to a real human asking a real question.

    What this means for your brand: Since every Perplexity crawl answers a real user query, you can think of PerplexityBot traffic as a proxy for 'how often real people are asking AI about brands in your space.' Our 1,699 requests/week translates to roughly 240 real-person brand queries per day being answered partly by our directory data.

    ByteDance

    15 requests (0.1%)
    Crawler Requests What It Does
    Bytespider 15 TikTok/Doubao crawler

    Minimal presence today, but worth watching. ByteDance's Bytespider is the crawler behind TikTok's AI features and their Doubao assistant.

    15 requests is noise, but it's noise that suggests TikTok's AI is starting to index brand data. If your audience skews younger or if TikTok is a channel for your market, this will matter within 12 months.

    What this means for your brand: 15 requests is a baseline, not a signal to ignore. TikTok is increasingly used for product research, especially in B2C and SMB-focused categories. If Bytespider's crawl volume increases over the next 3-6 months, it could signal TikTok's AI features gaining meaningful brand knowledge.

    The Bigger Picture

    Zooming out from AI crawlers, here's the full traffic landscape for a brand directory serving 5,829 profiles:

    Traffic Source Requests/Week Key Crawlers
    AI model crawlers 19,454 GPTBot, ClaudeBot, PerplexityBot, Amazonbot
    SEO tool crawlers 10,936 SemrushBot, MJ12bot, SERankingBot, AhrefsBot
    Search engine crawlers 6,730 Googlebot (4,530), YandexBot (1,403), bingbot (768)
    Human traffic (estimated) ~13,000 Direct, referral, organic

    The ratio that matters:

    AI crawlers are hitting our directory at nearly 3× the rate of search engine crawlers. AI models are consuming brand data faster and more aggressively than Google. This is the clearest infrastructure-level signal that AI is becoming a primary consumer of brand information — not a secondary channel.

    How Crawl Frequency Correlates with AI Accuracy

    An obvious question: do brands that get crawled more frequently have higher AI representation scores?

    From our directory data: yes, but with important caveats. Brands with strong structured data and consistent authoritative sources tend to get crawled more frequently — AI models learn that these pages are reliable and revisit them more often. It's a virtuous cycle: accuracy drives crawl frequency, which drives more up-to-date parametric knowledge, which drives accuracy.

    The inverse is also true. Phantom brands (AI representation 0–19) tend to have the lowest crawl frequencies. AI models never formed a strong entity representation, so they don't prioritize re-crawling those pages.

    The actionable insight: if you can get your page right once — accurate structured data, correct industry classification, consistent messaging — the AI crawlers will return often enough to keep it current. The hard part is the initial correction. The maintenance is largely self-sustaining.

    What This Means for Your robots.txt

    If you take one action from this report, audit your robots.txt for AI crawler permissions. Here's the minimum configuration we recommend:

    User-agent: GPTBot

    Allow: /

    User-agent: OAI-SearchBot

    Allow: /

    User-agent: ChatGPT-User

    Allow: /

    User-agent: ClaudeBot

    Allow: /

    User-agent: PerplexityBot

    Allow: /

    User-agent: Amazonbot

    Allow: /

    We've seen brands unknowingly blocking AI crawlers because their IT team added broad bot-blocking rules. One brand in our directory went from Phantom (AI representation 3) to Challenger (AI representation 54) in 3 weeks — the only change was fixing their robots.txt. That's how impactful discoverability is. For the full implementation guide, see our AI-Friendly robots.txt Guide.

    What's Coming Next

    This is the first installment in what will be a monthly series. Future reports will include:

    • Month-over-month crawl volume trends per platform
    • New crawler variants detected (AI companies frequently deploy new bot versions)
    • Correlation analysis between crawl frequency and AI representation score changes
    • Category-level crawl patterns (which brand categories do AI crawlers visit most?)

    Bookmark this page or read our full March 2026 report — the longitudinal data gets more valuable with every month.

    Related

    This is the first in our monthly "State of AI Brand Crawling" series. Read the full March 2026 report →

    Data source: Cloudflare server logs from Optimly's AI Brand Index infrastructure, week of March 22–28, 2026. All request counts are approximate and based on user-agent classification. Some crawler variants may be underrepresented if they use non-standard user-agent strings.

    Previously published under the name "AI Brand Directory" — renamed to "AI Brand Index" in 2026.