PINGDOM_CHECK

#ExtractSummit2026 The world's largest web scraping conference returns. Austin Oct 7–8 · Dublin Nov 10–11

Register now
Data Services
Login
Try Zyte APIContact Sales
  • Unblocking and Extraction

    Zyte API

    The ultimate API for web scraping. Avoid website bans and access a headless browser or AI Parsing

    Ban Handling

    Headless Browser

    AI Extraction

    SERP

    Enterprise

    DocumentationSupport

    Hosting and Deployment

    Scrapy Cloud

    Run, monitor, and control your Scrapy spiders however you want to.

    Coding Agent Add-Ons

    Agentic Web Data

    Plugins that give coding agents the context to build production Scrapy projects. Starts with Claude Code.

  • Data Services
  • Zyte API

    Zyte Data

    Scrapy Cloud

  • Browse

    • BlogArticles, podcasts, videos
    • Case studiesCustomer outcomes
    • White papersIn-depth reports
    • DocumentationGuides & API reference
    • EventsConferences, webinars, recordings

    Subscribe

    • NewsletterSwiftly delivered
    • Join our community2,000+ web scraping engineers
  • Product and E-commerce

    From e-commerce and online marketplaces

    Data for AI

    Collect and structure web data to feed AI

    Job Posting

    From job boards and recruitment websites

    Real Estate

    From Listings portals and specialist websites

    News and Article

    From online publishers and news websites

    Search

    Search engine results page data (SERP)

    Social Media

    From social media platforms online

  • Meet Zyte

    Our story, people and values

    Contact us

    Get in touch

    Support

    Knowledge base and raise support tickets

    Terms and Policies

    Accept our terms and policies

    Open Source

    Our open source projects and contributions

    Web Data Compliance

    Guidelines and resources for compliant web data collection

    Affiliate Program

    Join Zyte’s affiliate program and start earning commissions today

    Join the team building the future of web data
    We're Hiring
    Trust Center
    Security, compliance & certifications
Login
Try Zyte APIContact Sales
All articles
AI75, 75 articles
Data quality15, 15 articles
Developer interest60, 60 articles
Integration3, 3 articles
Open-source50, 50 articles
Proxies35, 35 articles
Scraping practice35, 35 articles
Scraping strategy48, 48 articles
Search results4, 4 articles
Web data75, 75 articles
Web scraping APIs49, 49 articles
Scrapy47, 47 articles
Scrapy Cloud26, 26 articles
Web Scraping Copilot11, 11 articles
Zyte API71, 71 articles
AI & Machine Learning3, 3 articles
Automotive3, 3 articles
E-commerce & retail35, 35 articles
Entertainment & Streaming2, 2 articles
Financial Services8, 8 articles
Government2, 2 articles
Market Research & Intelligence7, 7 articles
Media & publishing11, 11 articles
Real Estate2, 2 articles
Recruitment & HR3, 3 articles
Transportation & Logistics2, 2 articles
Travel & hospitality3, 3 articles
iPaaS2, 2 articles
Large language model29, 29 articles
MCP3, 3 articles
Python110, 110 articles
Scraping at Scale7, 7 articles
Scraping Fundamentals11, 11 articles
Web Scraping Industry Report20, 20 articles

Appearance

Discord Community
BlogData gathering for AIFour in 10 sites block AI bots with robots.txt
ArticleResearch / ReportData gathering for AILarge Language Models (LLMs)Access handlingFuture of the web

Four in 10 sites block AI bots with robots.txt

Nearly four in 10 of the world's top sites are now closed to well-behaved AI crawlers. Here's what operators actually do with robots.txt, and which AI agents they block or welcome.

Robert Andrews · Senior editor

September 7, 2026

Four in 10 sites block AI bots with robots.txt

The debate about AI companies and web content has generated an enormous amount of heat.

Publishers say AI companies are taking their content without compensation. AI companies say they are indexing the public web as search engines always have. Lawyers are arguing about what training data is, and what robots.txt actually means in a legal context.

Zyte's State of Web Access 2026 does not address those questions. But it does offer something the debate has largely lacked: a ground-level measurement of what site operators are actually doing about it, across 11,100 of the world's most popular landing pages.

How sites block AI with a text file

How sites handle AI crawlers in robots.txt: 12.7% name one to block, 27.2% block all traffic by wildcard, 1.8% explicitly allow

The headline numbers?

  • 12.7% of all sites now name at least one named AI crawler in a Disallow rule - an explicit, named policy decision targeting AI access specifically.
  • A further 27.2% block all automated traffic via wildcard, catching AI crawlers by default without naming them.
  • Only 1.8% actively welcome AI with explicit Allow rules.

The data reveals not a uniform response but a fractured one - with clear distinctions between training crawlers and search-linked agents, between content industries and service industries, and between blocking and welcoming that sometimes coexist on the same site.

GPTBot leads the block list, but the ratios tell a different story

Among the eleven named AI agents in our dataset, GPTBot - OpenAI's primary training crawler - faces the most named Disallow rules: 8.4% of all 11,100 sites.

Named AI crawlers by Disallow and Allow rate, with GPTBot facing the most blocks at 8.4% of sites

CCBot, operated by Common Crawl and used historically as a training data source, sits at 7.3%.

ClaudeBot (Anthropic) is at 6.3%, Google-Extended at 5.6%, and Bytespider (ByteDance) at 5.5%.

But the Disallow count alone obscures a more interesting pattern.

AI search is more favoured than AI crawling

Disallow-to-Allow ratio by AI agent: search-linked bots like OAI-SearchBot are far more tolerated than training crawlers

The Disallow:Allow ratio separates the agents into two groups with fundamentally different commercial positions.

OAI-SearchBot - the crawler behind ChatGPT's web search feature - is blocked at a 2:1 ratio: the most balanced in the dataset. ChatGPT-User and PerplexityBot are blocked at roughly 3:1.

These agents power AI search products that can send referral traffic to publishers - a dynamic sites recognise and some have decided is worth accommodating.

Training crawlers carry no such proposition. GPTBot is blocked at 5:1. CCBot at 12:1. Bytespider at 17:1. Sites that name these agents in robots.txt have almost uniformly one thing to say.

The training/search distinction

The divergence between training crawlers and search-linked agents reflects a commercial logic that publishers have articulated explicitly in the legal cases and licensing discussions of the past two years: search engines send traffic; AI seeks training content.

A site that blocks GPTBot but allows or tolerates OAI-SearchBot is drawing a precise distinction - between a crawler that builds a product competing with the publisher and one that distributes the publisher's content to users. That distinction shows up cleanly in the ratio data.

Google-Extended - Google's opt-out string for Gemini training and Vertex AI - sits at 5:1, in line with other pure training crawlers. Applebot-Extended at 10:1. Neither carries a search referral proposition compelling enough to shift the ratio. The pattern holds: agents tied to AI search products face balanced responses; agents associated with model training face near-uniform exclusion.

Newspapers are most likely to block AI agents

The industry distribution of explicit AI blocking is dominated by content producers.

Top sectors naming an AI crawler to restrict, led by newspapers at 64%

Newspapers lead: 64% of newspaper sites name at least one AI crawler in a Disallow rule. Publishing sits at 49%, mass media at 43%.

These are the sectors most directly threatened by AI systems trained on their output and capable of generating competing content at scale without the editorial infrastructure.

Sports at 33% and entertainment at 30% reflect similar dynamics - content that has commercial value beyond the original site, that can be repurposed or summarised, and that publishers have decided they prefer not to feed into training pipelines without compensation.

At the other end of the spectrum, market research and outsourcing register zero explicit AI blocking. Transport, logistics, printing, and utilities cluster at 2–3%. These sectors have no text asset that requires protection from AI training - their value is in services, not publishable content, and the economics of AI access policy simply don't apply.

Sectors least likely to name an AI crawler, with market research and outsourcing at zero

Many sectors actively welcome some AI bots

The most striking finding in the allow data is that newspapers - which lead on explicit AI blocking at 64% - also sit joint third on explicit AI allowing, at 10%.

Top sectors explicitly allowing an AI crawler, led by real estate and travel at 13%

Real estate and travel and tourism lead the allow side at 13% each, airlines at 10%.

This is no contradiction but the search-linked distinction in practice: a newspaper that blocks GPTBot and CCBot while explicitly allowing OAI-SearchBot and PerplexityBot is executing a precise commercial decision about which AI agents return value. Newspapers are the industry most actively engaged with AI access policy in both directions simultaneously.

Real estate and travel leading on explicit allows reflects a different calculation: these are inventory-heavy sectors where AI-powered search could drive significant referral and booking traffic. A travel site that allows PerplexityBot is betting that AI search will surface its inventory to users who then click through. The economics of that wager look different from the economics facing a news publisher.

Catch-all wildcard rules block the most AI crawlers

The 12.7% who name AI crawlers explicitly represent only one part of the picture. The larger group - 27.2% of all sites - blocks AI crawlers without naming them, via wildcard User-agent: * with a blanket Disallow.

These sites have not made a specific policy decision about AI. They have made a general decision to close their robots.txt to all automated access, and AI crawlers are caught as part of that.

Whether this reflects intentional AI exclusion or simply inherited configuration - a Disallow: / that predates the AI crawler debate by a decade - cannot be determined from the file itself.

The combined effective block rate is 39.9%. Nearly four in ten of the world's most popular sites are inaccessible to a well-behaved AI crawler that respects robots.txt.

Try Zyte API

Build your first scraper in minutes

Free trial, no credit card. From a single request to production in an afternoon.

Get started
Data gathering for AILarge Language Models (LLMs)Access handlingFuture of the web

Robert Andrews

Senior editor

Robert is a journalist and editor turned content strategist who eats and sleeps the web. Previously senior editor at Google, News Corp, ContentNext and others. As Zyte's senior editor, Robert covers the state of the data-access industry — legal developments affecting AI and scra…

More from this author

In this article

  • How sites block AI with a text file
  • GPTBot leads the block list, but the ratios tell a different story
  • AI search is more favoured than AI crawling
  • The training/search distinction
  • Newspapers are most likely to block AI agents
  • Many sectors actively welcome some AI bots
  • Catch-all wildcard rules block the most AI crawlers

Follow

Get the latest

Zyte and the data web in your inbox — or wherever you already are.

Subscribe

Or follow elsewhere

The Community · Newsletter

The best of Zyte and the data web, in your inbox.

One curated edition — new articles, product updates, and the stories shaping the data web. No noise.

Services

Zyte Data

Fully managed web data extraction, delivered to your spec.

Explore Zyte Data

Web Scraping API

Zyte API

Scrape any website at scale with automatic proxy rotation and ban handling.

Sign Up

Developers

Zyte Developers

Docs, tools, and a community to help you build and scale scrapers.

Join Us
    • Zyte API
    • Ban Handling
    • AI Extraction
    • SERP
    • Enterprise
    • Scrapy Cloud
    • Agentic Web Data
    • Pricing
    • Product & E-commerce
    • Data for AI
    • Job Posting
    • Real Estate
    • News & Articles
    • Search
    • Social Media
    • Blog
    • Learn
    • Case Studies
    • Webinars
    • White Papers
    • Join our community
    • Join our Affiliate Program
    • Documentation
    • Meet Zyte
    • Contact us
    • Jobs
    • Support
    • Terms and Policies
    • Trust Center
    • Do not sell
    • Cookie settings
    • Web Data Compliance
    • Open Source
    • What is Web Scraping
    • Web Scraping in Python: Ultimate Guide
    • Stop getting blocked, start scraping
  • Logo EWDCILogo Most Loved WorkplaceLogo Job TogetherISO 27001 SealMedal Leader Europe Winter 2025Fastest Implementation Winter 2025Logo Leader Winter 2025Grid Leader Spring 2025Grid Leader Summer 2025Leader Fall 2025Leader Winter 2026
    XFacebookInstagramYouTubeLinkedInDiscord

    © Zyte Group Limited 2026