PINGDOM_CHECK

Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

fable cover

Fable 5.1 shipped, and GLM-5.3-Flash turned out to be someone I already met

A personal take on Claude Fable 5.1 and GLM-5.3-Flash: real benchmarks, a live extraction test, and the model I guessed before Zhipu confirmed it.

Ayan Pahwa
zytexmarimo-cover

Web data in a reactive notebook: an introduction to marimo

marimo is a reactive Python notebook that reruns only affected cells. Learn to scrape web data with Zyte API, chart prices, and cache costly API calls.

Ayan Pahwa
Toy robot walking toward a wall
Future of the web

75% of the web uses robots.txt - here's how

Three in four of the world's top sites publish a robots.txt, yet few name individual crawlers. A look inside the web's advisory layer: coverage, sophistication and limits.

Robert Andrews
Browser automation scripts - Puppeteer and Playwright
Product Update

Introducing Zyte CDP support: Your browser automation, our infrastructure

Want to use traditional browser automation frameworks without infrastructure headaches? Tap our new CDP support to run scripts on Zyte’s powerful infrastructure.

Valter Sciarrillo
Text file struggling under the weight of European privacy burden.
Web data collection legality

Why Europe’s new AI scraping guidelines miss the mark

New guidelines on generative AI scraping aim to protect user privacy. But turning a rudimentary, 30-year-old web server standard into a legal barrier will disenfranchise users and create accidental monopolies.

Sally-Anne Hinfey
webfetch workflow

WebFetch is ‘lossy’ by design: give your coding agent a better fetch in one command

Your coding agent built-in webfetch tool is not the best and it's hampering your research and coding workflows, fix it with one CLI tool and never face blocks again.

Ayan Pahwa
SoWA retail industry analysis
Access handling

Revealed: How retailers use tech to block the bots

The web’s big general marketplaces exist to be browsed by everyone. But they are also the second most defended sector online, shutting the door on AI crawlers.

Robert Andrews
Theresia Tanzil, State of Web Access portrait
Scraping strategy

The web is being priced, not blocked

Zyte's State of Web Access research finds new barriers making web scraping more difficult. We interview the researcher who says that doesn't mean the opportunity is over.

Robert Andrews
The State of Web Access report cover

The State of Web Access

View the free State of Web Access report to learn just how difficult modern web scraping has become.

SoWA fsahion icon
Access handling

Fashion websites are the hardest to size up

While publishers fight over AI crawlers, fashion has built some of the most heavily defended shop windows on the web. So, which controls are in fashion, in fashion?

Robert Andrews
Dom and Neha talking about synthetic web and extract summit

The web is synthetic. Now look at the human trail

Domagoj Marić explores how AI, web scraping, and OSINT turn scattered personal data into profiles, scams, and security risks at Extract Summit.

Neha Setia Nagpal
State of Web Access - industries
Access handling

The State of Web Access: How different industries behave to bots

Zyte's large-scale audit of web access controls shows how different kinds of businesses exhibit different policies.

Robert Andrews