Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

How to ensure data quality in your Scrapy web scraping projects using Spidermon and Claude Code
Spidermon is an open-source monitoring framework for Scrapy. You attach it to your spider, define what "success" looks like, and it automatically checks your crawl results after the spider closes, flagging anything that doesn't meet your standards.

Why your API responses look like gibberish: the gzip decompression trap
The script was working. Requests were going out, responses were coming back with HTTP 200. But the response body was unreadable noise, a wall of binary characters that crashed the JSON parser and reported "no data found". No error code, no timeout, no network failure; just garbage where structured data should be.

Dawn of the autonomous data pipeline
Discover how autonomous, agent-driven data pipelines are transforming web scraping in 2026, enabling self-healing systems, API discovery, and end-to-end automation.

Are programming practices relevant anymore?
From LLM-powered extraction to agentic pipelines, here's how AI is reshaping every stage of the web scraping workflow in 2026 -- and what it means for your stack.

AI is the new engine for web scraping
From LLM-powered extraction to agentic pipelines, here's how AI is reshaping every stage of the web scraping workflow in 2026 -- and what it means for your stack.

Super-powers, toll booths and the new era of data collection
Explore how AI is transforming web scraping, the rise of advanced anti-bot systems, and what the future holds for data collection in an increasingly controlled internet.

Is your AI coding assistant stuck in the past?
AI coding tools like Copilot can suggest outdated code and practices. Learn why this happens and how developers can ensure they use the latest standards.

Data outcomes are top of the scraping stack
By 2026, scraping shifts from proxy management to API-first data outcomes. Learn why unified scraping APIs are replacing traditional stacks.

No-code web scraping workflows are here: Introducing the Zyte integration for Zapier
Discover Zyte’s Zapier integration to automate web data extraction and connect it to 10,000+ apps. Build powerful workflows without coding and turn web data into actionable insights instantly.

The Scrapy whisperer: Adrian Chaves on Web Scraping Copilot
An interview with Scrapy maintainer Adrian Chaves on Zyte’s Web Scraping Copilot, AI-generated parsing code, and building reliable scraping workflows.

How to parse HTML tables into structured data (CSV/Excel)
In this guide, you'll learn three things: how HTML tables are actually structured (so the parsing makes sense), how to extract clean tabular data using Python, and how to export it to CSV or Excel

Code is cheap, show me the talk: How copilots are re-engineering developers
Discover how AI copilots like Zyte’s Web Scraping Copilot are transforming developer workflows—making code a commodity and shifting value to problem-solving and prompting skills.




_HFpro5d6k3.png&w=256&q=75)
_E4PyVpfAxa.png&w=256&q=75)


-(1).png&w=1920&q=75)
-(1)_VZGHqxCgXV.png&w=1920&q=75)