PINGDOM_CHECK

Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Introducing Web Scraping Copilot 1.0: AI-accelerated web scraping inside VS Code
Announcement

Introducing Web Scraping Copilot 1.0: AI-accelerated web scraping inside VS Code

Discover Web Scraping Copilot 1.0, Zyte’s VS Code extension that uses AI to generate, test, and deploy production-ready Scrapy spiders faster while maintaining full developer control.

Mitch Holt10 min read
Zyte Blog — field notes from the world of data extraction
Use case

How to Test Web Scrapers During Development

Learn how to test web scrapers during development. Validate selectors, use HTML fixtures, and ensure reliable data extraction across changing websites.

Arnold Alexander10 min read
Zyte Blog — field notes from the world of data extraction
Use case

How Developers Debug Web Scraping Selectors

Learn how developers debug web scraping selectors. Discover common issues, testing techniques, and how to build reliable extraction logic for changing websites.

Arnold Alexander10 min read
Zyte Blog — field notes from the world of data extraction
Use case

Best VS Code Extensions for Web Scraping

Discover the best VS Code extensions for web scraping, including Python tools, HTTP clients, and AI-powered solutions to build and debug scrapers faster.

Arnold Alexander10 min read
Zyte Blog — field notes from the world of data extraction
Use case

How to Build a Web Scraper in VS Code (Step-by-Step)

Learn how to build a web scraper in VS Code using Scrapy and AI tools. Follow this step-by-step guide to create, test, and scale your scraping projects.

Arnold Alexander10 min read
Build your own MCP server: LLMs meets web data with Zyte API
Web scraping APIs

Build your own MCP server: LLMs meets web data with Zyte API

Learn how to build your own Model Context Protocol (MCP) server to connect LLMs with real-time web data using Zyte API, FastMCP, and the Docker MCP toolkit.

Ayan Pahwa10 min read
More data, more trouble: How a perfect corpus corrupted my AI dream
Data gathering for AI

More data, more trouble: How a perfect corpus corrupted my AI dream

A failed AI experiment reveals why adding more data doesn’t always improve LLM outputs. Learn when web scraping, RAG, and curated datasets actually make AI better.

Neha Setia Nagpal10 min read
Zyte Blog — field notes from the world of data extraction
Developer interest

Stop using Python requests for web scraping: Use these modern modules instead

While the 'Requests' library remains the default choice for many Python developers due to its reliability and extensive documentation, the Python HTTP landscape has evolved considerably. Modern alternatives now offer significant advantages, including built-in asynchronous support, HTTP/2 compatibility, enhanced performance, and up-to-date TLS handling.

Ayan Pahwa6 min read
Claude skills, MCP or Web Scraping Copilot: Which should you choose?
AI-assisted data extraction

Claude skills, MCP or Web Scraping Copilot: Which should you choose?

Compare Claude skills, MCP servers, and Web Scraping Copilot to understand when to use each for AI-powered web scraping, data extraction, and production pipelines with Zyte API.

Neha Setia Nagpal10 min read
Supercharging web scraping with Claude skills
AI-assisted data extraction

Supercharging web scraping with Claude skills

Learn how Claude skills can automate HTML fetching, AI parsing, selector generation, and structured data extraction to build faster, smarter web scraping workflows.

John Rooney10 min read
P1001083
Web data collection legality

Is your AI breaking the law? Legal experts’ advice for web scrapers

Legal experts discuss how AI, web scraping, copyright law, and the EU AI Act intersect—covering fair use, data provenance, and compliance risks for businesses.

Robert Andrews10 min read
Screenshot webpages with this Claude Skill and Zyte API

Screenshot webpages with this Claude Skill and Zyte API

John Rooney7 min read