PINGDOM_CHECK

Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Four sweet spots for AI in web scraping
Large Language Models (LLMs)

Four sweet spots for AI in web scraping

Discover how AI and LLMs are enhancing web scraping with smarter crawling, fuzzy data extraction, automated spider generation, and intelligent QA.

Theresia Tanzil7 min read
Kill your product - why sacrificing your cash cow can be the path to growth
Web data collection

Kill your product - why sacrificing your cash cow can be the path to growth

Letting go of the past is the best way to embrace the future. Retiring a flagship product isn’t a sign of failure; it’s a commitment to innovation.

6 min read
From script to system: 10 building blocks to scale web scraping
Web data collection

From script to system: 10 building blocks to scale web scraping

Scaling your business’ web data gathering – acquiring, monitoring and storing a growing amount of data from a growing number of sources over time – requires holistic planning.

Theresia Tanzil10 min read
Zyte Blog — field notes from the world of data extraction
Scraping practice

Scrape Web Pages and Files Using Python, wget, and Zyte

The command-line utility wget (pronounced "web-get") can download online files. This free network downloader may run in the background without user intervention.

Karlo Jeđud7 min read
New in Zyte: Scroll Control, Lower Costs, and More
Web data collection

New in Zyte: Scroll Control, Lower Costs, and More

As the web continues to evolve, Zyte API is evolving right alongside it—adding powerful new features and refinements designed to make data extraction smarter, faster, and more adaptable than ever.

Daniel Cave5 min read
The future of Scrapy: Smarter, faster and ready for AI-powered scraping
Open-source

The future of Scrapy: Smarter, faster and ready for AI-powered scraping

What does the future hold for the tool some describe as “the gift that revolutionised web scraping”?

Robert Andrews6 min read
Rise of the Data Vendor: How Outsourcing is Transforming Supply and Fuelling Businesses
Web data collection

Rise of the Data Vendor: How Outsourcing is Transforming Supply and Fuelling Businesses

With the emergence of managed data extraction vendors, businesses no longer need to gather web data themselves.

Robert Andrews6 min read
Quality, focus and scale: Three ways data outsourcing benefits businesses
Web data collection

Quality, focus and scale: Three ways data outsourcing benefits businesses

The Strategic Case for Buying Web Data: Quality, Focus, and Scale

Theresia Tanzil8 min read
What AI Builders Need to Know About the Training Data Copyright Debate
Large Language Models (LLMs)

What AI Builders Need to Know About the Training Data Copyright Debate

The generative AI gold rush is upon us, with astounding new products and capabilities that are fuelled by web data and raising questions about AI copyright.

Sanaea Daruwalla6 min read
Ten years since Scrapy 1.0: The stats and stories behind your favorite framework
Open-source

Ten years since Scrapy 1.0: The stats and stories behind your favorite framework

See what 10 years of Scrapy 1.0 has built — in milestones and metrics.

Cleber Alexandre5 min read
Zyte Blog — field notes from the world of data extraction
Traffic

Using curl with a Proxy for Web Scraping

When it comes to command-line tools for HTTP requests, few are as versatile and powerful as curl. Loved by developers and system administrators alike, curl makes fetching web resources straightforward.

Karlo Jeđud8 min read
What’s your data type? Solving the procurement problem
Web data collection

What’s your data type? Solving the procurement problem

Engagements with data suppliers break down when buyers don’t have a clear project concept. Understanding and articulating your needs is paramount. Meet the three types of data buyers. Which one are you?

Theresia Tanzil10 min read