Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Skinfer: Inferring JSON Schemas Made Easy
Skinfer: A Tool for Inferring JSON Schemas - Discover Skinfer, a powerful tool for inferring JSON schemas. Simplify data extraction from unstructured sources.

Handling JavaScript In Scrapy With Splash
Handling modern websites that entirely run on Javascript? In this article, learn how to use Splash to render JavaScript-based pages in your Scrapy spiders.

Zyte Crawls The Deep Web
Explore Memex - The Future of Web Archiving: Learn about Memex, an innovative web archiving system that opens new horizons for preserving digital content.

New Changes to Our Scrapy Cloud Platform: Enhanced Performance and Features
New Changes to Our Scrapy Cloud Platform - Stay up-to-date with the latest changes to Scrapy Cloud. Enhance your web scraping workflow with new features.

Introducing ScrapyRT: An API for Scrapy Spiders
Introducing ScrapyRT: An API for Scrapy Spiders - Make the most of your Scrapy spiders with ScrapyRT. Explore its functionalities as an API for seamless integration.

Looking Back at 2014: Highlights and Milestones
Looking Back at 2014 - Take a trip down memory lane and see the milestones and breakthroughs at Zyte in 2014.

XPath Tips From The Web Scraping Trenches
XPath is helpful for web scraping, allowing to write specifications more flexibly than CSS selectors. This tutorial is packed with XPath tips and examples.

Introducing Data Reviews: Unlocking Insights with Zyte
Introducing Data Reviews - Empower your decision-making with data reviews. Gain insights into data quality and credibility for your web scraping projects.

Extract Schema.Org Microdata with Scrapy Selectors
Web pages are full of data. Microdata markup helps machines understand pages. Schema.org supports a set of schemas for structured data markup on web pages.

Portia: The Open-Source Visual Web Scraper
As you can see, Portia allows you to visually configure what’s crawled and extracted in a very natural way. It provides immediate feedback, making the process

Optimizing Memory Usage Of Scikit-Learn Models Using Succinct Tries
We use the scikit-learn library for various machine-learning tasks at Zyte. For example, for text classification we'd typically build a statistical

Open source at Zyte
Open Source at Zyte, Now Zyte - Embrace the open-source movement at Zyte. Learn how we contribute to the community and promote transparency.




_HFpro5d6k3.png&w=256&q=75)
_E4PyVpfAxa.png&w=256&q=75)


-(1).png&w=1920&q=75)
-(1)_VZGHqxCgXV.png&w=1920&q=75)