PINGDOM_CHECK

#ExtractSummit2026 The world's largest web scraping conference returns. Austin Oct 7–8 · Dublin Nov 10–11.

Register now
Data Services
Pricing
Login
Try Zyte APIContact Sales
  • Unblocking and Extraction

    Zyte API

    The ultimate API for web scraping. Avoid website bans and access a headless browser or AI Parsing

    Ban Handling

    Headless Browser

    AI Extraction

    SERP

    Enterprise

    DocumentationSupport

    Hosting and Deployment

    Scrapy Cloud

    Run, monitor, and control your Scrapy spiders however you want to.

    Coding Agent Add-Ons

    Agentic Web Data

    Plugins that give coding agents the context to build production Scrapy projects. Starts with Claude Code.

  • Data Services
  • Pricing
  • Browse

    • BlogArticles, podcasts, videos
    • Case studiesCustomer outcomes
    • White papersIn-depth reports
    • DocumentationGuides & API reference
    • EventsConferences, webinars, recordings

    Subscribe

    • NewsletterSwiftly delivered
    • Discord communityExtract Data community
  • Product and E-commerce

    From e-commerce and online marketplaces

    Data for AI

    Collect and structure web data to feed AI

    Job Posting

    From job boards and recruitment websites

    Real Estate

    From Listings portals and specialist websites

    News and Article

    From online publishers and news websites

    Search

    Search engine results page data (SERP)

    Social Media

    From social media platforms online

  • Meet Zyte

    Our story, people and values

    Contact us

    Get in touch

    Support

    Knowledge base and raise support tickets

    Terms and Policies

    Accept our terms and policies

    Open Source

    Our open source projects and contributions

    Web Data Compliance

    Guidelines and resources for compliant web data collection

    Join the team building the future of web data
    We're Hiring
    Trust Center
    Security, compliance & certifications
Login
Try Zyte APIContact Sales
All articles
AI66, 66 articles
Data quality13, 13 articles
Developer interest57, 57 articles
Integration2, 2 articles
Open-source41, 41 articles
Proxies29, 29 articles
Scraping practice19, 19 articles
Scraping strategy29, 29 articles
Web data60, 60 articles
Web scraping APIs36, 36 articles
Scrapy47, 47 articles
Scrapy Cloud14, 14 articles
Web Scraping Copilot11, 11 articles
Zyte API57, 57 articles
AI & Machine Learning3, 3 articles
Automotive2, 2 articles
E-commerce & retail27, 27 articles
Entertainment & Streaming2, 2 articles
Financial Services8, 8 articles
Government2, 2 articles
Market Research & Intelligence3, 3 articles
Media & publishing8, 8 articles
Real Estate2, 2 articles
Recruitment & HR3, 3 articles
Transportation & Logistics2, 2 articles
Travel & hospitality2, 2 articles
iPaaS2, 2 articles
Large language model24, 24 articles
MCP3, 3 articles
Python88, 88 articles
Web Scraping Industry Report14, 14 articles

Appearance

Discord Community
BlogDeveloper interestWhat I Learned As A Google Summer Of Code Student At Zyte
ArticleDeveloper interest

What I Learned As A Google Summer Of Code Student At Zyte

Google Summer of Code (GSoC) was such a great experience for students like me. I learned so much about open source communities as well as contributing to

C

Chau Tung Lam Nguyen Bhatt

4 min read · September 12, 2018

What I Learned As A Google Summer Of Code Student At Zyte

What I learned as a Google Summer of Code student at Zyte

Google Summer of Code (GSoC) was such a great experience for students like me. I learned so much about open source communities as well as contributing to their complex projects. I also learned a great deal from my mentors, Konstantin and Cathal, about programming and software engineering practices. In my opinion, the most valuable lesson I got from GSoC was what it was like to be a Software Engineer, which prepared me to continue the pursuit of my dream career in technology.

What is GSoC?

GSoC is a program hosted by Google for students to spend their summer contributing to open source projects from various organizations. Fortunately, my proposal for a project called Scrapy in Python Software Foundation got accepted. Before I applied for the program, I had no idea what open source was, so I decided to take some small steps to get started on contributing. I spent 3 months studying smaller projects within the Scrapy organization to familiarize myself with the code. I found the application easier after I got used to contributing. I also received a lot of help from people within the community, especially from Konstantin, who was to be my GSoC mentor. From my experience, it was extremely important to have a willingness to learn and ask for help when anyone wanted to start participating in an open source project.

My Expectations

In the beginning, I did not know what to expect at GSoC due to the mixed review from students. I figured since it is an individual project for each student, their experiences must be different from one another. I also wondered how a student-built project could have significant impact on the community. However, I came out of GSoC with a completely different perspective. I learned that it was me who would make or break my project, and I was responsible to push myself harder. I also believed my project was a success because I got other contributors’ attention and, hopefully, I inspired more students to do the same thing!

My GSoC Journey

The first week of GSoC was the most confusing part of the program, because my original project proposal was too generic. It took more effort than I anticipated to build the product I planned on originally. However, it was okay, because it was a common mistake among students, since most of us never worked on a large scale project before.

I was first assigned to profile the Scrapy project spider to identify the components that had significant run time to implement speed improvement. It is a meticulous process, so I had to spend so much time on it. I was bored because I thought I did not join GSoC to analyze graphs. However, I learned that everything had to start small. I could spend less time on profiling and started coding instead, but I would be building something meaningless, since I wasn’t sure which component to be optimized. I believe it is an analogy to building a skyscraper with an insecure foundation.

After a lot of time profiling, I finally found out what component needed optimization, the URL parsing library from CPython urllib. I then had to profile again. I was getting impatient, since I couldn’t get my hands on coding yet. The potential libraries fell into 2 types: those that didn’t improve anything at all and those that were fast but not compatible with Scrapy. I felt hopeless from being stuck on the same problem for such a long time. However, we, student developers, will have to experience that eventually. I took a deep breath and continued with the task that I was assigned. Eventually, I decided to build a library from scratch to replace urllib, and I named the project Scurl (GitHub repository).

After a month developing Scurl, I got it to a stable stage. However, the Chromium source I used was from another project on Github, so Scurl would not last long if the Chromium source could not be updated. What if Chromium would release a patch to the components that the library uses? Or the source code would change completely in 20 years? My project would be thrown away.

Chromium is a gigantic project. Building the Chromium Source on an average machine is slower than a snail running half a mile (quoted from here). Working with just 2 components of the project was really difficult. Since I did not have any prior experience working on Chromium’s source, or C++, I spent a lot of time trying to track which source files I needed for my project. It was taking me too long to figure it out, and I thought of giving up several times. However, giving up was not an option, since I already spent two and a half months for the project. Despite the struggle, I learned how to update the Chromium source code, which allowed others to maintain this library with ease.

Lessons Learned

Overall, GSoC gave me a chance to be a better software developer. Not only did I have an opportunity to hone my programming skills, I also trained my mind to be ready for the career path that I chose. I am truly grateful for the experience that I had. I want to thank my mentors, Konstantin and Cathal for their help and support, and I hope that I have inspired others to go out of their comfort zones to do the same!

Special thanks to Tram Nguyen and Samuel Coveney for helping me edit this article!

Try Zyte API

Build your first scraper in minutes

Free trial, no credit card. From a single request to production in an afternoon.

Get started
Developer interest
C

Chau Tung Lam Nguyen Bhatt

More from this author

In this article

  • What is GSoC?
  • My Expectations
  • My GSoC Journey
  • Lessons Learned

Follow

Get the latest

Zyte and the data web in your inbox — or wherever you already are.

Subscribe

Or follow elsewhere

The Community · Newsletter

The best of Zyte and the data web, in your inbox.

One curated edition — new articles, product updates, and the stories shaping the data web. No noise.

Services

Zyte Data

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Talk to us

Web Scraping API

Zyte API

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Sign Up

Developers

Zyte Developers

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Join Us
    • Zyte API
    • Ban Handling
    • AI Extraction
    • SERP
    • Enterprise
    • Scrapy Cloud
    • Agentic Web Data
    • Pricing
    • Product & E-commerce
    • Data for AI
    • Job Posting
    • Real Estate
    • News & Articles
    • Search
    • Social Media
    • Blog
    • Learn
    • Case Studies
    • Webinars
    • White Papers
    • Join our community
    • Documentation
    • Meet Zyte
    • Contact us
    • Jobs
    • Support
    • Terms and Policies
    • Trust Center
    • Do not sell
    • Cookie settings
    • Web Data Compliance
    • Open Source
    • What is Web Scraping
    • Web Scraping in Python: Ultimate Guide
    • Stop getting blocked, start scraping
  • EWDCI logoMost loved workplace certificateZyte rewardISO 27001 iconG2 rewardG2 rewardG2 reward
    XFacebookInstagramYouTubeLinkedInDiscord

    © Zyte Group Limited 2026