OSS TanbouSign in with GitHub

Crawl websites and send extracted structured data through reusable pipelines

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
64,578
Primary language
Python
License
BSD-3-Clause
Repository last updated
Oct 4, 2026
On this page

Overview

Scrapy is a Python web crawling framework that separates link traversal, page extraction, and item validation or storage. It suits recurring collection and research that must move from a one-off script to a maintainable process.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Keep traversal and extraction rules in a Spider

A Spider receives responses from starting URLs, selects values and follow-up links with CSS selectors or XPath, and yields structured items. An asynchronous downloader handles concurrent requests, while Item Pipelines validate, deduplicate, and store results. Downloader Middleware and Spider Middleware add cross-cutting request, response, or item processing.

Sources: [1][2]

Best fit

Manage recurring collection as a reproducible project

It fits product catalogs, public-information monitoring, and research datasets that repeatedly traverse many pages under the same rules. Start with a small URL set, verify the required fields and pagination, and validate the expected item shape in a Pipeline before expanding coverage.

Sources: [2][3]

Before adoption

Control request rates and plan for target-site changes

Version 2.19.0 requires Python 3.10 or newer. Review the target site's terms, robots.txt, authentication, and personal-data handling, then tune CONCURRENT_REQUESTS, CONCURRENT_REQUESTS_PER_DOMAIN, and DOWNLOAD_DELAY to an appropriate load. Markup changes can break selectors, so representative-page checks and failure monitoring are operational requirements. GitHub reports BSD-3-Clause.

Sources: [4][5][6][7]

Official sources

  1. [1]Scrapy 2.19.0 overview(2026-10-04)
  2. [2]Scrapy 2.19.0 tutorial(2026-10-04)
  3. [3]Scrapy 2.19.0 architecture(2026-10-04)
  4. [4]Scrapy 2.19.0 settings(2026-10-04)
  5. [5]Scrapy 2.19.0 package metadata(2026-10-04)
  6. [6]scrapy/scrapy GitHub repository metadata(2026-10-04)
  7. [7]Scrapy 2.19.0 LICENSE(2026-10-04)
Supplemental curator note

Before scaling a crawl, review the site's terms and robots.txt, then estimate selector maintenance, concurrency, and download delays against a small representative sample.

Try it in 3 steps

  1. 1

    Create an isolated Python environment

    Using Python 3.10 or newer in a macOS or Linux POSIX shell, create and activate a disposable virtual environment. Continue in the same terminal and working directory.

    python3 -m venv scrapy-demo-env && . scrapy-demo-env/bin/activate
  2. 2

    Install Scrapy 2.19.0

    Install the reviewed release in the virtual environment so its project commands are available.

    python -m pip install "Scrapy==2.19.0"
  3. 3

    Create and verify a Spider

    Generate a project and an example.com Spider without requesting the site, then verify that Scrapy registers it as example.

    scrapy startproject tutorial && cd tutorial && scrapy genspider example example.com && test "$(scrapy list)" = "example"
Check the official README

Growth

Growth trends · Last 30 days

64,578 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
148
Open PRs
143

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • python
  • scraping
  • crawling
  • framework
  • crawler
  • hacktoberfest
  • web-scraping
  • web-scraping-python
Stars
64,578
Forks
11,989
Watchers
1,758
Open issues
151
Contributors
374
Owner type
Organization
Primary language
Python
Repository last updated
Oct 4, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?