OSS Tanbou

build production NLP pipelines in Python and Cython for tokenization, NER, tagging, parsing, text classification, and training

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
33,935
Primary language
Python
License
MIT
Repository last updated
Sep 30, 2026
On this page

Overview

spaCy is a Python/Cython NLP library designed for production use. It combines linguistically motivated tokenization with components for part-of-speech tagging, dependency parsing, named entity recognition, text classification, lemmatization, entity linking, and more, with tokenization and training support across 70+ languages.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Compose tokenization, NER, parsing, and classification into pipelines

spaCy applies trainable and non-trainable components after tokenization and collects linguistic analysis in a shared Doc object.

Sources: [1]

Combine trained pipelines with custom components

Language-specific trained pipelines can be extended with project rules or custom models, including components backed by PyTorch or TensorFlow.

Sources: [1]

Cover training, packaging, and deployment workflows

The README lists a production-ready training system, model packaging, deployment, and workflow management among the core capabilities.

Sources: [1]

Best fit

Fits applications processing large text collections through reproducible NLP pipelines

It is useful when tokenization, NER, classification, and related tasks should be trained, packaged, versioned, and reused as a consistent pipeline rather than independent scripts.

Sources: [1]

Before adoption

Manage spaCy and trained-pipeline compatibility separately

Library updates can require updated statistical models. Production deployments should pin spaCy together with pipeline packages and custom components.

Sources: [1]

3.8.16 declares Python >=3.9,<3.15

The release-v3.8.16 package metadata declares python_requires = >=3.9,<3.15 and classifiers through Python 3.14.

Sources: [2]

Keep transformer and GPU dependencies explicit

Transformer integration and CUDA support use optional dependencies. Separate CPU-only NLP from GPU/transformer environments unless those dependencies are intentionally required.

Sources: [2][1]

Official sources

  1. [1]spaCy 3.8.16 README(2026-10-04)
  2. [2]spaCy 3.8.16 package metadata(2026-10-04)
  3. [3]spaCy 3.8.16 release(2026-10-04)
  4. [4]spaCy MIT license(2026-10-04)
Supplemental curator note

spaCy itself and trained pipelines are versioned separately. After upgrading the library, validate model compatibility and pin the library, pipeline packages, and custom components together for production. Version 3.8.16 declares Python >=3.9,<3.15.

Try it in 3 steps

  1. 1

    Install spaCy 3.8.16 in a virtual environment

    Pin stable 3.8.16 without downloading a trained pipeline.

    python3 -m venv spacy-demo && . spacy-demo/bin/activate && pip install 'spacy==3.8.16'
  2. 2

    Create a tokenization sample with a blank English pipeline

    Exercise the tokenizer and Doc API without an external model.

    printf 'import spacy\nnlp = spacy.blank("en")\ndoc = nlp("spaCy processes text.")\nprint([t.text for t in doc])\n' > spacy-demo/demo.py
  3. 3

    Run the local tokenizer

    Confirm tokenization without network access or a GPU. Validate version compatibility separately when adding trained pipelines.

    . spacy-demo/bin/activate && python spacy-demo/demo.py
Check the official README

Growth

Growth trends · Last 30 days

33,935 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
1
Open PRs
70

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • natural-language-processing
  • data-science
  • machine-learning
  • python
  • cython
  • nlp
  • artificial-intelligence
  • ai
  • spacy
  • nlp-library
  • neural-network
  • neural-networks
Stars
33,935
Forks
4,734
Watchers
564
Open issues
178
Contributors
388
Owner type
Organization
Primary language
Python
License
MIT
Repository last updated
Sep 30, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?