On this page
Overview
spaCy is a Python/Cython NLP library designed for production use. It combines linguistically motivated tokenization with components for part-of-speech tagging, dependency parsing, named entity recognition, text classification, lemmatization, entity linking, and more, with tokenization and training support across 70+ languages.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Compose tokenization, NER, parsing, and classification into pipelines
spaCy applies trainable and non-trainable components after tokenization and collects linguistic analysis in a shared Doc object.
Sources: [1]
Combine trained pipelines with custom components
Language-specific trained pipelines can be extended with project rules or custom models, including components backed by PyTorch or TensorFlow.
Sources: [1]
Cover training, packaging, and deployment workflows
The README lists a production-ready training system, model packaging, deployment, and workflow management among the core capabilities.
Sources: [1]
Best fit
Fits applications processing large text collections through reproducible NLP pipelines
It is useful when tokenization, NER, classification, and related tasks should be trained, packaged, versioned, and reused as a consistent pipeline rather than independent scripts.
Sources: [1]
Before adoption
Manage spaCy and trained-pipeline compatibility separately
Library updates can require updated statistical models. Production deployments should pin spaCy together with pipeline packages and custom components.
Sources: [1]
3.8.16 declares Python >=3.9,<3.15
The release-v3.8.16 package metadata declares python_requires = >=3.9,<3.15 and classifiers through Python 3.14.
Sources: [2]
Official sources
- [1]spaCy 3.8.16 README(2026-10-04)
- [2]spaCy 3.8.16 package metadata(2026-10-04)
- [3]spaCy 3.8.16 release(2026-10-04)
- [4]spaCy MIT license(2026-10-04)
Supplemental curator note
spaCy itself and trained pipelines are versioned separately. After upgrading the library, validate model compatibility and pin the library, pipeline packages, and custom components together for production. Version 3.8.16 declares Python >=3.9,<3.15.
Try it in 3 steps
- 1
Install spaCy 3.8.16 in a virtual environment
Pin stable 3.8.16 without downloading a trained pipeline.
python3 -m venv spacy-demo && . spacy-demo/bin/activate && pip install 'spacy==3.8.16' - 2
Create a tokenization sample with a blank English pipeline
Exercise the tokenizer and Doc API without an external model.
printf 'import spacy\nnlp = spacy.blank("en")\ndoc = nlp("spaCy processes text.")\nprint([t.text for t in doc])\n' > spacy-demo/demo.py - 3
Run the local tokenizer
Confirm tokenization without network access or a GPU. Validate version compatibility separately when adding trained pipelines.
. spacy-demo/bin/activate && python spacy-demo/demo.py
Growth
Growth trends · Last 30 days
33,935 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 1
- Open PRs
- 70
Development activity is still being collected.
Built with
Categories and tags
Categories
GitHub data
GitHub dataView detailed GitHub data
GitHub Topics
- natural-language-processing
- data-science
- machine-learning
- python
- cython
- nlp
- artificial-intelligence
- ai
- spacy
- nlp-library
- neural-network
- neural-networks
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- MLflow28,248 Stars
2 shared tag(s) · Same language
Unify ML experiments with LLM and agent tracing, evaluation, prompts, and AI Gateway in one tracking and registry platform
Python - Apache Spark44,112 Stars
2 shared tag(s)
Process SQL, DataFrames, streams, and machine learning on one distributed engine
Scala - XGBoost28,822 Stars
2 shared tag(s)
Train gradient-boosted trees for tabular data across CPU, GPU, and distributed environments
C++ - LightGBM18,829 Stars
2 shared tag(s)
fast and memory-efficient gradient-boosted trees at scale
C++ - PyTorch103,630 Stars
3 shared tag(s) · 2 shared category(s) · Same language
tensor computation and automatic differentiation from Python
Python - Stable Diffusion WebUI164,915 Stars
3 shared tag(s) · 1 shared category(s) · Same language
run image generation, inpainting, LoRAs, and extensions from one Gradio browser UI
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.