OSS Tanbou

Paperless-ngx — OCR scans and PDFs into a searchable self-hosted document archive

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
45,262
Primary language
Python
License
GPL-3.0
Repository last updated
Sep 18, 2026

Overview

Paperless-ngx is a document management system that ingests scanned or digital documents, runs OCR when needed, indexes their contents for full-text search, and organizes them with tags, correspondents, document types, storage paths, matching, and workflows. The latest release checked is v3.1.3.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

OCR scanned documents and make their contents searchable

Documents without embedded text can be OCRed so invoices, contracts, notices, and other records become searchable instead of remaining opaque files.

Sources: [2][3]

Automate organization with metadata matching and workflows

Tags, correspondents, document types, and storage paths can be assigned through matching, while workflows can react to document changes or schedules. Mail ingestion is also supported.

Sources: [3][4]

For homes and teams that need invoices, contracts, and records to be searchable rather than filed manually

It fits mixed paper/PDF archives where users need metadata and content search rather than only folders.

Sources: [2][3]

Sensitive documents are stored without encryption, so trusted hosting and backups are essential

The README explicitly warns that sensitive documents are stored in clear text and Paperless-ngx should never be run on an untrusted host. Operators should add host/storage encryption, access controls, updates, and backups as separate controls. The project is GPL-3.0.

Sources: [2][5][6]

Official sources

  1. [1]paperless-ngx/paperless-ngx repository(2026-09-18)
  2. [2]Paperless-ngx README(2026-09-18)
  3. [3]Paperless-ngx usage documentation(2026-09-18)
  4. [4]Paperless-ngx configuration documentation(2026-09-18)
  5. [5]Paperless-ngx LICENSE(2026-09-18)
  6. [6]Paperless-ngx v3.1.3 release(2026-09-18)
Supplemental curator note

Paperless-ngx turns scans and PDFs into a searchable archive, but its README explicitly warns that sensitive documents are stored without encryption and should not be hosted on an untrusted machine.

Try it in 3 steps

  1. 1

    Get the source

    git clone --depth 1 https://github.com/paperless-ngx/paperless-ngx.git
  2. 2

    Enter the repository

    cd paperless-ngx
  3. 3

    Check the official steps

    Continue with the commands in the README Installation, Quick Start, or Getting Started section.

    find . -maxdepth 1 -iname 'README*' -exec sed -n '1,220p' {} \; -quit
Check the official README

Growth

Growth trends · Last 30 days

45,262 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
279
Open PRs
6

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • angular
  • archiving
  • django
  • dms
  • document-management
  • document-management-system
  • machine-learning
  • ocr
  • optical-character-recognition
  • pdf
  • ai
  • llm
Stars
45,262
Forks
3,121
Watchers
165
Open issues
0
Primary language
Python
License
GPL-3.0
Repository last updated
Sep 18, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.