OSS Tanbou

run OCR for images in 100+ languages from browsers and Node.js

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
38,750
Primary language
JavaScript
License
Apache-2.0
Repository last updated
May 17, 2026
On this page

Overview

Tesseract.js exposes a WebAssembly port of the Tesseract OCR engine to JavaScript. It runs in browsers and Node.js, creating workers that recognize images and return extracted text.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Create an OCR worker and recognize images from JavaScript

Use createWorker with a language, pass images to worker.recognize, and read the returned text. The README recommends reusing one worker for multiple images and terminating it once at the end.

Sources: [2]

Use the same OCR library in browsers and Node.js

The project supports webpack, ES modules, CDN script tags, and npm on Node.js, covering both client-side and server-side JavaScript workflows.

Sources: [2]

Best fit

Fits image-to-text features and JavaScript document-processing pipelines

It is useful for extracting text from scanned images, photos, or screenshots before passing the result into search, data entry, or downstream processing.

Sources: [2]

Before adoption

PDF handling and OCR model improvements are out of scope

The README explicitly states that Tesseract.js does not support PDF files and does not modify the Tesseract recognition model to improve accuracy. Those requirements need a different tool or workflow.

Sources: [2]

Tesseract.js v7 requires Node.js 16 or newer

The README states that v7 requires Node.js 16+. Since v6, output formats other than text are disabled by default, so formats such as hOCR must be explicitly enabled when needed.

Sources: [2]

Official sources

  1. [1]naptha/tesseract.js repository(2026-09-30)
  2. [2]Tesseract.js README(2026-09-30)
  3. [3]Tesseract.js Apache-2.0 license(2026-09-30)
Supplemental curator note

It makes OCR accessible from JavaScript, but PDF processing is outside its scope. Reuse workers for batches and explicitly enable any non-text outputs the application needs.

Try it in 3 steps

  1. 1

    Install Tesseract.js

    Add the official package to a browser or Node.js JavaScript project. Tesseract.js v7 requires Node.js 16 or newer on Node.js.

    npm install tesseract.js
  2. 2

    Create an English OCR worker

    After importing createWorker from tesseract.js, create one worker. Replace the language code when another supported language is required.

    const worker = await createWorker('eng');
  3. 3

    Recognize the image and terminate the worker

    Inspect the extracted text. Reuse a worker for multiple images and terminate it once at the end.

    const { data: { text } } = await worker.recognize(image); await worker.terminate();
Check the official README

Growth

Growth trends · Last 30 days

38,750 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
0
Open PRs
20

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • tesseract
  • webassembly
  • ocr
  • javascript
  • deep-learning
Stars
38,750
Forks
2,394
Watchers
486
Open issues
35
Contributors
79
Owner type
Organization
Primary language
JavaScript
License
Apache-2.0
Repository last updated
May 17, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?