OSS Tanbou

scale parallel analytics with NumPy/pandas-like collections and lazy task graphs

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
13,931
Primary language
Python
License
BSD-3-Clause
Repository last updated
Sep 29, 2026
On this page

Overview

Dask is a flexible parallel computing library for analytics. High-level collections such as Dask Array, DataFrame, and Bag build task graphs, and schedulers execute those graphs lazily on a single machine or a cluster.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Partition Array and DataFrame workloads into lazy task graphs

Dask collections split data into chunks or partitions and represent operations as task graphs rather than immediately executing them. Calling compute() asks a scheduler to evaluate the required graph.

Sources: [2][1]

Install only core features or add array, dataframe, distributed, and diagnostics extras

The package offers a lightweight core plus optional dependency groups for array, dataframe, distributed, diagnostics, and complete installations so environments can include only the required functionality.

Sources: [3][4]

Best fit

Fits NumPy and pandas workflows that need larger data or parallel execution

Dask provides collections close to familiar PyData APIs while introducing partitioning and task scheduling, supporting a path from multicore single-machine work to distributed clusters.

Sources: [2][1]

Before adoption

Tune partition sizes and task granularity to control scheduler overhead

Because execution is organized around task graphs and partitions, workloads with very many tiny tasks can spend a larger fraction of time in scheduling overhead. Choose chunk and partition sizes appropriate to the data and operations.

Sources: [2]

Dask 2026.8.0 requires Python 3.10+ and aligns its distributed extra to the same release line

The 2026.8.0 package metadata requires Python >=3.10 and specifies distributed >=2026.8.0,<2026.8.1 for the distributed extra, so cluster environments should keep Dask and Distributed on matching release lines.

Sources: [4][5]

Official sources

  1. [1]Dask 2026.8.0 README(2026-10-03)
  2. [2]Dask 2026.8.0 ten-minute guide(2026-10-03)
  3. [3]Dask 2026.8.0 installation guide(2026-10-03)
  4. [4]Dask 2026.8.0 package metadata(2026-10-03)
  5. [5]Dask 2026.8.0 release(2026-10-03)
Supplemental curator note

Dask can extend familiar NumPy and pandas-style workflows to partitioned data and larger-than-memory or parallel computation. Scheduler overhead can dominate small workloads, so inspect task graphs and partition sizes before moving to distributed execution.

Try it in 3 steps

  1. 1

    Create an isolated Python virtual environment

    Use Python 3.10 or newer as required by Dask 2026.8.0.

    python3 -m venv .venv && . .venv/bin/activate
  2. 2

    Install Dask Array 2026.8.0

    Pin core Dask and the Array dependency set, including NumPy.

    python -m pip install "dask[array]==2026.8.0"
  3. 3

    Compute a chunked array lazily

    Split 100 elements into four chunks, build the task graph, and execute it with compute. Add the matching distributed release line when moving to a cluster.

    python -c 'import dask.array as da; x=da.arange(100, chunks=25); print(x.sum().compute())'
Check the official README

Growth

Growth trends · Last 30 days

13,931 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
1
Open PRs
307

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • dask
  • python
  • pydata
  • numpy
  • pandas
  • scikit-learn
  • scipy
Stars
13,931
Forks
1,971
Watchers
203
Open issues
1,043
Contributors
416
Owner type
Organization
Primary language
Python
Repository last updated
Sep 29, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?