OSS Tanbou

share columnar data through a common in-memory format to reduce conversion and copying across analytics systems

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
17,172
Primary language
C++
License
Apache-2.0
Repository last updated
Oct 2, 2026
On this page

Overview

Apache Arrow is a language-independent columnar memory format and a set of libraries for efficient data interchange and in-memory analytics. Arrow Columnar Format, IPC, Flight, and implementations such as C++, Python, and R let data frames, databases, and analytics engines exchange data without repeatedly converting it into proprietary intermediate layouts.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Exchange data through a common columnar memory layout

Arrow Columnar Format defines a language-independent in-memory representation for numeric, string, nested, and other data types. When multiple systems understand the same layout, they can avoid many format conversions at process boundaries.

Sources: [1][2]

Use IPC and Flight across processes and networks

Arrow IPC serializes Arrow data and metadata for interchange between processes. Arrow Flight builds an RPC protocol on the IPC format for services such as storage systems and databases that need to move Arrow data over a network.

Sources: [1][2]

Work with the same data model from C++, Python, R, and other languages

This repository contains C++, Python, R, Ruby, and related implementations with column, table, compute, IPC, and Parquet capabilities. Several other implementations, including .NET, Go, Java, JavaScript, and Rust, are maintained in separate repositories.

Sources: [1]

Best fit

Fits ETL, analytics, and database boundaries where conversion cost matters

Arrow is relevant when large datasets move between systems such as Python and C++, data frames and query engines, or databases and analytics services. Longer stretches that can retain a shared columnar representation reduce the need for intermediate conversion.

Sources: [1][2]

Before adoption

It does not replace persistent storage formats or databases

Arrow primarily defines in-memory representation and interchange. Persistent file formats such as Parquet and databases serve different roles, and Arrow libraries commonly read and write Parquet as part of a larger pipeline.

Sources: [1]

Implementations and release cycles vary by language

The Arrow project is not delivered as one package version for every language. Go, Java, Rust, and other implementations have separate repositories and releases, so verify the official package and compatibility requirements for the language you use.

Sources: [1][3]

25.0.1 is a patch release following 25.0.0

The 25.0.1 release published on August 10, 2026 primarily contains fixes and targeted improvements after 25.0.0. Pin the relevant language package and review its release notes before adoption.

Sources: [4][3]

Official sources

  1. [1]Apache Arrow README(2026-10-03)
  2. [2]Arrow Columnar Format(2026-10-03)
  3. [3]Apache Arrow installation(2026-10-03)
  4. [4]Apache Arrow 25.0.1 release(2026-10-03)
  5. [5]Apache Arrow Apache-2.0 license(2026-10-03)
Supplemental curator note

When investigating analytics performance, measure conversion and serialization costs at system boundaries as well as individual operators. Arrow is a useful comparison point when those boundaries can share a common columnar representation.

Try it in 3 steps

  1. 1

    Install PyArrow 25.0.1

    Add the current stable PyArrow package to your Python environment. Activate a virtual environment first if you use one.

    python -m pip install pyarrow==25.0.1
  2. 2

    Create a small Arrow Table

    Create an in-memory columnar table and inspect its schema and values.

    python -c "import pyarrow as pa; t=pa.table({'name':['a','b'],'value':[1,2]}); print(t)"
  3. 3

    Write and read an IPC stream

    Serialize the table to Arrow IPC and read it back with the same schema.

    python -c "import pyarrow as pa; t=pa.table({'x':[1,2,3]}); sink=pa.BufferOutputStream(); w=pa.ipc.new_stream(sink,t.schema); w.write_table(t); w.close(); print(pa.ipc.open_stream(sink.getvalue()).read_all())"
Check the official README

Growth

Growth trends · Last 30 days

17,172 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
148
Open PRs
382

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • arrow
  • parquet
Stars
17,172
Forks
4,339
Watchers
348
Open issues
2,081
Contributors
371
Owner type
Organization
Primary language
C++
License
Apache-2.0
Repository last updated
Oct 2, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?