On this page
Overview
CatBoost is a machine-learning library based on gradient boosting over decision trees. It supports numerical and categorical features for classification, regression, ranking, and related tasks. Python, R, Java, and C++ workflows are available, with CPU, GPU, and multi-GPU training options.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Train directly on tabular data containing categorical features
CatBoost highlights support for both numerical and categorical features, reducing the need to make a separate one-hot encoding pipeline mandatory before GBDT training.
Sources: [1]
Use one library for classification, regression, and ranking
The project and Python package target classification, regression, ranking, and other machine-learning tasks through task-specific estimators, losses, and metrics.
Choose CPU, GPU, or distributed training for the workload
The README documents GPU and multi-GPU training and also points to distributed workflows through Apache Spark and the command-line interface.
Sources: [1]
Best fit
Fits GBDT evaluation on tabular datasets with many categorical values
It is useful for customer, product, event, and similar structured datasets where numerical and categorical columns are mixed and classification, regression, or ranking models are being compared.
Sources: [1]
Before adoption
Validate GPU and distributed environments separately from local CPU use
GPU and Spark execution add hardware, driver, runtime, and cluster requirements beyond a local CPU workflow. Validate the model and feature definitions locally before scaling the execution environment.
Sources: [1]
Source builds require CMake, Conan, Cython, and related tooling
The Python package pyproject declares CMake, Conan, Cython, NumPy, and other build dependencies. Prefer prebuilt packages for normal use and assemble the native toolchain only when source builds are required.
Official sources
- [1]CatBoost v1.2.10 README(2026-10-03)
- [2]CatBoost 1.2.10 Python setup.py(2026-10-03)
- [3]CatBoost 1.2.10 Python pyproject.toml(2026-10-03)
- [4]CatBoost 1.2.10 version metadata(2026-10-03)
- [5]CatBoost 1.2.10 release(2026-10-03)
- [6]CatBoost Apache-2.0 license(2026-10-03)
Supplemental curator note
CatBoost fits tabular workloads where categorical values should be handled without a separate encoding pipeline. Compare accuracy together with training time, inference latency, and model size on the real dataset before choosing CPU/GPU settings or selecting among GBDT libraries.
Try it in 3 steps
- 1
Install CatBoost 1.2.10 in a virtual environment
Install the prebuilt Python package with the stable release pinned.
python3 -m venv .venv && . .venv/bin/activate && python -m pip install catboost==1.2.10 - 2
Train a tiny classification model
Train and predict on a small CPU-only dataset without downloading external data.
python -c "from catboost import CatBoostClassifier; m=CatBoostClassifier(iterations=5,verbose=False); m.fit([[0],[1],[2],[3]],[0,0,1,1]); print(m.predict([[2.5]]))" - 3
Verify the installed version
Confirm the library version so it can be recorded together with model artifacts.
python -c "import catboost; print(catboost.__version__)"
Growth
Growth trends · Last 30 days
9,134 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 112
- Open PRs
- 63
Development activity is still being collected.
Built with
Categories and tags
GitHub data
GitHub dataView detailed GitHub data
GitHub Topics
- machine-learning
- decision-trees
- gradient-boosting
- gbm
- gbdt
- python
- r
- kaggle
- gpu-computing
- catboost
- tutorial
- categorical-features
- Stars
- 9,134
- Forks
- 1,343
- Watchers
- 187
- Open issues
- 673
- Contributors
- 189
- Owner type
- Organization
- Primary language
- C++
- License
- Apache-2.0
- Repository last updated
- Oct 3, 2026
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- ComfyUI135,903 Stars
2 shared tag(s) · 1 shared category(s)
save generative AI pipelines as visible graphs and rerun only what changed
Python - TensorFlow200,675 Stars
4 shared tag(s) · 1 shared category(s) · Same language
machine-learning infrastructure with important Python and dependency changes in 2.21
C++ - LightGBM18,829 Stars
4 shared tag(s) · 1 shared category(s) · Same language
fast and memory-efficient gradient-boosted trees at scale
C++ - scikit-learn67,463 Stars
3 shared tag(s) · 2 shared category(s)
a Python machine-learning library for composing preprocessing, training, model selection, and evaluation through estimators and pipelines
Python - pandas49,908 Stars
3 shared tag(s) · 2 shared category(s)
a Python data-analysis library for reshaping, joining, aggregating, and reading/writing labeled tabular and time-series data with Series and DataFrame
Python - statsmodels11,670 Stars
3 shared tag(s) · 2 shared category(s)
Estimate regressions, time series, and hypothesis tests with coefficients, uncertainty, and diagnostics
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.