Scikit-learn
Classical machine learning with one consistent fit/predict interface across every algorithm.
What it is
Scikit-learn is a Python library for machine learning that provides simple and efficient tools for data mining, analysis, and modeling. It includes algorithms for classification, regression, clustering, dimensionality reduction, and model evaluation.
Scikit-learn provides consistent APIs for different machine learning algorithms. You can fit models to data, predict outcomes, evaluate performance, and preprocess datasets with transformers and pipelines.
- Best known for
- The uniform estimator API — every model has .fit() and .predict()
- Licence
- BSD 3-clause
- Rule of thumb
- Start here for any tabular problem, even if you expect to move on
When to use it
The question documentation cannot answer for you — because it cannot recommend something else.
Reach for it when
- Tabular data problems — classification, regression, clustering, dimensionality reduction
- You want a strong baseline before considering deep learning
- Building reproducible pipelines with preprocessing and cross-validation
Look elsewhere when
- The problem involves images, audio or language, where deep learning dominates
- You are chasing maximum accuracy on structured data — gradient boosting libraries usually win
Installation
pip install scikit-learnGetting started
The smallest useful thing you can do with it, and what each part means.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
print(model.score(X_test, y_test))from sklearn.cluster import KMeans
import numpy as np
X = np.array([[1,2],[1,4],[1,0],[10,2],[10,4],[10,0]])
kmeans = KMeans(n_clusters=2, random_state=0).fit(X)
print(kmeans.labels_)Advanced usage
Where the library earns its place over a simpler alternative.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
pipeline = Pipeline([('scaler', StandardScaler()), ('svc', SVC())])
pipeline.fit(X_train, y_train)
print(pipeline.score(X_test, y_test))from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5)
print(scores)from sklearn.feature_selection import SelectKBest, f_classif
selector = SelectKBest(f_classif, k=2)
X_new = selector.fit_transform(X, y)
print(X_new[:5])Errors and fixes
The failures you are most likely to hit, and what actually resolves them.
- ConvergenceWarning
- Increase `max_iter` or scale features when fitting models like LogisticRegression.
- ValueError: Input contains NaN
- Handle missing values using `SimpleImputer` or other preprocessing techniques.
Best practices
- Scale or normalize features when using distance-based algorithms.
- Use pipelines to combine preprocessing and modeling steps.
- Split data into training, validation, and test sets to avoid overfitting.
- Leverage cross-validation for robust performance evaluation.
- Understand the assumptions of each algorithm before applying it.
Alternatives
Comparable options, and the reason you would pick one over the other.
Background
Why it exists, and what it was reacting to.
Scikit-learn was initially developed by David Cournapeau in 2007 and has since evolved into a widely used open-source machine learning library maintained by a large community. It is built on top of NumPy, SciPy, and Matplotlib, and is widely adopted in both academia and industry for machine learning tasks.
