Core ML
The Linear Model Family - From Least Squares to Bayesian Regression
Linear models predict a target as a weighted sum of features. They are fast, interpretable, and the strongest baseline you can build in minutes.
- Ordinary Least Squares - minimize residual sum of squares with a closed-form solution
- Ridge, Lasso, Elastic Net - add L2, L1, or mixed penalties to control shrinkage and sparsity
- LARS, OMP, Bayesian methods - sparse paths, forward selection, and uncertainty estimates
"All models are wrong, but some are useful." - George Box
Prerequisites
Before diving into the linear model family, you should be comfortable with:
- Linear Regression - OLS, residuals, and the normal equation (see Linear Regression)
- Regularization - why penalties prevent overfitting (see Regularization)
- Matrix Algebra - dot products, norms, and the shape conventions of X and y
- Cross-Validation - k-fold and leave-one-out for tuning hyperparameters
- Python and Scikit-learn - the fit and predict API
Learning Objectives
By the end of this tutorial, you will be able to:
- Write the general linear model and map its symbols to scikit-learn attributes
- Choose between OLS, Ridge, Lasso, and Elastic Net from data characteristics
- Explain how coordinate descent and soft-thresholding produce sparse coefficients
- Select regularization strength with cross-validation or information criteria (AIC, BIC)
- Describe LARS, orthogonal matching pursuit, and when each beats plain Lasso
- Use Bayesian Ridge and ARD when predictions need uncertainty estimates
- Estimate the computational cost of each solver before running it
- Avoid the classic pitfalls: unscaled features, un-tuned alpha, and ignored multicollinearity
The Linear Model
Every method in this guide predicts from the same form: the output is a linear combination of the input features.
The vector of weights is called coef_ and the bias term is called intercept_ in scikit-learn. Classification with linear models uses the same structure through a link function - see Logistic Regression for the classification counterpart.
Ordinary Least Squares
LinearRegression fits coefficients that minimize the residual sum of squares between observed and predicted targets.
from sklearn import linear_model
reg = linear_model.LinearRegression()
reg.fit([[0, 0], [1, 1], [2, 2]], [0, 1, 2])
print(reg.coef_) # [0.5 0.5]
print(reg.intercept_) # 0.0
OLS relies on the independence of features. When features are correlated, the design matrix becomes close to singular and coefficient estimates get unstable with high variance. This situation is called multicollinearity, and it usually appears when data was collected without an experimental design.
Non-Negative Least Squares
When coefficients must be non-negative - frequencies, prices, counts - set positive=True on LinearRegression and non-negative least squares are applied instead.
OLS Complexity
The solution comes from the singular value decomposition of X, costing:
That cost is the reason OLS becomes expensive when features grow into the tens of thousands - which is exactly where regularized solvers take over.
Ridge Regression
Ridge addresses the instability of OLS by penalizing the size of the coefficients.
The complexity parameter alpha controls shrinkage: the larger alpha, the stronger the shrinkage and the more robust the coefficients become to collinearity. Unlike Lasso, Ridge never sets coefficients exactly to zero.
from sklearn import linear_model
reg = linear_model.Ridge(alpha=0.5)
reg.fit([[0, 0], [0, 0], [1, 1]], [0, 0.1, 1])
print(reg.coef_) # [0.34545455 0.34545455]
print(reg.intercept_) # 0.13636364
Solver Selection
With solver="auto", Ridge checks these conditions top to bottom and picks the first match:
| Solver | Condition |
|---|---|
| lbfgs | positive=True was requested |
| cholesky | The input array X is not sparse |
| sparse_cg | None of the above conditions apply |
Ridge for Classification
RidgeClassifier converts binary targets to minus one and plus one, treats the problem as regression, and predicts the class from the sign of the output. For multiclass problems it becomes multi-output regression and picks the class with the highest score.
It can be much faster than LogisticRegression when the number of classes is large, because the projection matrix is computed only once. In practice it often reaches similar cross-validation scores while using a very different - and often cheaper - set of numerical solvers.
Setting Alpha: RidgeCV
RidgeCV builds cross-validation for alpha into the estimator and defaults to efficient leave-one-out cross-validation:
import numpy as np
from sklearn import linear_model
reg = linear_model.RidgeCV(alphas=np.logspace(-6, 6, 13))
reg.fit([[0, 0], [0, 0.1], [1, 1]], [0, -0.1, 1])
print(reg.alpha_) # 0.1
Passing cv=10 switches from leave-one-out to ordinary 10-fold cross-validation.
Lasso
The Lasso estimates sparse coefficients - it can set coefficients exactly to zero. That makes it both a regressor and a built-in feature selector, and it is fundamental to compressed sensing.
from sklearn import linear_model
reg = linear_model.Lasso(alpha=0.1)
reg.fit([[0, 0], [1, 1]], [0, 1])
print(reg.predict([[1, 1]])) # [0.8]
đĄ
Feature selection with Lasso - because the fitted model is sparse, you can drop every feature with a zero coefficient and retrain on the reduced set. For pure selection workflows, scikit-learn also exposes L1-based selectors in feature selection.
Coordinate Descent and Soft Thresholding
The Lasso solver optimizes one feature at a time, reducing the problem to a single dimension:
where the soft-thresholding function is:
Whenever alpha is at least as large as the absolute value of z, the update is exactly zero - that is where sparsity comes from. The solver cycles through features (selection="cyclic") or picks them randomly (selection="random") and stops when the duality gap falls below tol.
Setting the Regularization Parameter
Two practical routes exist for choosing alpha:
- Cross-validation:
LassoCVruns the full path for each candidate alpha. It is the most common choice, and the preferred one for high-dimensional data with collinear features - Information criteria:
LassoLarsICcomputes the path once and scores it with AIC or BIC, which is much cheaper than k-fold validation
Here L-hat is the maximum likelihood, d the number of parameters, and N the number of samples. Information criteria assume the candidate model is correct and degrade when there are more features than samples - in that regime, prefer LassoCV.
from sklearn import linear_model
# X: (n_samples, n_features), y: (n_samples,)
reg = linear_model.LassoCV(cv=5, random_state=0)
reg.fit(X, y)
print(reg.alpha_) # cross-validated regularization strength
print(reg.coef_) # sparse coefficient vector
Elastic-Net
ElasticNet trains with both L1 and L2 penalties, mixing them with the l1_ratio parameter (rho in the formula below).
Elastic-Net matters when multiple features are correlated. Lasso picks one of them almost at random, while Elastic-Net tends to select or drop the whole group together - and it inherits Ridge's stability under rotation of the feature space.
from sklearn import linear_model
reg = linear_model.ElasticNet(alpha=0.1, l1_ratio=0.7)
reg.fit([[0, 0], [1, 1], [2, 2]], [0, 1, 2])
print(reg.coef_)
Use ElasticNetCV to tune both alpha and l1_ratio by cross-validation.
Multi-Task Lasso and Elastic-Net
When y is a 2D array of shape (n_samples, n_tasks), MultiTaskLasso selects the same features for every task. The penalty mixes L1 across rows with L2 across tasks:
Plain Lasso applied to each task separately scatters non-zero coefficients across different columns; MultiTaskLasso produces full active columns, which is exactly what you want when fitting a time series model that requires an active feature to stay active. MultiTaskElasticNet and MultiTaskElasticNetCV provide the mixed-penalty versions.
Least Angle Regression (LARS)
LARS is a forward method for high-dimensional data. At each step it moves toward the feature most correlated with the current residual; when several features tie, it moves in the equiangular direction between them.
Advantages:
- Efficient when features greatly outnumber samples
- As fast as forward selection, with the same complexity order as OLS
- Produces a full piecewise-linear solution path, ideal for cross-validation
- Behaves intuitively: two equally correlated features grow at the same rate
Disadvantages:
- Because it refits residuals iteratively, it is sensitive to noise in small samples
LassoLars runs Lasso through the LARS algorithm and returns the exact piecewise-linear path:
from sklearn import linear_model
reg = linear_model.LassoLars(alpha=0.1)
reg.fit([[0, 0], [1, 1]], [0, 1])
print(reg.coef_) # [0.6 0. ]
The full coefficient path is available from lars_path and lars_path_gram, stored as coef_path_ with one column per regularization step - the first column is always zero.
Orthogonal Matching Pursuit
OrthogonalMatchingPursuit and orthogonal_mp approximate a linear model with a fixed budget of non-zero coefficients:
Alternatively it can target a residual tolerance instead of a coefficient count. OMP is greedy: each iteration adds the atom most correlated with the current residual, then recomputes the residual with an orthogonal projection on the chosen atoms - which makes it more accurate than plain matching pursuit at the same budget.
Use it when you need a strictly bounded number of active features and want deterministic, cheap selection.
Bayesian Regression
Bayesian linear models treat the regularization parameter as a random variable tuned from the data instead of a constant set by hand. Ridge regression with an L2 penalty is equivalent to a maximum-a-posteriori estimate under a Gaussian prior over the coefficients.
Advantages: the model adapts to the data, and regularization is estimated rather than hand-tuned. Disadvantages: fitting is slower because inference optimizes over the prior parameters too.
Bayesian Ridge Regression
BayesianRidge places a spherical Gaussian prior over the coefficients and gamma priors over the precisions. The parameters w, alpha, and lambda are estimated jointly by maximizing the log marginal likelihood.
from sklearn import linear_model
X = [[0., 0.], [1., 1.], [2., 2.], [3., 3.]]
Y = [0., 1., 2., 3.]
reg = linear_model.BayesianRidge()
reg.fit(X, Y)
print(reg.predict([[1, 0.]])) # [0.5]
print(reg.coef_) # slightly shrunk vs OLS
The weights differ slightly from OLS because of the Bayesian framework, but Bayesian Ridge is noticeably more robust to ill-posed problems.
Automatic Relevance Determination (ARD)
ARDRegression drops the spherical prior for an elliptic one: each coefficient gets its own precision, so some coefficients can collapse to near zero. The result is much sparser than Bayesian Ridge - ARD is also known as Sparse Bayesian Learning or the Relevance Vector Machine.
Use ARD when you suspect only a few features matter but still want posterior uncertainty on the rest.
Other Linear Models in the Toolbox
| Estimator | Loss | Best for |
|---|---|---|
SGDRegressor / SGDClassifier | hinge, log, squared error | very large datasets, online learning |
HuberRegressor | Huber (robust) | targets contaminated by outliers |
RANSACRegressor | random sample consensus | strong outliers in a linear relationship |
PassiveAggressiveRegressor | passive-aggressive update | streaming regression with constant learning rate |
PoissonRegressor, GammaRegressor, TweedieRegressor | generalized linear models | positive targets with count or skew distributions (rates, costs, durations) |
The generalized linear models use a log link by default, so predictions stay positive - the right choice when the target is a non-negative rate rather than an unbounded number.
Choosing the Right Linear Model
The matching selection table for quick reference:
| Situation | Recommended model |
|---|---|
| All features useful, collinear features | Ridge |
| Sparse model, feature selection | Lasso |
| Correlated groups, stable selection | Elastic Net |
| Multiple targets share features | MultiTaskLasso |
| Features far outnumber samples | LARS / LassoLars |
| Strict budget on non-zero coefficients | Orthogonal Matching Pursuit |
| Predictions need uncertainty | BayesianRidge / ARD |
| Huge dataset, streaming updates | SGD with elastic-net penalty |
| Positive skewed target (rates, costs) | Poisson, Gamma or Tweedie regressor |
How Coefficients Shrink
Regularization strength moves every coefficient along a path: OLS values on the left, heavy shrinkage on the right. Lasso paths reach zero at a finite alpha, which is where feature selection happens.
Two practical readings of this picture:
- Shrinkage is continuous in alpha - tune alpha on a logarithmic grid rather than guessing a single value
- Sparsity is a threshold event - the feature drops out the moment the threshold beats its correlation with the residual, so small alpha changes can flip the selected set near that point
Complexity at a Glance
| Method | Per-fit cost | Solver style |
|---|---|---|
| OLS | O(n_samples x n_features squared) | singular value decomposition |
| Ridge | same order as OLS | cholesky, lbfgs or sparse_cg |
| Lasso / Elastic Net | O(n_samples x n_features) per pass | coordinate descent over the path |
| LARS | same order as OLS for the full path | equiangular forward steps |
| OMP | O(n_nonzero x n_samples x n_features) | greedy orthogonal projections |
| Bayesian Ridge | OLS cost plus prior optimization | evidence maximization |
Rule of thumb: pick the cheapest solver that can express the penalty you need, and let cross-validation decide alpha - not manual trial and error.
Common Pitfalls
- Forgetting to scale features - penalties are scale-dependent, so standardize inputs (or use a Pipeline with StandardScaler) before Ridge, Lasso, or Elastic Net
- Leaving alpha at its default - the default is a starting point, not a tuned value; always cross-validate on a log grid
- Ignoring multicollinearity in OLS - unstable coefficients mean unstable predictions; Ridge or Elastic Net fixes this
- Expecting Ridge to select features - L2 shrinks but never zeroes; use L1 when you need sparsity
- Using information criteria with p far larger than n - AIC and BIC break in that regime; use LassoCV instead
- Regularizing the intercept - it should stay free; scikit-learn estimators exclude it by default
Interview Questions
Q1: What is the difference between Ridge and Lasso?
A: Ridge adds an L2 penalty, shrinking all coefficients toward zero but keeping every feature in the model - it stabilizes estimates under multicollinearity. Lasso adds an L1 penalty, which can drive coefficients exactly to zero, performing automatic feature selection. Use Ridge when all features matter, Lasso when you believe only a subset does.
Q2: Why does Lasso produce sparse solutions while Ridge does not?
A: The L1 constraint region is a diamond whose corners touch the axes, so the loss contours are likely to intersect at a corner where some coefficients are exactly zero. The L2 region is a circle, and intersections on the axes are rare - hence shrinkage without sparsity.
Q3: How do you choose alpha in production?
A: Cross-validate on a logarithmic grid - for example logspace from 1e-4 to 1e4 - with k-fold for large datasets and LassoLarsIC when the path must be computed only once. Wrap scaling and the model in a Pipeline so preprocessing happens inside each fold and cannot leak.
Q4: When would you pick Elastic Net over Lasso?
A: With correlated features. Lasso selects one of a correlated group almost arbitrarily and the choice is unstable across resamples; Elastic Net keeps the group together, producing stable selection and inheriting Ridge's robustness under rotation.
Q5: What is the computational difference between Lasso and LARS?
A: Lasso via coordinate descent sweeps features repeatedly until the duality gap converges, with a cost proportional to the number of passes. LARS constructs the entire solution path in a sequence of at most n_features equiangular steps with OLS-level cost per step, and LassoLars returns the exact piecewise-linear path - valuable when the full path, not one alpha, is the deliverable.
Q6: When do you reach for Bayesian Ridge instead of plain Ridge?
A: When you want the regularization strength estimated from data with uncertainty attached. Bayesian Ridge optimizes the log marginal likelihood over both coefficients and precisions, giving predictive intervals for free - useful for risk-sensitive decisions where knowing the confidence matters as much as the point estimate.
Q7: How does the SVM C parameter relate to alpha?
A: They are reciprocal: alpha equals 1 divided by C, or 1 divided by (n_samples times C), depending on the estimator's exact objective. Larger C (weaker regularization) corresponds to smaller alpha.
Key Takeaways
Further Reading
Official Documentation
- Scikit-learn Linear Models User Guide - the reference this lesson follows
- Scikit-learn API: sklearn.linear_model
Academic Papers
- "Regression Shrinkage and Selection via the Lasso" - Tibshirani (1996), the original Lasso paper
- "Regularization and Variable Selection via the Elastic Net" - Zou and Hastie (2005)
- "Least Angle Regression" - Efron, Hastie, Johnstone and Tibshirani (2004)
- "Regularization Paths for Generalized Linear Models by Coordinate Descent" - Friedman, Hastie and Tibshirani (2010)
- "Notes on Regularized Least Squares" - Rifkin and Lippert (2007) on Ridge and leave-one-out
Books
- "The Elements of Statistical Learning" - Hastie, Tibshirani, Friedman - Chapter 3
- "An Introduction to Statistical Learning" - James, Witten, Hastie, Tibshirani - Chapter 6
- "Pattern Recognition and Machine Learning" - Bishop - Bayesian regression and ARD
What to Learn Next
-> Linear Regression The OLS foundation these models extend, with assumptions and gradient descent.
-> Regularization Deeper treatment of why Ridge and Lasso prevent overfitting, with geometry.
-> Logistic Regression The linear model for classification, through the sigmoid link.