🎉 75% of content is free forever — Unlock Premium from $10/mo →
CW
NEWSLIVESearch All Content
đŸ’ŧ Servicesâ„šī¸ Aboutâœ‰ī¸ ContactView Pricing Plansfrom $10

Linear Models - OLS, Ridge, Lasso, Elastic Net, LARS and Bayesian Regression

Core MLLinear ModelsđŸŸĸ Free Lesson

Advertisement

Core ML

The Linear Model Family - From Least Squares to Bayesian Regression

Linear models predict a target as a weighted sum of features. They are fast, interpretable, and the strongest baseline you can build in minutes.

  • Ordinary Least Squares - minimize residual sum of squares with a closed-form solution
  • Ridge, Lasso, Elastic Net - add L2, L1, or mixed penalties to control shrinkage and sparsity
  • LARS, OMP, Bayesian methods - sparse paths, forward selection, and uncertainty estimates

"All models are wrong, but some are useful." - George Box


Prerequisites

Before diving into the linear model family, you should be comfortable with:

  • Linear Regression - OLS, residuals, and the normal equation (see Linear Regression)
  • Regularization - why penalties prevent overfitting (see Regularization)
  • Matrix Algebra - dot products, norms, and the shape conventions of X and y
  • Cross-Validation - k-fold and leave-one-out for tuning hyperparameters
  • Python and Scikit-learn - the fit and predict API

Learning Objectives

By the end of this tutorial, you will be able to:

  1. Write the general linear model and map its symbols to scikit-learn attributes
  2. Choose between OLS, Ridge, Lasso, and Elastic Net from data characteristics
  3. Explain how coordinate descent and soft-thresholding produce sparse coefficients
  4. Select regularization strength with cross-validation or information criteria (AIC, BIC)
  5. Describe LARS, orthogonal matching pursuit, and when each beats plain Lasso
  6. Use Bayesian Ridge and ARD when predictions need uncertainty estimates
  7. Estimate the computational cost of each solver before running it
  8. Avoid the classic pitfalls: unscaled features, un-tuned alpha, and ignored multicollinearity

The Linear Model

Every method in this guide predicts from the same form: the output is a linear combination of the input features.

The vector of weights is called coef_ and the bias term is called intercept_ in scikit-learn. Classification with linear models uses the same structure through a link function - see Logistic Regression for the classification counterpart.


Ordinary Least Squares

LinearRegression fits coefficients that minimize the residual sum of squares between observed and predicted targets.

python
from sklearn import linear_model

reg = linear_model.LinearRegression()
reg.fit([[0, 0], [1, 1], [2, 2]], [0, 1, 2])
print(reg.coef_)       # [0.5 0.5]
print(reg.intercept_)  # 0.0

OLS relies on the independence of features. When features are correlated, the design matrix becomes close to singular and coefficient estimates get unstable with high variance. This situation is called multicollinearity, and it usually appears when data was collected without an experimental design.

Non-Negative Least Squares

When coefficients must be non-negative - frequencies, prices, counts - set positive=True on LinearRegression and non-negative least squares are applied instead.

OLS Complexity

The solution comes from the singular value decomposition of X, costing:

That cost is the reason OLS becomes expensive when features grow into the tens of thousands - which is exactly where regularized solvers take over.


Ridge Regression

Ridge addresses the instability of OLS by penalizing the size of the coefficients.

The complexity parameter alpha controls shrinkage: the larger alpha, the stronger the shrinkage and the more robust the coefficients become to collinearity. Unlike Lasso, Ridge never sets coefficients exactly to zero.

python
from sklearn import linear_model

reg = linear_model.Ridge(alpha=0.5)
reg.fit([[0, 0], [0, 0], [1, 1]], [0, 0.1, 1])
print(reg.coef_)       # [0.34545455 0.34545455]
print(reg.intercept_)  # 0.13636364

Solver Selection

With solver="auto", Ridge checks these conditions top to bottom and picks the first match:

SolverCondition
lbfgspositive=True was requested
choleskyThe input array X is not sparse
sparse_cgNone of the above conditions apply

Ridge for Classification

RidgeClassifier converts binary targets to minus one and plus one, treats the problem as regression, and predicts the class from the sign of the output. For multiclass problems it becomes multi-output regression and picks the class with the highest score.

It can be much faster than LogisticRegression when the number of classes is large, because the projection matrix is computed only once. In practice it often reaches similar cross-validation scores while using a very different - and often cheaper - set of numerical solvers.

Setting Alpha: RidgeCV

RidgeCV builds cross-validation for alpha into the estimator and defaults to efficient leave-one-out cross-validation:

python
import numpy as np
from sklearn import linear_model

reg = linear_model.RidgeCV(alphas=np.logspace(-6, 6, 13))
reg.fit([[0, 0], [0, 0.1], [1, 1]], [0, -0.1, 1])
print(reg.alpha_)  # 0.1

Passing cv=10 switches from leave-one-out to ordinary 10-fold cross-validation.


Lasso

The Lasso estimates sparse coefficients - it can set coefficients exactly to zero. That makes it both a regressor and a built-in feature selector, and it is fundamental to compressed sensing.

python
from sklearn import linear_model

reg = linear_model.Lasso(alpha=0.1)
reg.fit([[0, 0], [1, 1]], [0, 1])
print(reg.predict([[1, 1]]))  # [0.8]

💡

Feature selection with Lasso - because the fitted model is sparse, you can drop every feature with a zero coefficient and retrain on the reduced set. For pure selection workflows, scikit-learn also exposes L1-based selectors in feature selection.

Coordinate Descent and Soft Thresholding

The Lasso solver optimizes one feature at a time, reducing the problem to a single dimension:

where the soft-thresholding function is:

Whenever alpha is at least as large as the absolute value of z, the update is exactly zero - that is where sparsity comes from. The solver cycles through features (selection="cyclic") or picks them randomly (selection="random") and stops when the duality gap falls below tol.

Setting the Regularization Parameter

Two practical routes exist for choosing alpha:

  • Cross-validation: LassoCV runs the full path for each candidate alpha. It is the most common choice, and the preferred one for high-dimensional data with collinear features
  • Information criteria: LassoLarsIC computes the path once and scores it with AIC or BIC, which is much cheaper than k-fold validation

Here L-hat is the maximum likelihood, d the number of parameters, and N the number of samples. Information criteria assume the candidate model is correct and degrade when there are more features than samples - in that regime, prefer LassoCV.

python
from sklearn import linear_model

# X: (n_samples, n_features), y: (n_samples,)
reg = linear_model.LassoCV(cv=5, random_state=0)
reg.fit(X, y)
print(reg.alpha_)  # cross-validated regularization strength
print(reg.coef_)   # sparse coefficient vector

Elastic-Net

ElasticNet trains with both L1 and L2 penalties, mixing them with the l1_ratio parameter (rho in the formula below).

Elastic-Net matters when multiple features are correlated. Lasso picks one of them almost at random, while Elastic-Net tends to select or drop the whole group together - and it inherits Ridge's stability under rotation of the feature space.

python
from sklearn import linear_model

reg = linear_model.ElasticNet(alpha=0.1, l1_ratio=0.7)
reg.fit([[0, 0], [1, 1], [2, 2]], [0, 1, 2])
print(reg.coef_)

Use ElasticNetCV to tune both alpha and l1_ratio by cross-validation.


Multi-Task Lasso and Elastic-Net

When y is a 2D array of shape (n_samples, n_tasks), MultiTaskLasso selects the same features for every task. The penalty mixes L1 across rows with L2 across tasks:

Plain Lasso applied to each task separately scatters non-zero coefficients across different columns; MultiTaskLasso produces full active columns, which is exactly what you want when fitting a time series model that requires an active feature to stay active. MultiTaskElasticNet and MultiTaskElasticNetCV provide the mixed-penalty versions.


Least Angle Regression (LARS)

LARS is a forward method for high-dimensional data. At each step it moves toward the feature most correlated with the current residual; when several features tie, it moves in the equiangular direction between them.

Advantages:

  • Efficient when features greatly outnumber samples
  • As fast as forward selection, with the same complexity order as OLS
  • Produces a full piecewise-linear solution path, ideal for cross-validation
  • Behaves intuitively: two equally correlated features grow at the same rate

Disadvantages:

  • Because it refits residuals iteratively, it is sensitive to noise in small samples

LassoLars runs Lasso through the LARS algorithm and returns the exact piecewise-linear path:

python
from sklearn import linear_model

reg = linear_model.LassoLars(alpha=0.1)
reg.fit([[0, 0], [1, 1]], [0, 1])
print(reg.coef_)  # [0.6 0. ]

The full coefficient path is available from lars_path and lars_path_gram, stored as coef_path_ with one column per regularization step - the first column is always zero.


Orthogonal Matching Pursuit

OrthogonalMatchingPursuit and orthogonal_mp approximate a linear model with a fixed budget of non-zero coefficients:

Alternatively it can target a residual tolerance instead of a coefficient count. OMP is greedy: each iteration adds the atom most correlated with the current residual, then recomputes the residual with an orthogonal projection on the chosen atoms - which makes it more accurate than plain matching pursuit at the same budget.

Use it when you need a strictly bounded number of active features and want deterministic, cheap selection.


Bayesian Regression

Bayesian linear models treat the regularization parameter as a random variable tuned from the data instead of a constant set by hand. Ridge regression with an L2 penalty is equivalent to a maximum-a-posteriori estimate under a Gaussian prior over the coefficients.

Advantages: the model adapts to the data, and regularization is estimated rather than hand-tuned. Disadvantages: fitting is slower because inference optimizes over the prior parameters too.

Bayesian Ridge Regression

BayesianRidge places a spherical Gaussian prior over the coefficients and gamma priors over the precisions. The parameters w, alpha, and lambda are estimated jointly by maximizing the log marginal likelihood.

python
from sklearn import linear_model

X = [[0., 0.], [1., 1.], [2., 2.], [3., 3.]]
Y = [0., 1., 2., 3.]
reg = linear_model.BayesianRidge()
reg.fit(X, Y)
print(reg.predict([[1, 0.]]))  # [0.5]
print(reg.coef_)               # slightly shrunk vs OLS

The weights differ slightly from OLS because of the Bayesian framework, but Bayesian Ridge is noticeably more robust to ill-posed problems.

Automatic Relevance Determination (ARD)

ARDRegression drops the spherical prior for an elliptic one: each coefficient gets its own precision, so some coefficients can collapse to near zero. The result is much sparser than Bayesian Ridge - ARD is also known as Sparse Bayesian Learning or the Relevance Vector Machine.

Use ARD when you suspect only a few features matter but still want posterior uncertainty on the rest.


Other Linear Models in the Toolbox

EstimatorLossBest for
SGDRegressor / SGDClassifierhinge, log, squared errorvery large datasets, online learning
HuberRegressorHuber (robust)targets contaminated by outliers
RANSACRegressorrandom sample consensusstrong outliers in a linear relationship
PassiveAggressiveRegressorpassive-aggressive updatestreaming regression with constant learning rate
PoissonRegressor, GammaRegressor, TweedieRegressorgeneralized linear modelspositive targets with count or skew distributions (rates, costs, durations)

The generalized linear models use a log link by default, so predictions stay positive - the right choice when the target is a non-negative rate rather than an unbounded number.


Choosing the Right Linear Model

Which Linear Model Should You Pick?Features X and target y readyDo you need every feature keptin the final model?yesnoDo you need uncertaintyestimates on predictions?Are there correlatedfeature groups?yesnoyesnoBayesian Ridge / ARDposterior uncertaintyRidge (L2)shrink, keep allElastic NetL1 plus L2 mixLasso (L1)sparse solutionFeatures far outnumber samples? Use LARS, LassoLars or OMP for fast sparse paths

The matching selection table for quick reference:

SituationRecommended model
All features useful, collinear featuresRidge
Sparse model, feature selectionLasso
Correlated groups, stable selectionElastic Net
Multiple targets share featuresMultiTaskLasso
Features far outnumber samplesLARS / LassoLars
Strict budget on non-zero coefficientsOrthogonal Matching Pursuit
Predictions need uncertaintyBayesianRidge / ARD
Huge dataset, streaming updatesSGD with elastic-net penalty
Positive skewed target (rates, costs)Poisson, Gamma or Tweedie regressor

How Coefficients Shrink

Regularization strength moves every coefficient along a path: OLS values on the left, heavy shrinkage on the right. Lasso paths reach zero at a finite alpha, which is where feature selection happens.

Coefficient Paths as Regularization Increasescoeflog(alpha) increasing0Lasso sets this coefficient to zeroRidge: shrinks toward zeroElastic Net: mixed pathLasso: reaches exactly zero

Two practical readings of this picture:

  1. Shrinkage is continuous in alpha - tune alpha on a logarithmic grid rather than guessing a single value
  2. Sparsity is a threshold event - the feature drops out the moment the threshold beats its correlation with the residual, so small alpha changes can flip the selected set near that point

Complexity at a Glance

MethodPer-fit costSolver style
OLSO(n_samples x n_features squared)singular value decomposition
Ridgesame order as OLScholesky, lbfgs or sparse_cg
Lasso / Elastic NetO(n_samples x n_features) per passcoordinate descent over the path
LARSsame order as OLS for the full pathequiangular forward steps
OMPO(n_nonzero x n_samples x n_features)greedy orthogonal projections
Bayesian RidgeOLS cost plus prior optimizationevidence maximization

Rule of thumb: pick the cheapest solver that can express the penalty you need, and let cross-validation decide alpha - not manual trial and error.


Common Pitfalls

  • Forgetting to scale features - penalties are scale-dependent, so standardize inputs (or use a Pipeline with StandardScaler) before Ridge, Lasso, or Elastic Net
  • Leaving alpha at its default - the default is a starting point, not a tuned value; always cross-validate on a log grid
  • Ignoring multicollinearity in OLS - unstable coefficients mean unstable predictions; Ridge or Elastic Net fixes this
  • Expecting Ridge to select features - L2 shrinks but never zeroes; use L1 when you need sparsity
  • Using information criteria with p far larger than n - AIC and BIC break in that regime; use LassoCV instead
  • Regularizing the intercept - it should stay free; scikit-learn estimators exclude it by default

Interview Questions

Q1: What is the difference between Ridge and Lasso?

A: Ridge adds an L2 penalty, shrinking all coefficients toward zero but keeping every feature in the model - it stabilizes estimates under multicollinearity. Lasso adds an L1 penalty, which can drive coefficients exactly to zero, performing automatic feature selection. Use Ridge when all features matter, Lasso when you believe only a subset does.

Q2: Why does Lasso produce sparse solutions while Ridge does not?

A: The L1 constraint region is a diamond whose corners touch the axes, so the loss contours are likely to intersect at a corner where some coefficients are exactly zero. The L2 region is a circle, and intersections on the axes are rare - hence shrinkage without sparsity.

Q3: How do you choose alpha in production?

A: Cross-validate on a logarithmic grid - for example logspace from 1e-4 to 1e4 - with k-fold for large datasets and LassoLarsIC when the path must be computed only once. Wrap scaling and the model in a Pipeline so preprocessing happens inside each fold and cannot leak.

Q4: When would you pick Elastic Net over Lasso?

A: With correlated features. Lasso selects one of a correlated group almost arbitrarily and the choice is unstable across resamples; Elastic Net keeps the group together, producing stable selection and inheriting Ridge's robustness under rotation.

Q5: What is the computational difference between Lasso and LARS?

A: Lasso via coordinate descent sweeps features repeatedly until the duality gap converges, with a cost proportional to the number of passes. LARS constructs the entire solution path in a sequence of at most n_features equiangular steps with OLS-level cost per step, and LassoLars returns the exact piecewise-linear path - valuable when the full path, not one alpha, is the deliverable.

Q6: When do you reach for Bayesian Ridge instead of plain Ridge?

A: When you want the regularization strength estimated from data with uncertainty attached. Bayesian Ridge optimizes the log marginal likelihood over both coefficients and precisions, giving predictive intervals for free - useful for risk-sensitive decisions where knowing the confidence matters as much as the point estimate.

Q7: How does the SVM C parameter relate to alpha?

A: They are reciprocal: alpha equals 1 divided by C, or 1 divided by (n_samples times C), depending on the estimator's exact objective. Larger C (weaker regularization) corresponds to smaller alpha.


Key Takeaways


Further Reading

Official Documentation

Academic Papers

  • "Regression Shrinkage and Selection via the Lasso" - Tibshirani (1996), the original Lasso paper
  • "Regularization and Variable Selection via the Elastic Net" - Zou and Hastie (2005)
  • "Least Angle Regression" - Efron, Hastie, Johnstone and Tibshirani (2004)
  • "Regularization Paths for Generalized Linear Models by Coordinate Descent" - Friedman, Hastie and Tibshirani (2010)
  • "Notes on Regularized Least Squares" - Rifkin and Lippert (2007) on Ridge and leave-one-out

Books

  • "The Elements of Statistical Learning" - Hastie, Tibshirani, Friedman - Chapter 3
  • "An Introduction to Statistical Learning" - James, Witten, Hastie, Tibshirani - Chapter 6
  • "Pattern Recognition and Machine Learning" - Bishop - Bayesian regression and ARD

What to Learn Next

-> Linear Regression The OLS foundation these models extend, with assumptions and gradient descent.

-> Regularization Deeper treatment of why Ridge and Lasso prevent overfitting, with geometry.

-> Logistic Regression The linear model for classification, through the sigmoid link.

—
☆☆☆☆☆
0 ratings

Rate & Feedback

Need Expert Machine Learning Help?

Get personalized tutoring, project support, or professional consulting.

Advertisement