Collect a dataset containing input and target variables.
Identify the dependent and independent variables.
Fit a regression line to the training data.
Calculate the difference between actual and predicted values.
Minimise the prediction error using an optimization method.
Use the trained model to predict values for new data.
After exploring the basic working of linear regression in machine learning, the topic becomes now more clear. Now let’s explore the concepts and methods involved in working with linear regression model in machine learning.
1. Dependent and Independent Variables
Two main types of variables are used in Linear regression:
For example, if you are predicting house prices based on area:
With multiple independent variables, the model could use area, number of bedrooms, location score, and property age to predict the price.
2. Linear Regression Equation
The linear regression in machine learning formula shows how the target variable is estimated from one or more input variables. For simple linear regression, the formula is:
y = β₀ + β₁x + ε
Where:
The equation for predictions, is commonly written as:
ŷ = β₀ + β₁x
Here, ŷ represents the predicted value.
For multiple linear regression in machine learning, the equation becomes:
ŷ = β₀ + β₁x₁ + β₂x₂ + ... + βₙxₙ
Each coefficient represents the expected change in the target variable when its corresponding input changes by one unit, assuming other variables remain constant.
3. How Is the Regression Line Calculated?
The regression line represents the relationship between the input and target variables. The algorithm searches for the line that provides the best fit to the available data.
Consider the equation:
ŷ = 10 + 2x
If x = 5, then:
ŷ = 10 + (2 × 5)
ŷ = 20
Therefore, the model predicts a value of 20.
The slope, 2, means that the predicted target increases by 2 units for every one-unit increase in x.
4. Least Squares Method in Linear Regression
The least squares method is commonly used to determine the best-fitting regression line. It minimises the sum of the squared differences between actual and predicted values.
The residual for each observation is:
Residual = Actual Value − Predicted Value
The sum of squared errors can be represented as:
SSE = Σ(yᵢ − ŷᵢ)²
The algorithm selects the coefficients that minimise this value. Squaring the errors prevents positive and negative errors from cancelling each other and gives greater weight to larger errors.
5. Cost Function in Linear Regression
A cost function measures how accurately the linear regression model predicts the target variable.
One commonly used cost function is Mean Squared Error (MSE):
MSE = (1/n) Σ(yᵢ − ŷᵢ)²
Where:
A lower MSE indicates that predictions are, on average, closer to the actual values.
6. Gradient Descent in Linear Regression
Gradient descent is an optimisation technique that can be used to minimise the cost function.
The algorithm starts with initial values for the model parameters and repeatedly updates them in the direction that reduces the error.
The general update rule is:
θ = θ − α∇J(θ)
Where:
A suitable learning rate helps the algorithm converge efficiently. A very small learning rate can make training slow, while a very large learning rate can cause the optimization process to overshoot the minimum.
Linear regression can be used in different ways depending on the data and the number of input variables. So now, let’s explore the different types of linear regression in machine learning.
Types of Linear Regression in Machine Learning
Linear regression can be classified based on the number of predictors and whether regularization is applied. Below are main the types of linear regression:
1. Simple Linear Regression in Machine Learning
Simple linear regression in machine learning is a method in which a dependent variable is predicted using a single independent variable by fitting a straight line to the data.
Its equation is:
y = β₀ + β₁x + ε
For example, a company could predict monthly sales based only on its advertising expenditure.
Suppose:
Sales = 50 + 4 × Advertising Spend
If advertising spend is 10 units:
Sales = 50 + (4 × 10) = 90
The model therefore predicts sales of 90 units.
Simple linear regression in machine learning is easy to understand and visualize because it involves one predictor and one target variable.
2. Multiple Linear Regression in Machine Learning
Multiple linear regression in machine learning is a method in which a continuous target is predicted using two or more independent variables.
Its equation is:
y = β₀ + β₁x₁ + β₂x₂ + ... + βₙxₙ + ε
For example, house prices could be predicted using:
A multiple regression in machine learning allows several factors to be considered at the same time. However, adding more variables does not always improve the model. Irrelevant or highly correlated variables can increase complexity and reduce model performance.
3. Regularised Linear Regression
Regularisation is a technique in which a penalty is added to the model’s objective function to limit large coefficients. It can help reduce overfitting, especially when a dataset contains many features.
Two common regularised approaches are Ridge and Lasso regression.
(a) Ridge Regression
Ridge regression is a method in which coefficients are generally reduced toward zero without being set exactly to zero. It can be useful when several predictors contribute to the target and multicollinearity is present.
A simplified objective function is:
MSE + λΣβⱼ²
Where λ controls the strength of regularisation.
(b) Lasso Regression
Lasso regression is a method in which some coefficients can be reduced to exactly zero, allowing the model to perform feature selection.
MSE + λΣ|βⱼ|
Ridge and Lasso are useful when a standard linear regression model is at risk of overfitting.
How to Perform Linear Regression in Machine Learning?
Performing linear regression in machine learning generally follows the same five-step workflow, regardless of which tool or programming language you use.
1. Define the Problem and Select Variables
First, identify the dependent variable you want to predict and the independent variable(s) that may influence it. Then, check whether a linear relationship is suitable for the data, as a straight-line model may give poor predictions for clearly non-linear relationships.
For example:
Problem: Predict house prices.
Target variable: House price
Features: Area, bedrooms, location score, and property age
The quality and relevance of the selected features can strongly influence the final model.
2. Prepare and Split the Dataset
Before training, clean the dataset and handle missing values, duplicate records, and incorrect entries.
The dataset is usually divided into:
A common approach is an 80:20 or 70:30 train-test split, although the appropriate split depends on the dataset.
3. Train the Linear Regression Model
Train the model using the training data. The algorithm calculates the coefficients that best fit the data by minimising the cost function using methods such as least squares or gradient descent. This is the stage where the model learns the relationship between the input variables and the target.
4. Evaluate the Model
After training, generate predictions using the test data and evaluate them using appropriate regression metrics such as:
R-squared
Mean Squared Error
Mean Absolute Error
Do not confuse what these metrics means, as we will discuss all these metrics in the next section below.
5. Interpret the Results
Review the model’s coefficients to understand the relationship between the input variables and the target. A positive coefficient increases the prediction, while a negative coefficient decreases it. The coefficient size shows the strength of the effect when the variables are on similar scales.
These are the five basic steps which you will need to perform while training your model by linear regression in machine learning. Now let’s see how to evaluate any machine learning linear regression model.
How to Evaluate a Linear Regression Model?
To understand how accurately a linear regression model in machine learning performs on unseen data, there are some metrics which can be used. You can evaluate your model by using metrics such as:
1. R-Squared
To measure how much of the variation in the dependent variable is explained by the model you can use R-squared (R²).
It can be expressed as:
R² = 1 − (SSres / SStot)
Where:
An R² value closer to 1 generally means that the model explains more of the variation in the target. However, a high R² does not always mean that the model is suitable or will perform well on new data.
2. Mean Squared Error
To calculate the average squared difference between actual and predicted values you can use Mean Squared Error.
MSE = (1/n) Σ(yᵢ − ŷᵢ)²
MAE is expressed in the same unit as the dependent variable, which makes it easier to interpret directly. For example, an MAE of 5,000 on a salary prediction model means the model's predictions are, on average, off by 5,000 currency units. MAE is also less sensitive to outliers than MSE, since it does not square the errors.
3. Mean Absolute Error
To measure the average absolute difference between actual and predicted values you can Mean Absolute Error.
MAE = (1/n) Σ|yᵢ − ŷᵢ|
For example, if the prediction errors are 2, 3, and 5:
MAE = (2 + 3 + 5) / 3 = 3.33
MAE is easy to interpret because it shows the average prediction error in the same units as the target variable.
How to Interpret Linear Regression Results
When interpreting a regression model, consider:
Intercept: Expected target value when all predictors are zero.
Coefficient: Expected change in the target for a one-unit change in a predictor, holding other variables constant.
R²: Proportion of target variation explained by the model.
MAE/MSE: Magnitude of prediction errors.
Residuals: Differences between actual and predicted values.
A coefficient's statistical significance may also be examined when using statistical regression analysis.
Assumptions of Linear Regression in Machine Learning
Before applying linear regression in machine learning, it is important to check whether the data meets certain conditions. These conditions help ensure that the model produces reliable and meaningful results.
1. Linearity
Linearity assumes that a straight-line relationship is present between the independent and dependent variables. If the relationship is curved, accurate predictions may not be produced by a linear model. A scatter plot can be used to check this assumption.
2. Independence of Observations
Independence of observations means that each data point is unrelated to the others. If one observation influences another, the results may be misleading. This assumption is often violated in time series data, where values are related to previous observations.
3. Homoscedasticity
Homoscedasticity means that the variance of the residuals remains roughly constant across all levels of the independent variable. If the variance changes, heteroscedasticity occurs, which may appear as a funnel or cone shape in a residual plot.
4. Normality of Residuals
Normality of residuals means that the model’s errors are approximately normally distributed. This assumption is mainly important for reliable confidence intervals and hypothesis tests. A histogram or Q-Q plot can be used to check it.
5. Absence of Multicollinearity
Absence of multicollinearity means that the independent variables are not strongly correlated with each other. High correlation between variables can make the model’s coefficients unstable and difficult to interpret. VIF is commonly used to detect multicollinearity.
Linear Regression Implementation in Python
In this section you will go through a full linear regression in machine learning python code example and a linear regression in machine learning python implementation, using scikit-learn, one of the most widely used Python libraries for machine learning.
The example predicts salary from years of experience, a classic simple linear regression in machine learning example.
Import the Required Python Libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
Load and Explore the Dataset
np.random.seed(42)
years_experience = np.round(np.random.uniform(0, 12, 60), 1)
salary = 35000 + years_experience * 9500 + np.random.normal(0, 8000, 60)
df = pd.DataFrame({
"YearsExperience": years_experience,
"Salary": salary
})
print(df.head())
print(df.describe())
This creates a sample dataset of 60 employees with their years of experience and salary, with a realistic amount of random noise added so the relationship is linear but not perfectly clean, similar to real-world data.
Visualize the Variables
plt.scatter(df["YearsExperience"], df["Salary"], color="steelblue")
plt.xlabel("Years of Experience")
plt.ylabel("Salary")
plt.title("Years of Experience vs Salary")
plt.show()
Expected Output: