Step 1: Define the Decision Boundary
For a linearly separable dataset, SVM represents the decision boundary using a hyperplane. Its general equation is:
wᵀx + b = 0
Here, w is the weight vector perpendicular to the hyperplane, x represents the input features, and b is the bias term.
In two dimensions, this hyperplane appears as a straight line. With more features, it becomes a boundary in a higher-dimensional feature space.
Step 2: Set the Class Constraints
For binary classification, the support vector machine assigns the two classes the labels +1 and −1.
The model aims to position the hyperplane so that observations from the positive class fall on one side and observations from the negative class fall on the other. This can be expressed as:
yᵢ(wᵀxᵢ + b) ≥ 1
This constraint ensures that correctly classified observations are positioned on the appropriate side of the boundary.
Step 3: Identify the Support Vectors
Once possible separating boundaries are considered, the support vector machine algorithm pays particular attention to the observations closest to the hyperplane. These are the support vectors.
Support vectors are important because they determine the position of the optimal boundary. Moving or removing these points can change the resulting hyperplane.
Step 4: Maximise the Margin
The margin represents the gap between the decision boundary and the closest observations from the two classes.
SVM aims to maximise this gap because a wider margin can provide better separation between classes. Mathematically, the margin is:
Margin = 2 / ||w||
Therefore, maximising the margin is equivalent to minimising the magnitude of the weight vector. The optimisation problem for a hard-margin support vector machine can be written as:
Minimise: ½||w||²
Subject to:
yᵢ(wᵀxᵢ + b) ≥ 1
The resulting hyperplane is the one that provides the maximum possible margin while satisfying the class constraints.
Step 5: Handle Overlapping Data With a Soft Margin
Real-world datasets are not always perfectly separable. Some observations may overlap or fall on the wrong side of the ideal boundary.
In such cases, support vector machine can use a soft margin, which allows controlled violations of the margin. Slack variables, represented by ξ (xi), quantify these violations.
This gives the model flexibility to handle noisy or overlapping data instead of forcing an unrealistic perfect separation.
Step 6: Control Errors Using the C Parameter
The C hyperparameter controls how strongly support vector machines penalize classification errors and margin violations.
Large C: Places a higher penalty on errors, encouraging the model to classify training observations correctly. This can produce a narrower margin and may increase overfitting risk.
Small C: Allows more margin violations in exchange for a wider margin, which can support better generalisation.
Thus, support vector machine in machine learning balances two objectives: creating a wide margin and limiting classification errors.
SVM can be used to solve different types of machine learning problems. Depending on the nature of the data and the task, SVMs can be implemented in different ways. Let us now look at the major types of Support Vector Machines.
Types of Support Vector Machines
Support Vector Machines can be categorised based on the prediction task and the type of decision boundary they use.
1. Support Vector Classification (SVC)
Support Vector Classification, or SVC, is the standard and most common form of support vector machine, used when the target variable is categorical. The algorithm finds a decision boundary that separates the classes while attempting to maximise the margin.
For example, an SVC model could classify:
2. Support Vector Regression (SVR)
Support Vector Regression, or SVR, adapts the same core idea, maximising a margin around a boundary, to continuous target variables instead of categories.
Rather than trying to keep data points as far as possible from a separating hyperplane, SVR tries to fit a function such that as many data points as possible fall within a specified margin, or tube, around that function, while still keeping the function as flat, or simple, as possible.
3. Linear SVM
A linear SVM is used when the classes in the data can be separated reasonably well by a straight line, or a flat hyperplane in higher dimensions, without needing any additional transformation of the data.
Linear SVM is computationally simpler and faster to train than kernel-based variants, and it is a sensible starting point whenever the underlying relationship between features and classes appears roughly linear.
4. Kernel-Based SVM
When a linear boundary cannot effectively separate the data, SVM can use kernel functions.
A support vector machine kernel allows the algorithm to work with nonlinear relationships by implicitly mapping the original data into a higher-dimensional feature space.
Common kernels include:
Kernel | Typical Use |
Linear | Data that can be separated using a linear boundary |
Polynomial | Problems involving polynomial relationships |
RBF | General nonlinear classification and regression |
Sigmoid | Certain nonlinear modelling problems |
The RBF, or Radial Basis Function, kernel is commonly used when the relationship between features and classes is nonlinear.
Hard Margin vs Soft Margin Support Vector Machine
Support Vector Machine in Machine Learning can use either a hard-margin or soft-margin approach depending on whether classification errors are permitted.
What Is Hard Margin SVM?
Hard-margin SVM requires all training observations to be correctly classified and assumes that the classes can be completely separated by a hyperplane. It attempts to maximise the margin while allowing no classification violations.
However, real-world datasets often contain noise, outliers, or overlapping classes. In such cases, a hard margin may not be practical.
What Is Soft Margin SVM?
Soft-margin SVM allows some observations to violate the margin or even be classified incorrectly.
A penalty parameter, commonly represented by C, controls the trade-off between maximising the margin and penalising classification errors.
It is also important to view support vector machine algorithm in action. In the next section we will implement a support vector machine in Python using Scikit-learn.
How to Implement Support Vector Machine in Python
In this section you will find a complete support vector machine python example using scikit-learn, classifying two overlapping groups of data points, a support vector machine example chosen specifically to show how SVM handles classes that are not perfectly separable.
Import the Required Libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix
Load and Explore the Dataset
np.random.seed(42)
n = 200
class0 = np.random.normal(loc=[3, 3], scale=1.6, size=(n // 2, 2))
class1 = np.random.normal(loc=[5.5, 5.5], scale=1.6, size=(n // 2, 2))
X = np.vstack([class0, class1])
y = np.array([0] (n // 2) + [1] (n // 2))
df = pd.DataFrame(X, columns=["Feature1", "Feature2"])
df["Label"] = y
print(df.head())
print(df["Label"].value_counts())
This creates two overlapping clusters of 100 points each, deliberately generated with enough spread that the classes are not perfectly separable, similar to how real-world classification data usually looks.
Prepare and Scale the Data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
Scaling the features before training matters more for SVM than for many other algorithms, since the algorithm's distance-based margin calculation is directly affected by the scale of each feature, and unscaled features can cause one variable to unfairly dominate the decision boundary.
Split the Dataset
The dataset was already split into training and test sets in the step above, keeping the class proportions balanced between both sets using stratify=y.
Train the SVM Model
model = SVC(kernel="rbf", C=1.0, gamma="scale")
model.fit(X_train_scaled, y_train)
print("Number of support vectors per class:", model.n_support_)
Here, kernel="rbf" selects the radial basis function kernel covered earlier, C=1.0 sets a moderate soft margin penalty, and gamma="scale" controls how far the influence of a single training example reaches, a parameter specific to the RBF kernel.
Make Predictions
y_pred = model.predict(X_test_scaled)
Evaluate the Model
python
print("Accuracy:", accuracy_score(y_test, y_pred))
print("Precision:", precision_score(y_test, y_pred))
print("Recall:", recall_score(y_test, y_pred))
print("F1-score:", f1_score(y_test, y_pred))
print("Confusion matrix:\n", confusion_matrix(y_test, y_pred))
Running this example produces an accuracy of roughly 0.90, a precision of roughly 0.94, a recall of roughly 0.85, and an F1-score of roughly 0.89, with 28 and 27 support vectors identified for the two classes respectively. These numbers reflect exactly the kind of tradeoff a soft margin SVM is designed to make: since the two clusters genuinely overlap at their edges, a small number of misclassifications are accepted in exchange for a stable, generalisable decision boundary.
Visualise the Decision Boundary
xx, yy = np.meshgrid(
np.linspace(X_train_scaled[:, 0].min() - 1, X_train_scaled[:, 0].max() + 1, 300),
np.linspace(X_train_scaled[:, 1].min() - 1, X_train_scaled[:, 1].max() + 1, 300)
)
Z = model.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
plt.contourf(xx, yy, Z, alpha=0.2, cmap="coolwarm")
plt.scatter(
X_train_scaled[:, 0], X_train_scaled[:, 1],
c=y_train, cmap="coolwarm", edgecolors="k"
)
plt.scatter(
model.support_vectors_[:, 0], model.support_vectors_[:, 1],
s=100, facecolors="none", edgecolors="black", linewidths=1.5,
label="Support Vectors"
)
plt.xlabel("Feature 1 (scaled)")
plt.ylabel("Feature 2 (scaled)")
plt.title("SVM Decision Boundary with RBF Kernel")
plt.legend()
plt.show()
Expected Output: