The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels
Learners will apply RBF and polynomial kernels to a non-linearly separable dataset, and understand that kernels implicitly project data to higher dimensions.
The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels is a free Machine Learning Academy lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Machine Learning Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
The Problem: Non-Linear Data
Many real-world classification problems are not linearly separable — no straight line (or hyperplane) can correctly separate the classes. For example, data arranged in concentric rings cannot be separated by any linear boundary. One approach is to manually create new features (e.g., x², x×y) that make the classes linearly separable in the augmented space. The kernel trick does this automatically and implicitly, without ever computing the coordinates in the high-dimensional space.
Feature Maps: Lifting Data to Higher Dimensions
A feature map φ(x) transforms an input vector into a higher-dimensional representation. For example, φ([x₁, x₂]) = [x₁², √2·x₁x₂, x₂²] maps 2D data to 3D. After this mapping, classes that overlapped in 2D may become linearly separable in 3D. The SVM then finds a maximum-margin hyperplane in the transformed space. The corresponding decision boundary in the original 2D space is a curve, giving the SVM non-linear classification ability.
The Kernel Trick: Avoiding Explicit Feature Maps
Computing φ(x) explicitly is expensive or even impossible (some feature maps produce infinite-dimensional vectors). The key insight is that the SVM dual formulation only needs dot products φ(xᵢ)·φ(xⱼ), not the individual feature vectors. A kernel function K(xᵢ, xⱼ) computes this dot product directly from the original inputs without ever constructing φ(xᵢ). This is the kernel trick: expensive high-dimensional dot products computed cheaply in input space.
Polynomial Kernel
The polynomial kernel is defined as K(xᵢ, xⱼ) = (γ · xᵢ·xⱼ + r)^d, where d is the polynomial degree, γ is a scaling factor, and r is the coef0 parameter. A degree-2 polynomial kernel implicitly creates all pairwise interactions (x₁x₂) and squared terms (x₁²). Higher degrees create more complex boundaries but risk overfitting. In scikit-learn, use SVC(kernel='poly', degree=3).
from sklearn.svm import SVC
from sklearn.datasets import make_moons
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = make_moons(n_samples=300, noise=0.15, random_state=42)
for degree in [2, 3, 5]:
model = make_pipeline(StandardScaler(), SVC(kernel='poly', degree=degree, C=5))
score = cross_val_score(model, X, y, cv=5).mean()
print(f'Polynomial degree={degree}: CV accuracy={score:.4f}')RBF Kernel: The Default Workhorse
The Radial Basis Function (RBF) kernel, also called the Gaussian kernel, is defined as K(xᵢ, xⱼ) = exp(-γ · ||xᵢ - xⱼ||²). It measures similarity based on distance: nearby points have kernel value close to 1, distant points close to 0. The RBF kernel corresponds to an infinite-dimensional feature map, giving the SVM unlimited expressive power. It is the default kernel in scikit-learn's SVC and works well on most datasets with proper tuning of C and γ.
from sklearn.svm import SVC
from sklearn.datasets import make_moons
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = make_moons(n_samples=300, noise=0.15, random_state=42)
model = make_pipeline(StandardScaler(), SVC(kernel='rbf', C=1.0, gamma='scale'))
scores = cross_val_score(model, X, y, cv=5)
print('RBF SVM CV accuracy:', round(scores.mean(), 4))The Gamma Parameter in RBF Kernel
The gamma parameter controls how far the influence of a single training example reaches. A small gamma makes each point's influence extend far — the decision boundary is smooth and the model underfits (high bias). A large gamma makes influence drop off steeply — the boundary wraps tightly around individual training points (high variance, overfitting). scikit-learn defaults: gamma='scale' (uses 1/(n_features × X.var())) or gamma='auto' (uses 1/n_features). Always tune C and gamma together.
from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = load_breast_cancer(return_X_y=True)
for gamma in [0.0001, 0.001, 0.01, 0.1, 1]:
model = make_pipeline(StandardScaler(), SVC(kernel='rbf', C=10, gamma=gamma))
score = cross_val_score(model, X, y, cv=5).mean()
print(f'gamma={gamma}: CV accuracy={score:.4f}')Sigmoid Kernel
The sigmoid kernel is K(xᵢ, xⱼ) = tanh(γ · xᵢ·xⱼ + r), which resembles the activation function of a two-layer neural network. It is not always a valid (positive semi-definite) kernel for all parameter values, meaning the SVM optimisation may not converge to a global minimum. The sigmoid kernel is rarely the best choice in practice — RBF almost always outperforms it — but it can be useful when interpretability of the neural-network analogy is valued.
Choosing a Kernel in Practice
A practical guide for kernel selection: use linear when you have many features (text, genomics) or when the data is already high-dimensional — adding more dimensions via kernels is unnecessary; use RBF as the default for low-to-medium dimensional tabular data — it is the most flexible and often best; use polynomial when you have explicit reason to believe polynomial feature interactions matter; avoid sigmoid unless experimenting. Always compare kernels with cross-validation on your specific dataset.
Kernel SVM Complexity and Scalability
The main weakness of kernel SVMs is scalability. Training requires solving a quadratic programming problem that scales as O(n²) to O(n³) in the number of training examples. For 100,000 examples, an RBF SVM can take hours or run out of memory. Solutions: (1) use LinearSVC for linear kernels, which scales to millions of examples; (2) use approximate kernel methods like Nystroem or RBFSampler that create explicit low-dimensional feature maps; (3) switch to gradient boosting or neural networks for truly large datasets.
Comparing Kernels on the Same Dataset
The correct way to select a kernel is to compare them all with cross-validation on your dataset. Different datasets favour different kernels. A linearly separable problem gets no benefit from RBF. A problem with complex local structure may need high gamma RBF. Always start with the linear kernel as a baseline, then try RBF with a grid search over C and gamma. If neither outperforms the other significantly, choose linear for interpretability and speed.
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = load_digits(return_X_y=True)
for kernel in ['linear', 'poly', 'rbf']:
model = make_pipeline(StandardScaler(), SVC(kernel=kernel, C=10))
score = cross_val_score(model, X, y, cv=3).mean()
print(f'Kernel={kernel:8s}: CV accuracy={score:.4f}')Mercer's Theorem and Valid Kernels
Not every function can be used as a kernel. A valid kernel must satisfy Mercer's condition: it must be symmetric (K(x,y) = K(y,x)) and produce a positive semi-definite Gram matrix for any set of inputs. This guarantees that the kernel corresponds to a valid dot product in some feature space, making the SVM optimisation problem convex (one global minimum). Custom kernels for DNA sequences, graphs, or text can be defined and passed to SVC(kernel='precomputed') as long as they satisfy Mercer's theorem.
Quick Check
Test your understanding of the Kernel Trick from this lesson.
Lesson Recap
In this lesson you learned: kernel functions implicitly compute dot products in high-dimensional feature spaces, the RBF kernel is the most versatile default with gamma controlling the influence radius, and kernel SVMs do not scale to large datasets so consider linear kernels or approximate methods first. Next up we explore tuning C and gamma simultaneously with a grid search.
Frequently asked questions
Is the “The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels” lesson free?
Yes — the full text of “The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels” is free to read here on the web, and the Machine Learning Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Machine Learning Academy course, upgrade to CoddyKit PRO.
What will I learn in “The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels”?
Learners will apply RBF and polynomial kernels to a non-linearly separable dataset, and understand that kernels implicitly project data to higher dimensions. You practise Machine Learning Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Machine Learning Academy?
No prior experience is required. Machine Learning Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Machine Learning Academy lesson?
Yes. Every Machine Learning Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Maximum Margin Classifier: Support Vectors and Hyperplane
- Soft Margin SVM and the C Parameter
- The Kernel Trick: RBF, Polynomial, and Sigmoid Kernels
- Tuning C and Gamma with a Grid Search