0Pricing
Machine Learning Academy · Pelajaran

Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid

Peserta didik akan menerapkan kernel RBF dan polinomial pada dataset yang tidak dapat dipisahkan secara linear, serta memahami bahwa kernel secara implisit memproyeksikan data ke dimensi yang lebih tinggi.

Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid adalah pelajaran Machine Learning Academy gratis di CoddyKit. Ini adalah pelajaran 3 dari 4. Kamu bisa membaca pelajaran lengkapnya di bawah secara gratis — lalu praktikkan langsung di browser dengan editor kode bawaan dan tutor AI 24/7. Ini adalah bagian dari jalur belajar Machine Learning Academy, dan progresmu tersinkronisasi di web dan aplikasi CoddyKit. Kursus Machine Learning Academy mencakup 4 pelajaran total.

Bagian dari pelajaran ini belum diterjemahkan dan ditampilkan dalam bahasa Inggris.

The Problem: Non-Linear Data

Many real-world classification problems are not linearly separable — no straight line (or hyperplane) can correctly separate the classes. For example, data arranged in concentric rings cannot be separated by any linear boundary. One approach is to manually create new features (e.g., x², x×y) that make the classes linearly separable in the augmented space. The kernel trick does this automatically and implicitly, without ever computing the coordinates in the high-dimensional space.

Feature Maps: Lifting Data to Higher Dimensions

A feature map φ(x) transforms an input vector into a higher-dimensional representation. For example, φ([x₁, x₂]) = [x₁², √2·x₁x₂, x₂²] maps 2D data to 3D. After this mapping, classes that overlapped in 2D may become linearly separable in 3D. The SVM then finds a maximum-margin hyperplane in the transformed space. The corresponding decision boundary in the original 2D space is a curve, giving the SVM non-linear classification ability.

The Kernel Trick: Avoiding Explicit Feature Maps

Computing φ(x) explicitly is expensive or even impossible (some feature maps produce infinite-dimensional vectors). The key insight is that the SVM dual formulation only needs dot products φ(xᵢ)·φ(xⱼ), not the individual feature vectors. A kernel function K(xᵢ, xⱼ) computes this dot product directly from the original inputs without ever constructing φ(xᵢ). This is the kernel trick: expensive high-dimensional dot products computed cheaply in input space.

Polynomial Kernel

The polynomial kernel is defined as K(xᵢ, xⱼ) = (γ · xᵢ·xⱼ + r)^d, where d is the polynomial degree, γ is a scaling factor, and r is the coef0 parameter. A degree-2 polynomial kernel implicitly creates all pairwise interactions (x₁x₂) and squared terms (x₁²). Higher degrees create more complex boundaries but risk overfitting. In scikit-learn, use SVC(kernel='poly', degree=3).

from sklearn.svm import SVC
from sklearn.datasets import make_moons
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = make_moons(n_samples=300, noise=0.15, random_state=42)
for degree in [2, 3, 5]:
    model = make_pipeline(StandardScaler(), SVC(kernel='poly', degree=degree, C=5))
    score = cross_val_score(model, X, y, cv=5).mean()
    print(f'Polynomial degree={degree}: CV accuracy={score:.4f}')

RBF Kernel: The Default Workhorse

The Radial Basis Function (RBF) kernel, also called the Gaussian kernel, is defined as K(xᵢ, xⱼ) = exp(-γ · ||xᵢ - xⱼ||²). It measures similarity based on distance: nearby points have kernel value close to 1, distant points close to 0. The RBF kernel corresponds to an infinite-dimensional feature map, giving the SVM unlimited expressive power. It is the default kernel in scikit-learn's SVC and works well on most datasets with proper tuning of C and γ.

from sklearn.svm import SVC
from sklearn.datasets import make_moons
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = make_moons(n_samples=300, noise=0.15, random_state=42)
model = make_pipeline(StandardScaler(), SVC(kernel='rbf', C=1.0, gamma='scale'))
scores = cross_val_score(model, X, y, cv=5)
print('RBF SVM CV accuracy:', round(scores.mean(), 4))

The Gamma Parameter in RBF Kernel

The gamma parameter controls how far the influence of a single training example reaches. A small gamma makes each point's influence extend far — the decision boundary is smooth and the model underfits (high bias). A large gamma makes influence drop off steeply — the boundary wraps tightly around individual training points (high variance, overfitting). scikit-learn defaults: gamma='scale' (uses 1/(n_features × X.var())) or gamma='auto' (uses 1/n_features). Always tune C and gamma together.

from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = load_breast_cancer(return_X_y=True)
for gamma in [0.0001, 0.001, 0.01, 0.1, 1]:
    model = make_pipeline(StandardScaler(), SVC(kernel='rbf', C=10, gamma=gamma))
    score = cross_val_score(model, X, y, cv=5).mean()
    print(f'gamma={gamma}: CV accuracy={score:.4f}')

Sigmoid Kernel

The sigmoid kernel is K(xᵢ, xⱼ) = tanh(γ · xᵢ·xⱼ + r), which resembles the activation function of a two-layer neural network. It is not always a valid (positive semi-definite) kernel for all parameter values, meaning the SVM optimisation may not converge to a global minimum. The sigmoid kernel is rarely the best choice in practice — RBF almost always outperforms it — but it can be useful when interpretability of the neural-network analogy is valued.

Choosing a Kernel in Practice

A practical guide for kernel selection: use linear when you have many features (text, genomics) or when the data is already high-dimensional — adding more dimensions via kernels is unnecessary; use RBF as the default for low-to-medium dimensional tabular data — it is the most flexible and often best; use polynomial when you have explicit reason to believe polynomial feature interactions matter; avoid sigmoid unless experimenting. Always compare kernels with cross-validation on your specific dataset.

Kernel SVM Complexity and Scalability

The main weakness of kernel SVMs is scalability. Training requires solving a quadratic programming problem that scales as O(n²) to O(n³) in the number of training examples. For 100,000 examples, an RBF SVM can take hours or run out of memory. Solutions: (1) use LinearSVC for linear kernels, which scales to millions of examples; (2) use approximate kernel methods like Nystroem or RBFSampler that create explicit low-dimensional feature maps; (3) switch to gradient boosting or neural networks for truly large datasets.

Comparing Kernels on the Same Dataset

The correct way to select a kernel is to compare them all with cross-validation on your dataset. Different datasets favour different kernels. A linearly separable problem gets no benefit from RBF. A problem with complex local structure may need high gamma RBF. Always start with the linear kernel as a baseline, then try RBF with a grid search over C and gamma. If neither outperforms the other significantly, choose linear for interpretability and speed.

from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = load_digits(return_X_y=True)
for kernel in ['linear', 'poly', 'rbf']:
    model = make_pipeline(StandardScaler(), SVC(kernel=kernel, C=10))
    score = cross_val_score(model, X, y, cv=3).mean()
    print(f'Kernel={kernel:8s}: CV accuracy={score:.4f}')

Mercer's Theorem and Valid Kernels

Not every function can be used as a kernel. A valid kernel must satisfy Mercer's condition: it must be symmetric (K(x,y) = K(y,x)) and produce a positive semi-definite Gram matrix for any set of inputs. This guarantees that the kernel corresponds to a valid dot product in some feature space, making the SVM optimisation problem convex (one global minimum). Custom kernels for DNA sequences, graphs, or text can be defined and passed to SVC(kernel='precomputed') as long as they satisfy Mercer's theorem.

Quick Check

Test your understanding of the Kernel Trick from this lesson.

Lesson Recap

In this lesson you learned: kernel functions implicitly compute dot products in high-dimensional feature spaces, the RBF kernel is the most versatile default with gamma controlling the influence radius, and kernel SVMs do not scale to large datasets so consider linear kernels or approximate methods first. Next up we explore tuning C and gamma simultaneously with a grid search.

Pertanyaan yang Sering Diajukan

Apakah pelajaran “Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid” gratis?

Ya — teks lengkap “Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid” gratis dibaca di sini di web. Untuk praktiknya secara interaktif (editor kode bawaan dan tutor AI 24/7) dan buka sisa kursus Machine Learning Academy, upgrade ke CoddyKit PRO. Kursus Machine Learning Academy mencakup 4 pelajaran total.

Apa yang akan aku pelajari di “Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid”?

Peserta didik akan menerapkan kernel RBF dan polinomial pada dataset yang tidak dapat dipisahkan secara linear, serta memahami bahwa kernel secara implisit memproyeksikan data ke dimensi yang lebih t… Kamu berlatih Machine Learning Academy dengan kode praktik yang langsung kamu jalankan di browser, dan tutor AI 24/7 menjawab pertanyaanmu saat kamu mengerjakan pelajaran ini.

Apakah aku perlu pengalaman untuk memulai Machine Learning Academy?

Tidak diperlukan pengalaman sebelumnya. Machine Learning Academy di CoddyKit dirancang untuk pemula hingga pelajar tingkat lanjut, jadi kamu bisa memulai di sini atau dari awal dan belajar sesuai kecepatan kamu sendiri. Ini adalah pelajaran 3 dari 4.

Berapa lama pelajaran “Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid” memakan waktu?

Sebagian besar pelajaran CoddyKit memakan waktu sekitar 5–10 menit. Setiap pelajaran ringkas dan interaktif, jadi kamu membuat kemajuan stabil dan melanjutkan dari tempat kamu tinggalkan di web dan aplikasi.

Bisakah aku menulis dan menjalankan kode dalam pelajaran Machine Learning Academy ini?

Ya. Setiap pelajaran Machine Learning Academy menyertakan editor kode bawaan, jadi kamu menulis dan menjalankan kode nyata langsung di browser dan mendapatkan umpan balik AI instan — tidak diperlukan penyiapan lokal.

Semua pelajaran dalam kursus ini

  1. Pengklasifikasi Margin Maksimum: Vektor Pendukung dan Hiperbidang
  2. SVM Margin Lunak dan Parameter C
  3. Trik Kernel: Kernel RBF, Polinomial, dan Sigmoid
  4. Menyetel C dan Gamma dengan Pencarian Grid
← Kembali ke Machine Learning Academy