0Pricing
Machine Learning Academy · Lektion

SVM mit Soft Margin und der Parameter C

Beobachten Sie, wie ein höheres C die Randbreite verringert und Fehlklassifikationen stärker bestraft, während ein kleines C mehr Verletzungen für einen breiteren und robusteren Rand zulässt.

SVM mit Soft Margin und der Parameter C ist eine kostenlose Machine Learning Academy-Lektion auf CoddyKit. Dies ist Lektion 2 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des Machine Learning Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der Machine Learning Academy-Kurs umfasst insgesamt 4 Lektionen.

Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.

The Problem with Hard Margins

The hard-margin SVM requires that every training example be correctly classified with a margin of at least 1, leaving no room for error. In practice, real datasets are almost never perfectly linearly separable. Noise, mislabelled examples, and genuine class overlap mean that a hard margin is either impossible to satisfy or results in a boundary so contorted to avoid every violation that it overfits the training data. We need a principled way to allow some mistakes while still maximising the margin.

Slack Variables: Allowing Violations

The soft-margin SVM introduces slack variables ξᵢ ≥ 0 (xi, pronounced 'ksi'), one per training example, that measure how much a point violates the margin. If ξᵢ = 0, the point is correctly classified outside the margin. If 0 < ξᵢ < 1, the point is inside the margin but correctly classified. If ξᵢ > 1, the point is misclassified. The new objective minimises ||w||²/2 + C × Σξᵢ, balancing margin width against total violation.

The C Parameter: Penalty for Violations

C is the regularisation penalty: it controls how much the SVM penalises each margin violation. A large C imposes a heavy penalty, forcing the model to misclassify as few points as possible at the expense of a narrower margin — this leads to low bias, high variance (risk of overfitting). A small C accepts more violations in exchange for a wider, smoother margin — this leads to higher bias, lower variance (more robust to noise). Finding the right C requires cross-validation.

Visualising the Effect of C

With a very small C (e.g. 0.001), the SVM produces a wide margin with many misclassified training points — the boundary is smooth and generalises well but underfits if the data is cleanly separable. With a very large C (e.g. 1000), the boundary bends to correctly classify almost every training point, producing a narrow margin that may overfit. The optimal C sits between these extremes. This trade-off mirrors the bias-variance trade-off seen in all regularised models.

from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
for C in [0.001, 0.01, 0.1, 1, 10, 100]:
    model = make_pipeline(StandardScaler(), SVC(kernel='linear', C=C))
    score = cross_val_score(model, X, y, cv=5).mean()
    print(f'C={C:6}: CV accuracy={score:.4f}')

Number of Support Vectors vs C

As C decreases (more regularisation, wider margin), more training examples violate the margin and become support vectors. As C increases (less regularisation, narrower margin), fewer examples lie on or inside the margin, so fewer support vectors are needed. You can observe this by inspecting svm.n_support_. A model with many support vectors relies on more training examples to define its boundary — it tends to be more robust but also more complex.

from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
for C in [0.01, 0.1, 1, 10, 100]:
    svm = SVC(kernel='linear', C=C).fit(X_scaled, y)
    print(f'C={C:5}: support vectors = {svm.n_support_}, total = {sum(svm.n_support_)}')

Hinge Loss: The SVM Loss Function

The soft-margin SVM minimises the hinge loss: max(0, 1 - y × (w·x + b)) for each training example, plus the L2 regularisation term ||w||²/(2C). Hinge loss is zero when a point is correctly classified beyond the margin (the point 'has no loss'). As a point moves toward or across the boundary, loss increases linearly. This makes SVMs less sensitive to outliers than squared-error loss, which would penalise far-wrong predictions quadratically.

import numpy as np
# Hinge loss for a single example: y in {-1, +1}, score = decision function value
def hinge_loss(y, score):
    return max(0, 1 - y * score)

# Correctly classified, far beyond margin
print('Correct, margin=2:', hinge_loss(1, 3))    # 0
# Inside margin, still correct
print('Inside margin:', hinge_loss(1, 0.5))       # 0.5
# Misclassified
print('Misclassified:', hinge_loss(1, -1))        # 2

LinearSVC for Large Datasets

scikit-learn provides LinearSVC as a faster alternative to SVC(kernel='linear') for large datasets. It uses the LIBLINEAR optimiser (primal or dual coordinate descent) instead of LIBSVM's quadratic programming solver. For datasets with tens of thousands of examples, LinearSVC can be 10-100x faster while producing nearly identical results. It does not support predict_proba() natively, but you can apply Platt scaling via CalibratedClassifierCV.

from sklearn.svm import LinearSVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = load_breast_cancer(return_X_y=True)
model = make_pipeline(StandardScaler(), LinearSVC(C=1.0, max_iter=5000))
scores = cross_val_score(model, X, y, cv=5)
print('LinearSVC CV:', scores.mean().round(4), '+/-', scores.std().round(4))

Choosing C: Grid Search Strategy

The optimal C value spans many orders of magnitude, so always search on a logarithmic scale: [0.0001, 0.001, 0.01, 0.1, 1, 10, 100, 1000]. Linear spacing misses the important range of small values. Use GridSearchCV with 5-fold CV to evaluate each C. For SVM specifically, the search is one-dimensional (or two-dimensional with gamma for RBF kernel), making it computationally feasible even with a fine grid.

from sklearn.svm import SVC
from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import load_breast_cancer
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
model = make_pipeline(StandardScaler(), SVC(kernel='linear'))
param_grid = {'svc__C': np.logspace(-3, 3, 7)}
grid = GridSearchCV(model, param_grid, cv=5)
grid.fit(X, y)
print('Best C:', grid.best_params_['svc__C'])
print('Best CV score:', round(grid.best_score_, 4))

Soft Margin with Non-Linear Kernels

The soft-margin concept applies equally to non-linear kernels (RBF, polynomial). When you use an RBF kernel, C still controls the tolerance for margin violations, but now the boundary can be a curved surface in the original input space. A small C with an RBF kernel creates a very smooth, nearly circular boundary. A large C with a small gamma creates an extremely complex boundary that tightly wraps around every training cluster. Both extremes overfit in their own way.

Soft Margin Summary and Intuition

Think of the soft-margin SVM as a tradeoff dial. Turn C up: the model becomes aggressive, punishing every violation, squeezing the margin to fit training data. Turn C down: the model becomes lenient, accepting violations, widening the margin, prioritising generalisation. The right C is the one that strikes the balance your specific dataset needs. This is not unique to SVMs — regularisation appears in ridge regression (alpha), logistic regression (C), and neural networks (weight decay), always trading bias for variance.

Comparing SVM to Logistic Regression

Both logistic regression and the soft-margin SVM find linear decision boundaries, but they optimise different loss functions. Logistic regression minimises log-loss, which penalises all misclassifications continuously. SVM minimises hinge loss, which is zero for correctly classified points outside the margin and linear for violations. In practice: SVMs often outperform logistic regression on low-dimensional datasets with clear margins, while logistic regression is preferred when well-calibrated probability estimates are needed or when the dataset is large (millions of examples), where it is faster to train.

from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X, y = load_breast_cancer(return_X_y=True)
for name, model in [('SVM', SVC(kernel='linear', C=1.0)), ('LogReg', LogisticRegression(max_iter=1000))]:
    pipe = make_pipeline(StandardScaler(), model)
    score = cross_val_score(pipe, X, y, cv=5).mean()
    print(f'{name}: CV accuracy={score:.4f}')

Quick Check

Test your understanding of the Soft Margin SVM and C parameter from this lesson.

Lesson Recap

In this lesson you learned: the soft-margin SVM allows controlled margin violations via slack variables, C is the regularisation penalty that balances margin width against training errors, and always search C on a logarithmic scale using cross-validation. Next up we explore the kernel trick that extends SVMs to non-linear boundaries.

Häufig gestellte Fragen

Ist die Lektion „SVM mit Soft Margin und der Parameter C“ kostenlos?

Ja — der vollständige Text von „SVM mit Soft Margin und der Parameter C“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des Machine Learning Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der Machine Learning Academy-Kurs umfasst insgesamt 4 Lektionen.

Was lerne ich in „SVM mit Soft Margin und der Parameter C“?

Beobachten Sie, wie ein höheres C die Randbreite verringert und Fehlklassifikationen stärker bestraft, während ein kleines C mehr Verletzungen für einen breiteren und robusteren Rand zulässt. Du übst Machine Learning Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.

Brauche ich Erfahrung, um Machine Learning Academy zu starten?

Keine Vorkenntnisse erforderlich. Machine Learning Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 2 von 4.

Wie lange dauert die Lektion „SVM mit Soft Margin und der Parameter C“?

Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.

Kann ich in dieser Machine Learning Academy-Lektion Code schreiben und ausführen?

Ja. Jede Machine Learning Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.

Alle Lektionen in diesem Kurs

  1. Klassifikator mit maximalem Rand: Support-Vektoren und Hyperplane
  2. SVM mit Soft Margin und der Parameter C
  3. Der Kernel-Trick: RBF-, polynomiale und Sigmoid-Kernel
  4. C und Gamma mit einer Gittersuche optimieren
← Zurück zu Machine Learning Academy