SVM à marge souple et paramètre C
Observez comment l’augmentation de C réduit la largeur de la marge et pénalise les mauvaises classifications, tandis qu’une petite valeur de C autorise davantage de violations pour obtenir une marge plus large et robuste.
SVM à marge souple et paramètre C est une leçon Machine Learning Academy gratuite sur CoddyKit. Ceci est la leçon 2 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Machine Learning Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Machine Learning Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
The Problem with Hard Margins
The hard-margin SVM requires that every training example be correctly classified with a margin of at least 1, leaving no room for error. In practice, real datasets are almost never perfectly linearly separable. Noise, mislabelled examples, and genuine class overlap mean that a hard margin is either impossible to satisfy or results in a boundary so contorted to avoid every violation that it overfits the training data. We need a principled way to allow some mistakes while still maximising the margin.
Slack Variables: Allowing Violations
The soft-margin SVM introduces slack variables ξᵢ ≥ 0 (xi, pronounced 'ksi'), one per training example, that measure how much a point violates the margin. If ξᵢ = 0, the point is correctly classified outside the margin. If 0 < ξᵢ < 1, the point is inside the margin but correctly classified. If ξᵢ > 1, the point is misclassified. The new objective minimises ||w||²/2 + C × Σξᵢ, balancing margin width against total violation.
The C Parameter: Penalty for Violations
C is the regularisation penalty: it controls how much the SVM penalises each margin violation. A large C imposes a heavy penalty, forcing the model to misclassify as few points as possible at the expense of a narrower margin — this leads to low bias, high variance (risk of overfitting). A small C accepts more violations in exchange for a wider, smoother margin — this leads to higher bias, lower variance (more robust to noise). Finding the right C requires cross-validation.
Visualising the Effect of C
With a very small C (e.g. 0.001), the SVM produces a wide margin with many misclassified training points — the boundary is smooth and generalises well but underfits if the data is cleanly separable. With a very large C (e.g. 1000), the boundary bends to correctly classify almost every training point, producing a narrow margin that may overfit. The optimal C sits between these extremes. This trade-off mirrors the bias-variance trade-off seen in all regularised models.
from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
import numpy as np
X, y = load_breast_cancer(return_X_y=True)
for C in [0.001, 0.01, 0.1, 1, 10, 100]:
model = make_pipeline(StandardScaler(), SVC(kernel='linear', C=C))
score = cross_val_score(model, X, y, cv=5).mean()
print(f'C={C:6}: CV accuracy={score:.4f}')Number of Support Vectors vs C
As C decreases (more regularisation, wider margin), more training examples violate the margin and become support vectors. As C increases (less regularisation, narrower margin), fewer examples lie on or inside the margin, so fewer support vectors are needed. You can observe this by inspecting svm.n_support_. A model with many support vectors relies on more training examples to define its boundary — it tends to be more robust but also more complex.
from sklearn.svm import SVC
from sklearn.datasets import load_breast_cancer
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
import numpy as np
X, y = load_breast_cancer(return_X_y=True)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
for C in [0.01, 0.1, 1, 10, 100]:
svm = SVC(kernel='linear', C=C).fit(X_scaled, y)
print(f'C={C:5}: support vectors = {svm.n_support_}, total = {sum(svm.n_support_)}')Hinge Loss: The SVM Loss Function
The soft-margin SVM minimises the hinge loss: max(0, 1 - y × (w·x + b)) for each training example, plus the L2 regularisation term ||w||²/(2C). Hinge loss is zero when a point is correctly classified beyond the margin (the point 'has no loss'). As a point moves toward or across the boundary, loss increases linearly. This makes SVMs less sensitive to outliers than squared-error loss, which would penalise far-wrong predictions quadratically.
import numpy as np
# Hinge loss for a single example: y in {-1, +1}, score = decision function value
def hinge_loss(y, score):
return max(0, 1 - y * score)
# Correctly classified, far beyond margin
print('Correct, margin=2:', hinge_loss(1, 3)) # 0
# Inside margin, still correct
print('Inside margin:', hinge_loss(1, 0.5)) # 0.5
# Misclassified
print('Misclassified:', hinge_loss(1, -1)) # 2LinearSVC for Large Datasets
scikit-learn provides LinearSVC as a faster alternative to SVC(kernel='linear') for large datasets. It uses the LIBLINEAR optimiser (primal or dual coordinate descent) instead of LIBSVM's quadratic programming solver. For datasets with tens of thousands of examples, LinearSVC can be 10-100x faster while producing nearly identical results. It does not support predict_proba() natively, but you can apply Platt scaling via CalibratedClassifierCV.
from sklearn.svm import LinearSVC
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = load_breast_cancer(return_X_y=True)
model = make_pipeline(StandardScaler(), LinearSVC(C=1.0, max_iter=5000))
scores = cross_val_score(model, X, y, cv=5)
print('LinearSVC CV:', scores.mean().round(4), '+/-', scores.std().round(4))Choosing C: Grid Search Strategy
The optimal C value spans many orders of magnitude, so always search on a logarithmic scale: [0.0001, 0.001, 0.01, 0.1, 1, 10, 100, 1000]. Linear spacing misses the important range of small values. Use GridSearchCV with 5-fold CV to evaluate each C. For SVM specifically, the search is one-dimensional (or two-dimensional with gamma for RBF kernel), making it computationally feasible even with a fine grid.
from sklearn.svm import SVC
from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import load_breast_cancer
import numpy as np
X, y = load_breast_cancer(return_X_y=True)
model = make_pipeline(StandardScaler(), SVC(kernel='linear'))
param_grid = {'svc__C': np.logspace(-3, 3, 7)}
grid = GridSearchCV(model, param_grid, cv=5)
grid.fit(X, y)
print('Best C:', grid.best_params_['svc__C'])
print('Best CV score:', round(grid.best_score_, 4))Soft Margin with Non-Linear Kernels
The soft-margin concept applies equally to non-linear kernels (RBF, polynomial). When you use an RBF kernel, C still controls the tolerance for margin violations, but now the boundary can be a curved surface in the original input space. A small C with an RBF kernel creates a very smooth, nearly circular boundary. A large C with a small gamma creates an extremely complex boundary that tightly wraps around every training cluster. Both extremes overfit in their own way.
Soft Margin Summary and Intuition
Think of the soft-margin SVM as a tradeoff dial. Turn C up: the model becomes aggressive, punishing every violation, squeezing the margin to fit training data. Turn C down: the model becomes lenient, accepting violations, widening the margin, prioritising generalisation. The right C is the one that strikes the balance your specific dataset needs. This is not unique to SVMs — regularisation appears in ridge regression (alpha), logistic regression (C), and neural networks (weight decay), always trading bias for variance.
Comparing SVM to Logistic Regression
Both logistic regression and the soft-margin SVM find linear decision boundaries, but they optimise different loss functions. Logistic regression minimises log-loss, which penalises all misclassifications continuously. SVM minimises hinge loss, which is zero for correctly classified points outside the margin and linear for violations. In practice: SVMs often outperform logistic regression on low-dimensional datasets with clear margins, while logistic regression is preferred when well-calibrated probability estimates are needed or when the dataset is large (millions of examples), where it is faster to train.
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X, y = load_breast_cancer(return_X_y=True)
for name, model in [('SVM', SVC(kernel='linear', C=1.0)), ('LogReg', LogisticRegression(max_iter=1000))]:
pipe = make_pipeline(StandardScaler(), model)
score = cross_val_score(pipe, X, y, cv=5).mean()
print(f'{name}: CV accuracy={score:.4f}')Quick Check
Test your understanding of the Soft Margin SVM and C parameter from this lesson.
Lesson Recap
In this lesson you learned: the soft-margin SVM allows controlled margin violations via slack variables, C is the regularisation penalty that balances margin width against training errors, and always search C on a logarithmic scale using cross-validation. Next up we explore the kernel trick that extends SVMs to non-linear boundaries.
Apprends Python avec un tuteur IA — gratuit
Écris et exécute du vrai code dans ton navigateur, obtiens de l'aide instantanée d'un tuteur IA disponible 24h/24, et reprends là où tu t'es arrêté sur le web ou dans l'app.
- Cours
- 30
- Leçons
- 120
Questions Fréquemment Posées
La leçon « SVM à marge souple et paramètre C » est-elle gratuite ?
Oui — le texte complet de « SVM à marge souple et paramètre C » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Machine Learning Academy, passe à CoddyKit PRO. Le cours Machine Learning Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « SVM à marge souple et paramètre C » ?
Observez comment l’augmentation de C réduit la largeur de la marge et pénalise les mauvaises classifications, tandis qu’une petite valeur de C autorise davantage de violations pour obtenir une marge… Tu pratiques Machine Learning Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer Machine Learning Academy ?
Aucune expérience préalable n'est requise. Machine Learning Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 2 sur 4.
Combien de temps prend la leçon « SVM à marge souple et paramètre C » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon Machine Learning Academy ?
Oui. Chaque leçon Machine Learning Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Classificateur à marge maximale : vecteurs de support et hyperplan
- SVM à marge souple et paramètre C
- L’astuce du noyau : noyaux RBF, polynomial et sigmoïde
- Régler C et gamma avec une recherche sur grille