Ensembles par vote : vote dur ou vote souple
Combinez un KNN, une régression logistique et un arbre de décision dans un VotingClassifier, puis comparez la majorité du vote dur à la moyenne des probabilités du vote souple.
Ensembles par vote : vote dur ou vote souple est une leçon Machine Learning Academy gratuite sur CoddyKit. Ceci est la leçon 4 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Machine Learning Academy, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Machine Learning Academy comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
What Is a Voting Ensemble?
A Voting Ensemble combines the predictions of several different model types — such as a logistic regression, a k-nearest neighbours classifier, and a decision tree — to produce a single final prediction. Unlike bagging which uses many copies of the same model type, a voting ensemble leverages diversity in model architecture. Because different models make different kinds of errors, their combination can outperform any individual member, especially on datasets where no single algorithm dominates.
Hard Voting: Majority Rules
In hard voting, each model votes for a class label and the class that receives the most votes wins. If three models predict [cat, cat, dog], the ensemble predicts cat. Hard voting is simple and interpretable, but it treats all models as equally reliable and ignores confidence levels. A model that is barely 51% confident votes the same as one that is 99% confident, which can lead to suboptimal decisions when model confidences differ greatly.
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score
X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
estimators=[
('lr', LogisticRegression(max_iter=1000)),
('knn', KNeighborsClassifier()),
('dt', DecisionTreeClassifier())
],
voting='hard'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Hard Voting CV:', scores.mean().round(4))Soft Voting: Probability Averaging
In soft voting, each model outputs class probabilities rather than hard labels. The ensemble averages the probabilities across models and picks the class with the highest average probability. For example, if three models assign probabilities [0.9, 0.1], [0.7, 0.3], and [0.6, 0.4] to two classes, the average is [0.73, 0.27] and class 0 wins. Soft voting typically outperforms hard voting because it uses richer information about each model's confidence.
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score
X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
estimators=[
('lr', LogisticRegression(max_iter=1000)),
('knn', KNeighborsClassifier()),
('svc', SVC(probability=True)) # probability=True required for soft vote
],
voting='soft'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Soft Voting CV:', scores.mean().round(4))Requirement: All Models Must Support predict_proba
Soft voting requires every model in the ensemble to provide probability estimates via predict_proba(). Most scikit-learn classifiers support this natively (logistic regression, random forest, KNN, naive Bayes). However, SVC does not output probabilities by default — you must set probability=True in the constructor, which adds Platt scaling (a calibration step) and increases training time. If any model cannot produce probabilities, you must fall back to hard voting.
Weighting Individual Models
Both hard and soft voting support the weights parameter, which lets you give stronger influence to more accurate models. For example, if logistic regression has 92% accuracy and KNN has 85%, you might weight them as [2, 1]. In hard voting, each vote is replicated according to its weight. In soft voting, each model's probability vector is multiplied by its weight before averaging. Choosing weights based on cross-validation accuracy is a simple and effective approach.
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
X, y = load_breast_cancer(return_X_y=True)
voter = VotingClassifier(
estimators=[
('lr', LogisticRegression(max_iter=1000)),
('knn', KNeighborsClassifier(n_neighbors=5)),
('dt', DecisionTreeClassifier(max_depth=5))
],
voting='soft',
weights=[2, 1, 1] # Trust LR twice as much
)
print('Weighted soft vote CV:', cross_val_score(voter, X, y, cv=5).mean().round(4))Comparing Hard vs Soft vs Individual Models
Running a direct comparison across individual models, hard voting, and soft voting on the same dataset reveals the ensemble effect clearly. In most cases, soft voting exceeds hard voting, and both exceed the weakest individual model. However, the ensemble may not always beat the strongest single model — it depends on how much the models' errors are correlated. If all models fail on the same examples, combining them offers no benefit.
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
X, y = load_breast_cancer(return_X_y=True)
estimators = [('lr', LogisticRegression(max_iter=1000)), ('knn', KNeighborsClassifier()), ('dt', DecisionTreeClassifier())]
for name, est in estimators:
print(f'{name}: {cross_val_score(est, X, y, cv=5).mean():.4f}')
for v in ['hard', 'soft']:
vc = VotingClassifier(estimators=estimators, voting=v)
print(f'Voting ({v}): {cross_val_score(vc, X, y, cv=5).mean():.4f}')Choosing Models for Diversity
The key to a powerful voting ensemble is model diversity. Combining three logistic regressions with different seeds adds almost no value — they will all make the same mistakes. The most effective ensembles pair models with different inductive biases: a linear model (logistic regression), an instance-based model (KNN), a tree-based model (random forest), and optionally a kernel-based model (SVM). Each sees the data through a different lens and makes distinct errors that cancel out under averaging.
VotingRegressor for Continuous Targets
VotingRegressor applies the same idea to regression: average the numeric predictions from multiple regressors. There is no concept of hard vs soft voting for regression — the output is always a weighted or unweighted average of the individual predictions. Combining a Ridge regression (linear), a Random Forest (tree ensemble), and an SVR (kernel method) often beats any single model on diverse tabular datasets.
from sklearn.ensemble import VotingRegressor, RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.svm import SVR
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import cross_val_score
import numpy as np
X, y = fetch_california_housing(return_X_y=True)
vr = VotingRegressor(estimators=[
('ridge', make_pipeline(StandardScaler(), Ridge())),
('rf', RandomForestRegressor(n_estimators=50, random_state=42)),
('svr', make_pipeline(StandardScaler(), SVR()))
])
rmse = np.sqrt(-cross_val_score(vr, X, y, scoring='neg_mean_squared_error', cv=3).mean())
print('VotingRegressor RMSE:', round(rmse, 4))Voting Ensembles in Production
Voting ensembles are easy to deploy because all member models can be serialised alongside the VotingClassifier wrapper using joblib.dump(). Inference requires running every member model, so prediction latency scales linearly with the number of members. For latency-sensitive applications, limit the ensemble to 2-3 strong, fast models rather than many slow ones. Always profile inference time as part of your deployment evaluation.
When Voting Ensembles Fail to Help
Voting ensembles disappoint when: (1) all member models share the same blind spots and make correlated errors; (2) one model is vastly superior to the others and the weaker models drag down the ensemble; (3) the task has very low noise so a single well-tuned model already achieves near-perfect performance; or (4) the dataset is too small and extra model capacity just overfits. Diagnose by comparing individual model error patterns on a validation set — if the errors are highly correlated, try replacing some models with more diverse architectures.
Stacking vs Voting
Voting uses fixed, hand-chosen aggregation (majority vote or average). Stacking (or stacked generalisation) goes further: a meta-learner is trained to optimally combine the outputs of the base models. The base models' out-of-fold predictions become the input features for the meta-learner. This adds flexibility — the meta-learner can learn that model A is reliable on easy examples while model B is better on hard ones. Stacking typically outperforms voting but requires more careful implementation to avoid data leakage.
Quick Check
Test your understanding of Voting Ensemble concepts from this lesson.
Lesson Recap
In this lesson you learned: Voting Ensembles combine diverse model types to reduce correlated errors, hard voting uses majority labels while soft voting averages probability estimates, and model diversity is the key ingredient for an effective voting ensemble. Next up we explore Support Vector Machines and the maximum-margin classifier.
Questions Fréquemment Posées
La leçon « Ensembles par vote : vote dur ou vote souple » est-elle gratuite ?
Oui — le texte complet de « Ensembles par vote : vote dur ou vote souple » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Machine Learning Academy, passe à CoddyKit PRO. Le cours Machine Learning Academy comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Ensembles par vote : vote dur ou vote souple » ?
Combinez un KNN, une régression logistique et un arbre de décision dans un VotingClassifier, puis comparez la majorité du vote dur à la moyenne des probabilités du vote souple. Tu pratiques Machine Learning Academy avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer Machine Learning Academy ?
Aucune expérience préalable n'est requise. Machine Learning Academy sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 4 sur 4.
Combien de temps prend la leçon « Ensembles par vote : vote dur ou vote souple » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon Machine Learning Academy ?
Oui. Chaque leçon Machine Learning Academy inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Agrégation par bootstrap (bagging) expliquée
- Sélection aléatoire des caractéristiques : l’astuce des forêts aléatoires
- Erreur hors sac : validation gratuite dans la forêt
- Ensembles par vote : vote dur ou vote souple