0Pricing
Machine Learning Academy · Aula

Conjuntos por votação: voto rígido vs. voto suave

Combine um KNN, uma regressão logística e uma árvore de decisão em um VotingClassifier e compare a maioria por voto rígido com a média de probabilidades por voto suave.

Conjuntos por votação: voto rígido vs. voto suave é uma aula grátis de Machine Learning Academy no CoddyKit. Esta é a aula 4 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Machine Learning Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Machine Learning Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

What Is a Voting Ensemble?

A Voting Ensemble combines the predictions of several different model types — such as a logistic regression, a k-nearest neighbours classifier, and a decision tree — to produce a single final prediction. Unlike bagging which uses many copies of the same model type, a voting ensemble leverages diversity in model architecture. Because different models make different kinds of errors, their combination can outperform any individual member, especially on datasets where no single algorithm dominates.

Hard Voting: Majority Rules

In hard voting, each model votes for a class label and the class that receives the most votes wins. If three models predict [cat, cat, dog], the ensemble predicts cat. Hard voting is simple and interpretable, but it treats all models as equally reliable and ignores confidence levels. A model that is barely 51% confident votes the same as one that is 99% confident, which can lead to suboptimal decisions when model confidences differ greatly.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score

X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier()),
        ('dt', DecisionTreeClassifier())
    ],
    voting='hard'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Hard Voting CV:', scores.mean().round(4))

Soft Voting: Probability Averaging

In soft voting, each model outputs class probabilities rather than hard labels. The ensemble averages the probabilities across models and picks the class with the highest average probability. For example, if three models assign probabilities [0.9, 0.1], [0.7, 0.3], and [0.6, 0.4] to two classes, the average is [0.73, 0.27] and class 0 wins. Soft voting typically outperforms hard voting because it uses richer information about each model's confidence.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score

X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier()),
        ('svc', SVC(probability=True))  # probability=True required for soft vote
    ],
    voting='soft'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Soft Voting CV:', scores.mean().round(4))

Requirement: All Models Must Support predict_proba

Soft voting requires every model in the ensemble to provide probability estimates via predict_proba(). Most scikit-learn classifiers support this natively (logistic regression, random forest, KNN, naive Bayes). However, SVC does not output probabilities by default — you must set probability=True in the constructor, which adds Platt scaling (a calibration step) and increases training time. If any model cannot produce probabilities, you must fall back to hard voting.

Weighting Individual Models

Both hard and soft voting support the weights parameter, which lets you give stronger influence to more accurate models. For example, if logistic regression has 92% accuracy and KNN has 85%, you might weight them as [2, 1]. In hard voting, each vote is replicated according to its weight. In soft voting, each model's probability vector is multiplied by its weight before averaging. Choosing weights based on cross-validation accuracy is a simple and effective approach.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score

X, y = load_breast_cancer(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier(n_neighbors=5)),
        ('dt', DecisionTreeClassifier(max_depth=5))
    ],
    voting='soft',
    weights=[2, 1, 1]  # Trust LR twice as much
)
print('Weighted soft vote CV:', cross_val_score(voter, X, y, cv=5).mean().round(4))

Comparing Hard vs Soft vs Individual Models

Running a direct comparison across individual models, hard voting, and soft voting on the same dataset reveals the ensemble effect clearly. In most cases, soft voting exceeds hard voting, and both exceed the weakest individual model. However, the ensemble may not always beat the strongest single model — it depends on how much the models' errors are correlated. If all models fail on the same examples, combining them offers no benefit.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score

X, y = load_breast_cancer(return_X_y=True)
estimators = [('lr', LogisticRegression(max_iter=1000)), ('knn', KNeighborsClassifier()), ('dt', DecisionTreeClassifier())]
for name, est in estimators:
    print(f'{name}: {cross_val_score(est, X, y, cv=5).mean():.4f}')
for v in ['hard', 'soft']:
    vc = VotingClassifier(estimators=estimators, voting=v)
    print(f'Voting ({v}): {cross_val_score(vc, X, y, cv=5).mean():.4f}')

Choosing Models for Diversity

The key to a powerful voting ensemble is model diversity. Combining three logistic regressions with different seeds adds almost no value — they will all make the same mistakes. The most effective ensembles pair models with different inductive biases: a linear model (logistic regression), an instance-based model (KNN), a tree-based model (random forest), and optionally a kernel-based model (SVM). Each sees the data through a different lens and makes distinct errors that cancel out under averaging.

VotingRegressor for Continuous Targets

VotingRegressor applies the same idea to regression: average the numeric predictions from multiple regressors. There is no concept of hard vs soft voting for regression — the output is always a weighted or unweighted average of the individual predictions. Combining a Ridge regression (linear), a Random Forest (tree ensemble), and an SVR (kernel method) often beats any single model on diverse tabular datasets.

from sklearn.ensemble import VotingRegressor, RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.svm import SVR
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = fetch_california_housing(return_X_y=True)
vr = VotingRegressor(estimators=[
    ('ridge', make_pipeline(StandardScaler(), Ridge())),
    ('rf', RandomForestRegressor(n_estimators=50, random_state=42)),
    ('svr', make_pipeline(StandardScaler(), SVR()))
])
rmse = np.sqrt(-cross_val_score(vr, X, y, scoring='neg_mean_squared_error', cv=3).mean())
print('VotingRegressor RMSE:', round(rmse, 4))

Voting Ensembles in Production

Voting ensembles are easy to deploy because all member models can be serialised alongside the VotingClassifier wrapper using joblib.dump(). Inference requires running every member model, so prediction latency scales linearly with the number of members. For latency-sensitive applications, limit the ensemble to 2-3 strong, fast models rather than many slow ones. Always profile inference time as part of your deployment evaluation.

When Voting Ensembles Fail to Help

Voting ensembles disappoint when: (1) all member models share the same blind spots and make correlated errors; (2) one model is vastly superior to the others and the weaker models drag down the ensemble; (3) the task has very low noise so a single well-tuned model already achieves near-perfect performance; or (4) the dataset is too small and extra model capacity just overfits. Diagnose by comparing individual model error patterns on a validation set — if the errors are highly correlated, try replacing some models with more diverse architectures.

Stacking vs Voting

Voting uses fixed, hand-chosen aggregation (majority vote or average). Stacking (or stacked generalisation) goes further: a meta-learner is trained to optimally combine the outputs of the base models. The base models' out-of-fold predictions become the input features for the meta-learner. This adds flexibility — the meta-learner can learn that model A is reliable on easy examples while model B is better on hard ones. Stacking typically outperforms voting but requires more careful implementation to avoid data leakage.

Quick Check

Test your understanding of Voting Ensemble concepts from this lesson.

Lesson Recap

In this lesson you learned: Voting Ensembles combine diverse model types to reduce correlated errors, hard voting uses majority labels while soft voting averages probability estimates, and model diversity is the key ingredient for an effective voting ensemble. Next up we explore Support Vector Machines and the maximum-margin classifier.

Perguntas Frequentes

A aula “Conjuntos por votação: voto rígido vs. voto suave” é grátis?

Sim — o texto completo de “Conjuntos por votação: voto rígido vs. voto suave” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Machine Learning Academy, atualize para CoddyKit PRO. O curso de Machine Learning Academy inclui 4 aulas no total.

O que vou aprender em “Conjuntos por votação: voto rígido vs. voto suave”?

Combine um KNN, uma regressão logística e uma árvore de decisão em um VotingClassifier e compare a maioria por voto rígido com a média de probabilidades por voto suave. Você pratica Machine Learning Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Machine Learning Academy?

Nenhuma experiência prévia é necessária. Machine Learning Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 4 de 4.

Quanto tempo leva a aula “Conjuntos por votação: voto rígido vs. voto suave”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Machine Learning Academy?

Sim. Cada aula de Machine Learning Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Agregação por bootstrap (Bagging) explicada
  2. Seleção aleatória de atributos: o truque da floresta aleatória
  3. Erro fora da amostra: validação gratuita dentro da floresta
  4. Conjuntos por votação: voto rígido vs. voto suave
← Voltar para Machine Learning Academy