0Pricing
Machine Learning Academy · Lección

Ensembles de votación: voto duro frente a voto blando

Combine KNN, regresión logística y un árbol de decisión en un VotingClassifier, y compare la mayoría del voto duro con el promedio de probabilidades del voto blando.

Ensembles de votación: voto duro frente a voto blando es una lección gratuita de Machine Learning Academy en CoddyKit. Esta es la lección 4 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Machine Learning Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Machine Learning Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

What Is a Voting Ensemble?

A Voting Ensemble combines the predictions of several different model types — such as a logistic regression, a k-nearest neighbours classifier, and a decision tree — to produce a single final prediction. Unlike bagging which uses many copies of the same model type, a voting ensemble leverages diversity in model architecture. Because different models make different kinds of errors, their combination can outperform any individual member, especially on datasets where no single algorithm dominates.

Hard Voting: Majority Rules

In hard voting, each model votes for a class label and the class that receives the most votes wins. If three models predict [cat, cat, dog], the ensemble predicts cat. Hard voting is simple and interpretable, but it treats all models as equally reliable and ignores confidence levels. A model that is barely 51% confident votes the same as one that is 99% confident, which can lead to suboptimal decisions when model confidences differ greatly.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score

X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier()),
        ('dt', DecisionTreeClassifier())
    ],
    voting='hard'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Hard Voting CV:', scores.mean().round(4))

Soft Voting: Probability Averaging

In soft voting, each model outputs class probabilities rather than hard labels. The ensemble averages the probabilities across models and picks the class with the highest average probability. For example, if three models assign probabilities [0.9, 0.1], [0.7, 0.3], and [0.6, 0.4] to two classes, the average is [0.73, 0.27] and class 0 wins. Soft voting typically outperforms hard voting because it uses richer information about each model's confidence.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score

X, y = load_iris(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier()),
        ('svc', SVC(probability=True))  # probability=True required for soft vote
    ],
    voting='soft'
)
scores = cross_val_score(voter, X, y, cv=5)
print('Soft Voting CV:', scores.mean().round(4))

Requirement: All Models Must Support predict_proba

Soft voting requires every model in the ensemble to provide probability estimates via predict_proba(). Most scikit-learn classifiers support this natively (logistic regression, random forest, KNN, naive Bayes). However, SVC does not output probabilities by default — you must set probability=True in the constructor, which adds Platt scaling (a calibration step) and increases training time. If any model cannot produce probabilities, you must fall back to hard voting.

Weighting Individual Models

Both hard and soft voting support the weights parameter, which lets you give stronger influence to more accurate models. For example, if logistic regression has 92% accuracy and KNN has 85%, you might weight them as [2, 1]. In hard voting, each vote is replicated according to its weight. In soft voting, each model's probability vector is multiplied by its weight before averaging. Choosing weights based on cross-validation accuracy is a simple and effective approach.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score

X, y = load_breast_cancer(return_X_y=True)
voter = VotingClassifier(
    estimators=[
        ('lr', LogisticRegression(max_iter=1000)),
        ('knn', KNeighborsClassifier(n_neighbors=5)),
        ('dt', DecisionTreeClassifier(max_depth=5))
    ],
    voting='soft',
    weights=[2, 1, 1]  # Trust LR twice as much
)
print('Weighted soft vote CV:', cross_val_score(voter, X, y, cv=5).mean().round(4))

Comparing Hard vs Soft vs Individual Models

Running a direct comparison across individual models, hard voting, and soft voting on the same dataset reveals the ensemble effect clearly. In most cases, soft voting exceeds hard voting, and both exceed the weakest individual model. However, the ensemble may not always beat the strongest single model — it depends on how much the models' errors are correlated. If all models fail on the same examples, combining them offers no benefit.

from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score

X, y = load_breast_cancer(return_X_y=True)
estimators = [('lr', LogisticRegression(max_iter=1000)), ('knn', KNeighborsClassifier()), ('dt', DecisionTreeClassifier())]
for name, est in estimators:
    print(f'{name}: {cross_val_score(est, X, y, cv=5).mean():.4f}')
for v in ['hard', 'soft']:
    vc = VotingClassifier(estimators=estimators, voting=v)
    print(f'Voting ({v}): {cross_val_score(vc, X, y, cv=5).mean():.4f}')

Choosing Models for Diversity

The key to a powerful voting ensemble is model diversity. Combining three logistic regressions with different seeds adds almost no value — they will all make the same mistakes. The most effective ensembles pair models with different inductive biases: a linear model (logistic regression), an instance-based model (KNN), a tree-based model (random forest), and optionally a kernel-based model (SVM). Each sees the data through a different lens and makes distinct errors that cancel out under averaging.

VotingRegressor for Continuous Targets

VotingRegressor applies the same idea to regression: average the numeric predictions from multiple regressors. There is no concept of hard vs soft voting for regression — the output is always a weighted or unweighted average of the individual predictions. Combining a Ridge regression (linear), a Random Forest (tree ensemble), and an SVR (kernel method) often beats any single model on diverse tabular datasets.

from sklearn.ensemble import VotingRegressor, RandomForestRegressor
from sklearn.linear_model import Ridge
from sklearn.svm import SVR
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = fetch_california_housing(return_X_y=True)
vr = VotingRegressor(estimators=[
    ('ridge', make_pipeline(StandardScaler(), Ridge())),
    ('rf', RandomForestRegressor(n_estimators=50, random_state=42)),
    ('svr', make_pipeline(StandardScaler(), SVR()))
])
rmse = np.sqrt(-cross_val_score(vr, X, y, scoring='neg_mean_squared_error', cv=3).mean())
print('VotingRegressor RMSE:', round(rmse, 4))

Voting Ensembles in Production

Voting ensembles are easy to deploy because all member models can be serialised alongside the VotingClassifier wrapper using joblib.dump(). Inference requires running every member model, so prediction latency scales linearly with the number of members. For latency-sensitive applications, limit the ensemble to 2-3 strong, fast models rather than many slow ones. Always profile inference time as part of your deployment evaluation.

When Voting Ensembles Fail to Help

Voting ensembles disappoint when: (1) all member models share the same blind spots and make correlated errors; (2) one model is vastly superior to the others and the weaker models drag down the ensemble; (3) the task has very low noise so a single well-tuned model already achieves near-perfect performance; or (4) the dataset is too small and extra model capacity just overfits. Diagnose by comparing individual model error patterns on a validation set — if the errors are highly correlated, try replacing some models with more diverse architectures.

Stacking vs Voting

Voting uses fixed, hand-chosen aggregation (majority vote or average). Stacking (or stacked generalisation) goes further: a meta-learner is trained to optimally combine the outputs of the base models. The base models' out-of-fold predictions become the input features for the meta-learner. This adds flexibility — the meta-learner can learn that model A is reliable on easy examples while model B is better on hard ones. Stacking typically outperforms voting but requires more careful implementation to avoid data leakage.

Quick Check

Test your understanding of Voting Ensemble concepts from this lesson.

Lesson Recap

In this lesson you learned: Voting Ensembles combine diverse model types to reduce correlated errors, hard voting uses majority labels while soft voting averages probability estimates, and model diversity is the key ingredient for an effective voting ensemble. Next up we explore Support Vector Machines and the maximum-margin classifier.

Preguntas frecuentes

¿La lección «Ensembles de votación: voto duro frente a voto blando» es gratis?

Sí — el texto completo de «Ensembles de votación: voto duro frente a voto blando» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Machine Learning Academy, actualiza a CoddyKit PRO. El curso de Machine Learning Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Ensembles de votación: voto duro frente a voto blando»?

Combine KNN, regresión logística y un árbol de decisión en un VotingClassifier, y compare la mayoría del voto duro con el promedio de probabilidades del voto blando. Practicas Machine Learning Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Machine Learning Academy?

No se requiere experiencia previa. Machine Learning Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 4 de 4.

¿Cuánto tiempo toma la lección «Ensembles de votación: voto duro frente a voto blando»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Machine Learning Academy?

Sí. Cada lección de Machine Learning Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Explicación del agregamiento bootstrap (Bagging)
  2. Selección aleatoria de características: el truco de Random Forest
  3. Error out-of-bag: validación gratuita dentro del bosque
  4. Ensembles de votación: voto duro frente a voto blando
← Volver a Machine Learning Academy