Por qué se reserva un conjunto de prueba
Estimar el rendimiento con datos no vistos.
Por qué se reserva un conjunto de prueba es una lección gratuita de Data Science Academy en CoddyKit. Esta es la lección 1 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Data Science Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Data Science Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
The Real Question
A model that memorizes your data looks brilliant on it. The real question is how it does on data it has never seen.
Hold Some Data Back
So you set aside part of your data and never train on it. This locked-away slice is your test set, kept for the very end.
Train Here, Judge There
The model learns only from the training set. You then judge it on the untouched test set to see how it truly generalizes. 🎯
Why Not Score on Training
Scoring on the same rows it learned from is like grading a test with the answer key open. That number flatters the model and overstates its skill.
Generalization Is the Goal
You do not care how well it fits old data. You care about generalization: making good predictions on tomorrow's fresh, unseen rows.
Meet Overfitting
When a model nails training data but flops on the test set, it is overfitting. It memorized noise instead of learning the real pattern.
A Common Split
A simple, popular choice is to train on about 80% of rows and test on the remaining 20%. More data to learn, enough left to judge fairly.
Touch It Only Once
The test set is sacred. If you keep peeking and tweaking until the score looks good, you have quietly leaked it into your decisions.
Where the Split Happens
In scikit-learn, one helper does the splitting for you. It shuffles and carves your data into train and test parts in a single call.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)The Honest Number
The score on the test set is your honest estimate of real-world performance. Trust it more than any glowing training score.
More Than a Formality
Holding out data is not red tape. It is the one habit that stops you from shipping a model that only ever worked on paper.
Quick Check
Why do you keep a separate test set?
Recap
You split data, learn on the train part, and judge on a sacred test set. That untouched slice is your honest read on real-world skill. 🎯
Preguntas frecuentes
¿La lección «Por qué se reserva un conjunto de prueba» es gratis?
Sí — el texto completo de «Por qué se reserva un conjunto de prueba» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Data Science Academy, actualiza a CoddyKit PRO. El curso de Data Science Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Por qué se reserva un conjunto de prueba»?
Estimar el rendimiento con datos no vistos. Practicas Data Science Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Data Science Academy?
No se requiere experiencia previa. Data Science Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 1 de 4.
¿Cuánto tiempo toma la lección «Por qué se reserva un conjunto de prueba»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Data Science Academy?
Sí. Cada lección de Data Science Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Por qué se reserva un conjunto de prueba
- Usar train_test_split correctamente
- Validación cruzada K-Fold
- Evitar la fuga de datos desde el principio