0Pricing
NLP Academy · Lección

Entrenamiento con características TF-IDF

Entrenar un clasificador en scikit-learn

Entrenamiento con características TF-IDF es una lección gratuita de NLP Academy en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de NLP Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de NLP Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

From Text to Numbers First

A model cannot read raw sentences. You first convert documents into TF-IDF vectors, then train logistic regression on those numbers. 🔢

Fit the Vectorizer

The TfidfVectorizer learns your vocabulary and weighting from the training texts. One call turns a list of strings into a feature matrix.

from sklearn.feature_extraction.text import TfidfVectorizer
vec = TfidfVectorizer()
X = vec.fit_transform(train_texts)

Labels Stay Separate

Your features X come from the text, but your labels y come from you. Each document needs its correct class, like spam or not-spam.

Split Before You Train

Hold out a test set so you can judge fairly. train_test_split keeps some data unseen until the very end of evaluation.

from sklearn.model_selection import train_test_split
X_tr, X_te, y_tr, y_te = train_test_split(X, y)

Fit Means Learn

Calling fit tells logistic regression to find the word weights that best separate your classes on the training data.

from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=1000)
clf.fit(X_tr, y_tr)

Raise max_iter If It Warns

Text has many features, so the solver may need more steps. Bumping max_iter clears the common convergence warning you will see.

Predict on New Vectors

To classify fresh text, transform it with the same fitted vectorizer, then call predict. Never refit the vectorizer on test data.

preds = clf.predict(X_te)

Transform, Do Not Fit, on Test

Test text uses transform, not fit_transform. Refitting would leak test information and quietly inflate your scores.

Check the Accuracy

A quick score call gives accuracy on the held-out set. It is your first signal that training actually worked.

print(clf.score(X_te, y_te))

Same Steps Stay in Sync

Train and prediction must share the exact same vocabulary. Using one fitted vectorizer everywhere keeps feature columns aligned.

A Pipeline Saves You

Wrapping the vectorizer and model in a Pipeline chains transform and predict automatically, so you never forget a step.

from sklearn.pipeline import make_pipeline
pipe = make_pipeline(TfidfVectorizer(), LogisticRegression())

Quick Check

How should you prepare test text before predicting?

Recap: Vectorize Then Fit

Vectorize text with TF-IDF, split, then fit logistic regression. Reuse the same vectorizer for test data, ideally inside a pipeline. ✅

Preguntas frecuentes

¿La lección «Entrenamiento con características TF-IDF» es gratis?

Sí — el texto completo de «Entrenamiento con características TF-IDF» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de NLP Academy, actualiza a CoddyKit PRO. El curso de NLP Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Entrenamiento con características TF-IDF»?

Entrenar un clasificador en scikit-learn Practicas NLP Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar NLP Academy?

No se requiere experiencia previa. NLP Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.

¿Cuánto tiempo toma la lección «Entrenamiento con características TF-IDF»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de NLP Academy?

Sí. Cada lección de NLP Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Por qué la regresión logística triunfa con texto
  2. Entrenamiento con características TF-IDF
  3. Inspección de los coeficientes más fuertes
  4. Ajuste de la intensidad de la regularización
← Volver a NLP Academy