Skalowanie i wybieranie cech
Zachowaj cechy, które rzeczywiście pomagają
Skalowanie i wybieranie cech to bezpłatna lekcja NLP Academy na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej NLP Academy, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs NLP Academy zawiera 4 lekcji w sumie.
Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.
Too Many Features Hurt
Text models can explode to tens of thousands of features. Many add only noise, so trimming the list often improves both speed and accuracy.
Why Scale at All
Some models compare feature magnitudes directly. If one feature ranges 0 to 1000, it can drown out the rest unless you scale them first.
Standardizing Numbers
StandardScaler shifts each feature to zero mean and unit variance. Now every numeric feature speaks on the same scale.
from sklearn.preprocessing import StandardScaler
scaled = StandardScaler().fit_transform(numeric_features)Careful With Sparse Data
Centering a sparse TF-IDF matrix fills it with nonzeros and wastes memory. Use MaxAbsScaler, which scales without destroying sparsity.
from sklearn.preprocessing import MaxAbsScaler
scaled = MaxAbsScaler().fit_transform(tfidf_matrix)Drop the Dead Weight
Features that barely vary tell the model nothing. A VarianceThreshold filter removes near-constant columns in one quick pass.
Pick the Most Useful
SelectKBest keeps the top features by a scoring test like chi-squared. You ask for the best k and it drops the rest.
from sklearn.feature_selection import SelectKBest, chi2
best = SelectKBest(chi2, k=2000).fit_transform(X, y)Let the Model Choose
An L1-penalized model pushes useless weights to zero on its own. This built-in selection is often called embedded feature selection.
Fit on Train Only
Always fit scalers and selectors on training data alone, then apply them to test data. Mixing them causes leakage and rosy fake scores.
Keep It in a Pipeline
Putting scaling and selection inside a Pipeline runs them in the right order every time. It also blocks accidental leakage during validation.
from sklearn.pipeline import Pipeline
pipe = Pipeline([("select", SelectKBest(chi2, k=2000)), ("clf", model)])Fewer Features, Faster Model
A leaner feature set trains faster, needs less memory, and is easier to explain. Smaller is often better once noise is gone. ⚡
Measure, Do Not Guess
Try a few values of k and compare validation scores. Let the numbers, not a hunch, decide how many features to keep.
Quick Check
Which scaler keeps a sparse matrix sparse?
Recap
Scale numeric features and select the useful ones with SelectKBest or L1. Use MaxAbsScaler for sparse data, and fit everything inside a pipeline. ✅
Często zadawane pytania
Czy lekcja „Skalowanie i wybieranie cech” jest bezpłatna?
Tak — pełny tekst „Skalowanie i wybieranie cech” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu NLP Academy, przejdź na CoddyKit PRO. Kurs NLP Academy zawiera 4 lekcji w sumie.
Co nauczysz się w „Skalowanie i wybieranie cech”?
Zachowaj cechy, które rzeczywiście pomagają Ćwiczysz NLP Academy z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć NLP Academy?
Nie wymagamy żadnego doświadczenia. NLP Academy w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Skalowanie i wybieranie cech”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji NLP Academy?
Tak. Każda lekcja NLP Academy zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Poza bag-of-words
- N-gramy znakowe dla większej odporności
- Łączenie wielu typów cech
- Skalowanie i wybieranie cech