دمج أنواع متعددة من السمات
اجمع السمات النصية والرقمية
دمج أنواع متعددة من السمات درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
One Type Is Rarely Enough
TF-IDF captures words, your hand-built numbers capture style. The best models often combine several feature types into one input.
Text and Numbers Together
Imagine a review's TF-IDF vector plus its length and star rating. Stacking these gives the model both word and numeric signal at once.
The Shape Problem
TF-IDF outputs a sparse matrix, while your custom features are a small dense array. You must join them along the same rows.
Side by Side, Not Stacked
We glue features as new columns for the same documents, not new rows. This horizontal join is called concatenation.
Stacking Sparse Matrices
SciPy offers hstack to place matrices side by side efficiently. It keeps everything sparse, so memory stays under control.
from scipy.sparse import hstack
combined = hstack([tfidf_matrix, numeric_features])ColumnTransformer to the Rescue
scikit-learn's ColumnTransformer applies different steps to different columns. It is the clean way to route text and numbers through one pipeline.
Wiring It Up
You give ColumnTransformer a list of named transformers and the columns each one handles. It builds a single feature matrix for you.
from sklearn.compose import ColumnTransformer
ct = ColumnTransformer([("text", TfidfVectorizer(), "review"), ("num", "passthrough", ["length"])])FeatureUnion for Parallel Steps
When several transformers read the same input, FeatureUnion runs them in parallel and joins the outputs. It is built for combining feature extractors.
Mind the Scales
Raw word counts and a 0-to-1000 length live on very different ranges. Mixing them often calls for scaling so no feature dominates.
Let Data Decide
Adding feature types should be a measured experiment. Compare a validation score before and after to confirm the combo actually helps.
More Is Not Always Better
Throwing in every feature can add noise and slow training. Aim for a small, well-chosen mix over a giant kitchen sink. 🧹
Quick Check
How should TF-IDF and numeric features be merged?
Recap
Strong models combine text and numeric features by concatenating columns. Use hstack, ColumnTransformer, or FeatureUnion, and scale before mixing. ✅
الأسئلة الشائعة
هل درس «دمج أنواع متعددة من السمات» مجاني؟
نعم — نص درس «دمج أنواع متعددة من السمات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.
ماذا ستتعلم في «دمج أنواع متعددة من السمات»؟
اجمع السمات النصية والرقمية تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟
لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «دمج أنواع متعددة من السمات»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟
نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- ما بعد حقيبة الكلمات
- N-Gram للمحارف من أجل المتانة
- دمج أنواع متعددة من السمات
- توسيع نطاق السمات واختيارها