مسارات scikit-learn من البداية إلى النهاية
صِل أداة تحويل النص إلى متجهات بالنموذج بسلاسة
مسارات scikit-learn من البداية إلى النهاية درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 2 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
What Is a Pipeline?
A pipeline chains your steps into one object: clean, vectorize, then classify. Call it once and every step runs in the right order. 🔗
Why Glue Steps Together
Doing vectorizing and training by hand invites bugs. A pipeline keeps the steps locked together so they never drift out of sync.
Import the Pieces
You need a vectorizer and a classifier. Import the Pipeline class plus the two components you want to chain.
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegressionBuild the Pipeline
List your steps as named pairs. The first transforms text into features, the last makes the prediction.
pipe = Pipeline([
("tfidf", TfidfVectorizer()),
("clf", LogisticRegression())
])Fit on Raw Text
Call fit with your raw text and labels. The pipeline vectorizes and trains in a single line, no manual steps.
pipe.fit(X_train, y_train)Predict in One Call
To classify new text, call predict. The same vectorizer is reused automatically, so train and test stay perfectly consistent.
preds = pipe.predict(X_test)No Data Leakage
Pipelines learn the vocabulary only from training data. This stops data leakage, where test info sneaks into training and inflates your score.
Score the Pipeline
Use score to get accuracy on held-out data in one call. The pipeline handles vectorizing the test text for you.
acc = pipe.score(X_test, y_test)
print(acc)Tune the Whole Chain
Grid search can tune any step. Name parameters with the step name plus double underscore to reach inside a component.
params = {"tfidf__ngram_range": [(1,1), (1,2)],
"clf__C": [0.1, 1, 10]}Cross-Validate Safely
Pass the whole pipeline into cross-validation. Each fold re-fits the vectorizer, so every score reflects truly unseen data.
from sklearn.model_selection import cross_val_score
scores = cross_val_score(pipe, X, y, cv=5)One Object to Ship
The best part: a fitted pipeline is a single object. Save and load it as one unit, and inference matches training exactly.
Quick Check
Think about the main safety benefit a pipeline gives you.
Recap
You chained a vectorizer and classifier into one pipeline, fit on raw text, predicted in a line, tuned the chain, and avoided leakage. 🎉
الأسئلة الشائعة
هل درس «مسارات scikit-learn من البداية إلى النهاية» مجاني؟
نعم — نص درس «مسارات scikit-learn من البداية إلى النهاية» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.
ماذا ستتعلم في «مسارات scikit-learn من البداية إلى النهاية»؟
صِل أداة تحويل النص إلى متجهات بالنموذج بسلاسة تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟
لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 2 من أصل 4.
كم من الوقت يستغرق درس «مسارات scikit-learn من البداية إلى النهاية»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟
نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- هيكلة مشروع NLP حقيقي
- مسارات scikit-learn من البداية إلى النهاية
- حفظ النموذج وتحميله
- التنبؤ على نص جديد تمامًا