العد باستخدام CountVectorizer
حوّل المستندات إلى مصفوفة تعداد
العد باستخدام CountVectorizer درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Meet CountVectorizer
The CountVectorizer from scikit-learn builds your vocabulary and counts words for you in just a few lines. 🚀
from sklearn.feature_extraction.text import CountVectorizerCreate the Vectorizer
First make an instance. With no arguments it uses sensible defaults for tokenizing and lowercasing text.
vectorizer = CountVectorizer()Fit Learns the Vocabulary
Calling fit scans your documents and discovers every unique word, building the vocabulary automatically.
vectorizer.fit(docs)Transform Produces Counts
Then transform turns each document into its count vector, giving you a numeric matrix of word frequencies.
X = vectorizer.transform(docs)Fit and Transform Together
The shortcut fit_transform does both steps at once on your training text, which is the common workflow.
X = vectorizer.fit_transform(docs)See the Vocabulary
Inspect the learned words and their indices with get_feature_names_out. These are your matrix columns.
print(vectorizer.get_feature_names_out())Output Is a Sparse Matrix
To save memory, the result is a sparse matrix that stores only the non-zero counts, not every zero slot.
Peek at the Numbers
Convert to a dense array to actually read the counts. Use toarray for small examples only.
print(X.toarray())Lowercasing Is Automatic
By default it lowercases text and splits on word boundaries, so The and the count as the same token.
Built-In Cleaning Options
You can pass stop_words, min_df, or max_features to prune the vocabulary right inside the vectorizer.
CountVectorizer(stop_words="english", max_features=1000)Reuse on New Text
Never refit on test data. Call only transform so new documents use the exact same vocabulary you trained.
X_new = vectorizer.transform(new_docs)Quick Check
Which call learns the vocabulary and counts in one step?
Recap: CountVectorizer
You used CountVectorizer to fit a vocabulary, transform text into counts, inspect features, and reuse it on new data. 🎉
الأسئلة الشائعة
هل درس «العد باستخدام CountVectorizer» مجاني؟
نعم — نص درس «العد باستخدام CountVectorizer» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.
ماذا ستتعلم في «العد باستخدام CountVectorizer»؟
حوّل المستندات إلى مصفوفة تعداد تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟
لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «العد باستخدام CountVectorizer»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟
نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- لماذا تحتاج النماذج إلى أرقام لا كلمات
- إنشاء مفردات
- العد باستخدام CountVectorizer
- قراءة مصفوفة المستندات والمصطلحات