سمات N-Gram في scikit-learn
أضف العبارات إلى أداة تحويل النص إلى متجهات
سمات N-Gram في scikit-learn درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 3 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Let the Library Do It
You could hand-roll n-grams, but scikit-learn builds them for you. The trick is one small argument on your vectorizer.
The ngram_range Knob
Every text vectorizer accepts ngram_range, a tuple of min and max sizes. It tells the vectorizer which n-grams to count.
from sklearn.feature_extraction.text import CountVectorizer
vec = CountVectorizer(ngram_range=(1, 2))Unigrams Are the Default
Leave it alone and ngram_range is (1, 1), pure single words. That is the plain bag-of-words you already know.
Adding Bigrams
Set ngram_range to (1, 2) and the vectorizer keeps single words and every adjacent pair, all in one feature space.
vec = CountVectorizer(ngram_range=(1, 2))
X = vec.fit_transform(["this is not good"])Bigrams Only
Want pairs alone? Use (2, 2). Now single words vanish and only bigrams survive as features.
vec = CountVectorizer(ngram_range=(2, 2))Inspect the Vocabulary
After fitting, peek at the learned features with get_feature_names_out. You will see both words and joined phrases listed.
print(vec.get_feature_names_out())
# ['is not', 'not good', 'this is']Phrases Joined by Space
scikit-learn names each bigram by joining its words with a single space, like not good. That string becomes one column.
Same Knob, TF-IDF
The same ngram_range works on TfidfVectorizer too. You get n-gram phrases that are also weighted by how distinctive they are.
from sklearn.feature_extraction.text import TfidfVectorizer
vec = TfidfVectorizer(ngram_range=(1, 2))Now Negation Survives
With bigrams on, not good lands in its own column. Your classifier can finally learn that this pair signals a negative review.
Watch the Feature Count
Adding bigrams can multiply your columns dramatically. The matrix stays sparse, but the vocabulary grows fast.
print(X.shape) # many more columns than unigrams alonePair It With min_df
To tame the blow-up, combine ngram_range with min_df. Dropping rare phrases keeps only the n-grams that repeat usefully.
vec = CountVectorizer(ngram_range=(1, 2), min_df=2)Quick Check
Which ngram_range gives you single words plus adjacent pairs in one vectorizer?
Recap: N-Grams in scikit-learn
One ngram_range tuple turns any vectorizer into an n-gram machine. Inspect features, then trim with min_df. Next you will choose the right range. 🎯
الأسئلة الشائعة
هل درس «سمات N-Gram في scikit-learn» مجاني؟
نعم — نص درس «سمات N-Gram في scikit-learn» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.
ماذا ستتعلم في «سمات N-Gram في scikit-learn»؟
أضف العبارات إلى أداة تحويل النص إلى متجهات تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟
لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 3 من أصل 4.
كم من الوقت يستغرق درس «سمات N-Gram في scikit-learn»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟
نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- لماذا تفقد الكلمات المفردة معناها
- شرح الثنائيات والثلاثيات
- سمات N-Gram في scikit-learn
- اختيار نطاق N-Gram المناسب