0Pricing
NLP Academy · 课时

scikit-learn 中的 N-Gram 特征

将短语添加到向量化器中

scikit-learn 中的 N-Gram 特征 是 CoddyKit 上的免费 NLP Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 NLP Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 NLP Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Let the Library Do It

You could hand-roll n-grams, but scikit-learn builds them for you. The trick is one small argument on your vectorizer.

The ngram_range Knob

Every text vectorizer accepts ngram_range, a tuple of min and max sizes. It tells the vectorizer which n-grams to count.

from sklearn.feature_extraction.text import CountVectorizer
vec = CountVectorizer(ngram_range=(1, 2))

Unigrams Are the Default

Leave it alone and ngram_range is (1, 1), pure single words. That is the plain bag-of-words you already know.

Adding Bigrams

Set ngram_range to (1, 2) and the vectorizer keeps single words and every adjacent pair, all in one feature space.

vec = CountVectorizer(ngram_range=(1, 2))
X = vec.fit_transform(["this is not good"])

Bigrams Only

Want pairs alone? Use (2, 2). Now single words vanish and only bigrams survive as features.

vec = CountVectorizer(ngram_range=(2, 2))

Inspect the Vocabulary

After fitting, peek at the learned features with get_feature_names_out. You will see both words and joined phrases listed.

print(vec.get_feature_names_out())
# ['is not', 'not good', 'this is']

Phrases Joined by Space

scikit-learn names each bigram by joining its words with a single space, like not good. That string becomes one column.

Same Knob, TF-IDF

The same ngram_range works on TfidfVectorizer too. You get n-gram phrases that are also weighted by how distinctive they are.

from sklearn.feature_extraction.text import TfidfVectorizer
vec = TfidfVectorizer(ngram_range=(1, 2))

Now Negation Survives

With bigrams on, not good lands in its own column. Your classifier can finally learn that this pair signals a negative review.

Watch the Feature Count

Adding bigrams can multiply your columns dramatically. The matrix stays sparse, but the vocabulary grows fast.

print(X.shape)  # many more columns than unigrams alone

Pair It With min_df

To tame the blow-up, combine ngram_range with min_df. Dropping rare phrases keeps only the n-grams that repeat usefully.

vec = CountVectorizer(ngram_range=(1, 2), min_df=2)

Quick Check

Which ngram_range gives you single words plus adjacent pairs in one vectorizer?

Recap: N-Grams in scikit-learn

One ngram_range tuple turns any vectorizer into an n-gram machine. Inspect features, then trim with min_df. Next you will choose the right range. 🎯

常见问题解答

「scikit-learn 中的 N-Gram 特征」课时是免费的吗?

是的 — 「scikit-learn 中的 N-Gram 特征」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 NLP Academy 课程的其余内容,请升级到 CoddyKit PRO。 NLP Academy 课程共包含 4 节课。

「scikit-learn 中的 N-Gram 特征」这节课中我会学到什么?

将短语添加到向量化器中 你通过在浏览器中直接运行的动手代码来练习 NLP Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 NLP Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 NLP Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「scikit-learn 中的 N-Gram 特征」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 NLP Academy 课中编写并运行代码吗?

能。每节 NLP Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 单个词语为何会失去意义
  2. 详解二元词组与三元词组
  3. scikit-learn 中的 N-Gram 特征
  4. 选择合适的 N-Gram 范围
← 返回 NLP Academy