0Pricing
NLP Academy · درس

تقسيم النص إلى رموز باستخدام NLTK

استخدم أداة تقسيم حقيقية للنصوص غير المنتظمة

تقسيم النص إلى رموز باستخدام NLTK درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

A Real Tokenizer

Time to upgrade. NLTK is a classic Python library that gives you a proper tokenizer for messy, real-world text. 🛠️

Install and Import

You install it once with pip, then import it in your script. From there, NLTK's word_tokenize is one function call away.

from nltk import word_tokenize

One Time Setup

NLTK ships extra data separately. Before tokenizing, you download the punkt model once, and then it just works.

import nltk
nltk.download("punkt")

Tokenize a Sentence

Now pass any string to word_tokenize. It returns a clean list of word and punctuation tokens, ready to count or filter.

word_tokenize("I love cats!")
# ['I', 'love', 'cats', '!']

Punctuation Split Out

Notice the win: the exclamation mark is now its own token. So cats and cats! finally count as the same word.

Contractions Handled

NLTK is smart about contractions. It splits don't into do and n't, keeping the hidden negation visible to your code.

word_tokenize("don't")
# ['do', "n't"]

Why It Is Smarter

Under the hood, NLTK follows linguistic rules learned from real text. That is why it beats a plain whitespace split every time.

Sentences Too

NLTK also segments sentences. Pair sent_tokenize with word_tokenize to split a document into sentences, then each into words.

from nltk import sent_tokenize
sent_tokenize("Hi there. Bye now.")

Combine the Two

A common pattern loops over sentences and tokenizes each. This gives you a tidy list of lists, one token list per sentence.

Not the Only Option

NLTK is great for learning, but it is not alone. Libraries like spaCy offer faster tokenizers you will meet later on.

From Raw Text to Tokens

You now have the full move: raw text in, a clean token list out. This is the foundation every later NLP step builds on.

Quick Check

How does NLTK improve on naive splitting?

Recap

You used NLTK to tokenize real text: install, download punkt, then call word_tokenize. It splits punctuation and contractions cleanly for you.

الأسئلة الشائعة

هل درس «تقسيم النص إلى رموز باستخدام NLTK» مجاني؟

نعم — نص درس «تقسيم النص إلى رموز باستخدام NLTK» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.

ماذا ستتعلم في «تقسيم النص إلى رموز باستخدام NLTK»؟

استخدم أداة تقسيم حقيقية للنصوص غير المنتظمة تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟

لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «تقسيم النص إلى رموز باستخدام NLTK»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟

نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. ما الرمز المميّز فعلًا؟
  2. التقسيم عند المسافات البيضاء وحدوده
  3. أساسيات تقسيم الجمل
  4. تقسيم النص إلى رموز باستخدام NLTK
← العودة إلى NLP Academy