0Pricing
NLP Academy · درس

حيل تقسيم النص بالـ Regex

قسّم النص بالطريقة التي تريدها تمامًا

حيل تقسيم النص بالـ Regex درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Tokenizing With Patterns

You can tokenize text by describing what a token looks like, then letting regex find every piece that fits your definition.

Match Words Directly

The pattern \w+ with findall grabs every run of word characters, giving you a clean list of words and ignoring the spaces between them.

re.findall("\w+", "Hello, world!")

Split Instead of Match

The re.split function breaks text wherever a pattern appears. Splitting on whitespace turns a sentence into a list of rough tokens fast.

re.split("\s+", "one  two three")

Split on Multiple Separators

With a character class, re.split can break on many separators at once, like spaces, commas, and semicolons in a single sweep.

re.split("[ ,;]+", "a, b;c d")

Keep Punctuation as Tokens

Sometimes punctuation matters. A pattern like \w+|[^\w\s] captures words and lone symbols separately, so nothing is silently dropped.

The Alternation Operator

The pipe means or. The pattern cat|dog matches either word, letting one regex describe several token shapes you care about.

re.findall("cat|dog", "a dog, a cat")

Match Numbers and Decimals

Numbers need their own rule. The pattern \d+\.?\d* matches whole numbers and decimals, so prices and amounts stay together as one token.

re.findall("\d+\.?\d*", "buy 3 for 4.50")

Handle Contractions

Naive splitting wrecks words like do not in its short form. A smarter token pattern can keep apostrophes inside a word where they belong.

Word Boundaries Help

The \b anchor marks a word boundary, the edge between a word and a non-word. It helps you grab whole words without grabbing neighbors.

Build a Token Pattern

A practical tokenizer often combines rules with alternation: match URLs, then numbers, then words, then symbols, in priority order.

Compile for Speed

If you reuse a pattern a lot, re.compile turns it into a reusable object. It reads cleaner and runs faster across many strings.

tok = re.compile("\w+")

Quick Check

You want to break a string anywhere one or more spaces appear. Which function fits best?

Recap: Regex Tokenizers

You used findall, split, alternation, and compile to turn raw text into exactly the tokens you want. Regex gives you full control. 🎯

الأسئلة الشائعة

هل درس «حيل تقسيم النص بالـ Regex» مجاني؟

نعم — نص درس «حيل تقسيم النص بالـ Regex» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.

ماذا ستتعلم في «حيل تقسيم النص بالـ Regex»؟

قسّم النص بالطريقة التي تريدها تمامًا تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟

لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «حيل تقسيم النص بالـ Regex»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟

نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. التعبيرات النمطية في 5 دقائق
  2. العثور على عناوين البريد الإلكتروني وعناوين URL
  3. مجموعات الالتقاط والاستبدالات
  4. حيل تقسيم النص بالـ Regex
← العودة إلى NLP Academy