0Pricing
NLP Academy · درس

قراءة مصفوفة المستندات والمصطلحات

افهم الصفوف والأعمدة والتخلخل

قراءة مصفوفة المستندات والمصطلحات درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

What Is the DTM?

A document-term matrix is a grid of counts. Each row is a document and each column is a word from your vocabulary. 🧮

Rows Are Documents

One row holds the full count vector for a single document, summarizing how many times each word appeared in it.

Columns Are Terms

Each column tracks one word across every document, so reading down a column shows where that word shows up.

A Cell Is a Count

The value at row i, column j is how many times word j appeared in document i, a single count.

Check the Shape

The matrix shape tells you how many documents and how many unique words you are working with.

print(X.shape)  # (n_documents, n_words)

Most Cells Are Zero

Any single document uses only a few of the thousands of words, so the matrix is mostly zeros, called sparse.

Why Sparsity Matters

Storing only non-zero values keeps huge matrices in memory. This is why scikit-learn returns a sparse format.

Label the Columns

Pair the array with feature names to make it readable. A DataFrame turns raw counts into a clear table.

import pandas as pd
df = pd.DataFrame(X.toarray(), columns=names)

Read One Document

Look at a single row to see that document profile: which words it uses and how often. That row is its fingerprint.

Compare Two Documents

Similar documents have similar rows. Comparing vectors lets you measure how close two texts are in content.

From Matrix to Model

This matrix is the input you feed a classifier. The DTM is the features, and your labels are the targets.

Quick Check

In a document-term matrix, what does one cell hold?

Recap: The DTM

You read the document-term matrix: rows as documents, columns as words, cells as counts, and saw why it is sparse. 🎉

الأسئلة الشائعة

هل درس «قراءة مصفوفة المستندات والمصطلحات» مجاني؟

نعم — نص درس «قراءة مصفوفة المستندات والمصطلحات» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.

ماذا ستتعلم في «قراءة مصفوفة المستندات والمصطلحات»؟

افهم الصفوف والأعمدة والتخلخل تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟

لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «قراءة مصفوفة المستندات والمصطلحات»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟

نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. لماذا تحتاج النماذج إلى أرقام لا كلمات
  2. إنشاء مفردات
  3. العد باستخدام CountVectorizer
  4. قراءة مصفوفة المستندات والمصطلحات
← العودة إلى NLP Academy