تقييم التوليد باستخدام ROUGE وBLEU
قِس جودة التلخيص والترجمة
تقييم التوليد باستخدام ROUGE وBLEU درس مجاني في NLP Academy على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في NLP Academy، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة NLP Academy 4 دروس في المجموع.
بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.
Why Score Generation?
Summaries and translations are free-form text, so there is no single right answer. We need metrics to compare output against reference text. 📏
Compare to a Reference
Both ROUGE and BLEU compare your generated text to one or more human references. More overlap with the reference means a higher score.
BLEU for Translation
BLEU measures how many short word sequences in your output also appear in the reference. It is the classic metric for machine translation.
BLEU Rewards Precision
BLEU is precision-focused: it asks how much of your output matches the reference. A brevity penalty stops models from cheating with very short text.
ROUGE for Summaries
ROUGE is the go-to metric for summarization. It checks how much of the reference content your summary managed to recover.
ROUGE Rewards Recall
ROUGE leans on recall: did your summary capture the important words from the reference? ROUGE-1 counts single-word overlap.
ROUGE-L and Sequences
ROUGE-L looks at the longest matching word sequence, rewarding output that keeps the reference's order, not just its words.
Computing ROUGE
The evaluate library makes scoring a one-liner. Load rouge and pass your predictions with their references.
import evaluate
rouge = evaluate.load("rouge")
print(rouge.compute(predictions=preds, references=refs))Computing BLEU
BLEU works the same way. Note each prediction needs a list of references, since several translations can be correct.
bleu = evaluate.load("bleu")
print(bleu.compute(predictions=preds, references=refs))Higher Is Better
Both scores rise with overlap. They are useful for comparing models, but a single absolute number means little on its own.
Metrics Miss Meaning
These metrics count word overlap, not true meaning. A perfect paraphrase using different words can still score low, so pair them with human review.
Quick Check
Which metric is recall-focused and standard for summarization?
Recap
You learned to score generated text: BLEU for translation precision, ROUGE for summary recall, and why both still need human judgment. 🎯
الأسئلة الشائعة
هل درس «تقييم التوليد باستخدام ROUGE وBLEU» مجاني؟
نعم — نص درس «تقييم التوليد باستخدام ROUGE وBLEU» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة NLP Academy، انتقل إلى CoddyKit PRO. تتضمن دورة NLP Academy 4 دروس في المجموع.
ماذا ستتعلم في «تقييم التوليد باستخدام ROUGE وBLEU»؟
قِس جودة التلخيص والترجمة تتمرن على NLP Academy مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.
هل أحتاج إلى خبرة سابقة لأبدأ NLP Academy؟
لا تُشترط خبرة سابقة. NLP Academy على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.
كم من الوقت يستغرق درس «تقييم التوليد باستخدام ROUGE وBLEU»؟
معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.
هل يمكنني كتابة وتشغيل أكواد في درس NLP Academy هذا؟
نعم. كل درس في NLP Academy يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.
جميع الدروس في هذه الدورة
- الملخصات الاستخراجية مقابل التلخيصية
- التلخيص باستخدام نموذج Seq2Seq
- الترجمة الآلية عمليًا
- تقييم التوليد باستخدام ROUGE وBLEU