0Pricing
NLP Academy · Lezione

Valutare la generazione con ROUGE e BLEU

Misurare la qualità di riepiloghi e traduzioni

Valutare la generazione con ROUGE e BLEU è una lezione NLP Academy gratuita su CoddyKit. Questa è la lezione 4 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento NLP Academy, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso NLP Academy include 4 lezioni in totale.

Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.

Why Score Generation?

Summaries and translations are free-form text, so there is no single right answer. We need metrics to compare output against reference text. 📏

Compare to a Reference

Both ROUGE and BLEU compare your generated text to one or more human references. More overlap with the reference means a higher score.

BLEU for Translation

BLEU measures how many short word sequences in your output also appear in the reference. It is the classic metric for machine translation.

BLEU Rewards Precision

BLEU is precision-focused: it asks how much of your output matches the reference. A brevity penalty stops models from cheating with very short text.

ROUGE for Summaries

ROUGE is the go-to metric for summarization. It checks how much of the reference content your summary managed to recover.

ROUGE Rewards Recall

ROUGE leans on recall: did your summary capture the important words from the reference? ROUGE-1 counts single-word overlap.

ROUGE-L and Sequences

ROUGE-L looks at the longest matching word sequence, rewarding output that keeps the reference's order, not just its words.

Computing ROUGE

The evaluate library makes scoring a one-liner. Load rouge and pass your predictions with their references.

import evaluate
rouge = evaluate.load("rouge")
print(rouge.compute(predictions=preds, references=refs))

Computing BLEU

BLEU works the same way. Note each prediction needs a list of references, since several translations can be correct.

bleu = evaluate.load("bleu")
print(bleu.compute(predictions=preds, references=refs))

Higher Is Better

Both scores rise with overlap. They are useful for comparing models, but a single absolute number means little on its own.

Metrics Miss Meaning

These metrics count word overlap, not true meaning. A perfect paraphrase using different words can still score low, so pair them with human review.

Quick Check

Which metric is recall-focused and standard for summarization?

Recap

You learned to score generated text: BLEU for translation precision, ROUGE for summary recall, and why both still need human judgment. 🎯

Domande Frequenti

La lezione «Valutare la generazione con ROUGE e BLEU» è gratuita?

Sì — il testo completo di «Valutare la generazione con ROUGE e BLEU» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso NLP Academy, passa a CoddyKit PRO. Il corso NLP Academy include 4 lezioni in totale.

Cosa imparerò in «Valutare la generazione con ROUGE e BLEU»?

Misurare la qualità di riepiloghi e traduzioni Eserciti NLP Academy con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.

Ho bisogno di esperienza per iniziare NLP Academy?

Non è richiesta alcuna esperienza precedente. NLP Academy su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 4 di 4.

Quanto tempo richiede la lezione «Valutare la generazione con ROUGE e BLEU»?

La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.

Posso scrivere ed eseguire codice in questa lezione NLP Academy?

Sì. Ogni lezione NLP Academy include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.

Tutte le lezioni di questo corso

  1. Riepiloghi estrattivi e astrattivi
  2. Riassumere con un modello Seq2Seq
  3. La traduzione automatica nella pratica
  4. Valutare la generazione con ROUGE e BLEU
← Torna a NLP Academy