Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

BLEU and ROUGE are what type of metrics?

BLEU and ROUGE are similarity metrics used to evaluate generated text by comparing it to reference text. BLEU measures how much the candidate text shares exact n-gram sequences with the reference, focusing on precision across n-grams and usually applying a brevity penalty to avoid outputs that are too short. ROUGE emphasizes recall, assessing how much of the reference content appears in the generated text, with variants like ROUGE-N (n-gram overlap) and ROUGE-L (longest common subsequence). These metrics are common in tasks like machine translation and text summarization as automated indicators of quality. They are not fairness metrics (such as demographic parity or predictive rate parity) and they are not perplexity, which gauges how well a model predicts the next word rather than how closely the output matches a reference.

BLEU and ROUGE are similarity metrics used to evaluate generated text by comparing it to reference text. BLEU measures how much the candidate text shares exact n-gram sequences with the reference, focusing on precision across n-grams and usually applying a brevity penalty to avoid outputs that are too short. ROUGE emphasizes recall, assessing how much of the reference content appears in the generated text, with variants like ROUGE-N (n-gram overlap) and ROUGE-L (longest common subsequence). These metrics are common in tasks like machine translation and text summarization as automated indicators of quality. They are not fairness metrics (such as demographic parity or predictive rate parity) and they are not perplexity, which gauges how well a model predicts the next word rather than how closely the output matches a reference.